The technical partnership combines AMD Helios rackscale solutions with the Cerebras Wafer-Scale Engine to optimize disaggregated inference workflows. This joint solution is expected to deliver up to 5x higher tokens per second per watt (T/s/W) for latency-sensitive applications. Cerebras plans to deploy AMD Helios in its data centers, with initial availability through Cerebras Cloud in the second half of 2026.
AMD and Cerebras are creating a specialized inference platform for real-time agentic AI, addressing the growing demand for heterogeneous compute.