The next phase of AI infrastructure focuses on rack-scale system composition, optimizing different compute resources for various phases of agentic AI workflows, moving beyond a narrow focus on accelerator scale. This shift requires specialized CPUs for prefill, decode, and agent orchestration, as inference evolves from single-pass model calls to multi-step agentic pipelines. Arm's architectural strengths in heterogeneous compute, high concurrency, and power-constrained scaling are becoming particularly relevant for this new system-level approach.
AI infrastructure spending will increasingly prioritize system-level orchestration and specialized CPU capacity over raw accelerator count, shifting value to heterogeneous rack designs.