Perplexity AI CEO Aravind Srinivas predicts a Fable 5-quality model will be 3-4 times less expensive in under six months, with an Opus 4.8-grade model running locally within 12 months. This follows a Stanford report finding GPT-3.5-level inference costs have fallen 280 times in under two years, suggesting a Jevons paradox for intelligence.
Rapidly falling inference costs and local model deployment will drive a surge in multi-agent system adoption, shifting compute demand patterns.