The cost of GPT-4-class model output has fallen 55x to $0.40 per million tokens since late 2022, driven by models like DeepSeek R1 offering 97% discounts against OpenAI's o1-preview. Conversely, frontier models like OpenAI's GPT-5.5 doubled in price to $5 input and $30 output per million tokens, with Google's Gemini Flash 3.5 and Anthropic's Claude Sonnet 5 also increasing effective costs. Companies are now spending 10-20% of labor costs on tokens for agentic tasks, with 15-30% of users accounting for over 50% of total AI spend, often without correlating to productivity gains.
AI inference costs are splitting, with commodity models driving down basic task costs while complex agentic workloads push up spending on frontier models, impacting enterprise balance sheets.