The new AA-Briefcase benchmark evaluates models on long-horizon knowledge work tasks spanning multi-week projects and thousands of input files. It highlights that while per-token intelligence costs are declining, the overall cost of completing complex agentic tasks is increasing due to larger context requirements, more turns, and increased tool calls. Key factors driving this rising cost per task include token price, the number of turns, token efficiency, and prompt caching hit rates.
The rising cost of complex agentic tasks, despite falling token prices, creates a demand-pull for more efficient models and compute infrastructure.