
AI token prices have fallen sharply, but companies are spending more as complex agentic tasks consume vastly more tokens. Coforge says unit costs dropped about 100-fold in seven months while token use…
AI token prices have fallen sharply, but companies are spending more as complex agentic tasks consume vastly more tokens. Coforge says unit costs dropped about 100-fold in seven months while token use rose 8,000-fold. Open-weight models now account for 35% of token consumption on its platforms, up from 11-12% in early 2025.

Clients are exhausting annual token budgets in months or even days, forcing IT firms to rethink pricing. Coforge advocates measuring cost per unit of work rather than raw token cost, using a mix of models. Hexaware now includes token-based pricing in every proposal. McKinsey's survey found 93% of enterprises exceeded AI budgets, and Menlo Ventures data shows enterprise LLM spending tripled in 12 months to end-2025. Gartner forecasts agentic workflow inference costs will rise more than fivefold through 2028.
The two reportings align on the core facts: token costs are plummeting while consumption is exploding. The Hindu foregrounds the structural economics of the AI industry, centering on Coforge's unit-cost framework and broader OECD/Gartner data. Livemint is client-side and brokerage-focused, foregrounding the budget crisis and new pricing mechanisms from both Coforge and Hexaware. Neither is partisan or critical, both are neutral industry trend reports. The balanced reading is clear: what looks like deflation is concentrating into higher per-task costs as complexity grows. The real pressure point is the budget ceiling: every client that blew its annual token plan in days is now forcing IT firms to redesign pricing.
Coverage: 2 sources, 2 neutral
Sources (2): thehindu.com (neutral report), livemint.com (neutral report)
This brief was synthesised by AI from the 2 sources linked above, so one read covers every framing they carry.