VentureBeat described on 7 September 2026 a problem that has already earned its own name: tokenmaxxing, meaning surging token consumption without the ROI to justify it. The catch is that almost nobody can demonstrate what the spending buys.
Key takeaways
- Gartner forecast: $207 billion on agentic software in 2026, a 139% increase
- Uber exhausted its annual AI coding budget by April 2026, four months after rollout
- Uber's new cap: $1,500 per employee per month
- Everlaw: a project costing $3,500 in tokens cut delivery from 9.5 to 2.5 engineer-months
- Routing layers are already offered by Databricks, Amazon Bedrock, Azure AI Foundry and Merge
Uber burned a year's budget in four months
Uber gave engineers access to Claude Code in December 2025. By April 2026 the entire annual AI coding budget was gone. Token consumption climbed, but it could not be tied to a better product.
There was no link yet between that overblown usage and shipping actually better products.
Andrew Macdonald, President and Chief Operating Officer at Uber.
Everlaw shows the other side
The counterexample is Everlaw, where the arithmetic closes. Max Christoff, its chief technology officer, cites a Java infrastructure project: $3,500 spent on tokens cut delivery from 9.5 to 2.5 engineer-months?Engineer-month: A unit of effort: one engineer working for one month. It allows projects with different team sizes to be compared on the same scale.. On a larger product the cost reached $27,000 to $40,000, but scope fell from 90–100 engineer-months to 19.
| Dimension | Uber | Everlaw |
|---|---|---|
| Scope | company-wide rollout from December 2025 | individual engineering projects |
| Baseline set in advance | no | yes |
| Token cost | annual budget exhausted by April 2026 | $3,500 and $27,000–40,000 |
| Outcome | no demonstrated link to product quality | 9.5 → 2.5 and 90–100 → 19 engineer-months |
| Response | $1,500 per employee per month cap | — |
The difference is not the tool. It is that someone set a baseline in advance. Without one, a token invoice is undecidable.
Symbol meaning
- …
- the project's actual saving
- …
- estimated cost of the task without AI, fixed before work starts
- …
- the bill for tokens consumed
- …
- cost of the remaining engineering work
The routing layer as an answer
The most common proposal is model routing — matching the model and the harness to the task, instead of pushing everything through the most expensive one available.
How a routing layer works
Databricks shipped Smart Routing in its Unity AI Gateway, promising automatic selection of model and harness for each coding task. Comparable mechanisms exist in Amazon Bedrock and Azure AI Foundry, and Merge competes in the same space.
Why it matters
Uber and Everlaw used the same tools and reached opposite conclusions, because only one of them established beforehand what the task would have cost without AI.
At a forecast of $207 billion this stops being financial hygiene and becomes the condition for keeping budgets next year.
What's next
- Uber's $1,500 per-employee cap is a hard test — it will show whether limiting usage lowers quality or only the bill
- Gartner's forecast assumes 139% year-on-year growth, so pressure to prove returns rises faster than the spending itself
- None of the routing vendors described here publishes comparable data on real customer savings yet
Sources
- VentureBeat — Companies are spending millions rewiring how AI gets used. Almost none can prove it's working
- Databricks Documentation — Route coding tasks with Smart Routing





