Uber Cuts AI Costs 52% Per Session Despite 9x Usage Surge
The ride-hailing giant stabilized spending through model routing, token caps, and open-weight alternatives after exhausting its Q1 budget.
Uber has achieved a dramatic reduction in AI costs even as internal usage has exploded, demonstrating how large enterprises can scale artificial intelligence deployments without proportional budget increases.
Weekly AI agent requests at Uber have grown 9.4 times since February, yet the company's total AI spending has remained flat since April, according to Uday Kiran Medisetty, a distinguished engineer at Uber. Cost per 1,000 requests dropped nearly 34% from the April peak, while cost per session fell 52% from June highs, details first reported by Axios.
Why it matters
Uber's experience offers a roadmap for enterprises struggling with unpredictable AI expenses. The company's ability to maintain stable costs while dramatically increasing usage challenges the assumption that AI adoption necessarily means runaway spending, and shows that architectural choices matter as much as raw model performance.
The cost control playbook
Uber deployed several technical strategies to rein in expenses. The company now uses a router system that directs tasks to the most cost-effective model capable of handling them, reserving expensive frontier models only for complex work. Simpler tasks flow to less expensive alternatives.
The engineering team also implemented a 400,000-token cap on interactive sessions, even when using models that support up to one million tokens. This prevents unnecessary consumption during typical development workflows.
Visibility proved crucial: engineers can now see real-time costs for their AI sessions directly in their terminals, creating immediate feedback on spending patterns. Uber also extended its prompt cache duration from five minutes to one hour after observing that engineers frequently leave sessions idle for longer periods.
Scale without proportional spending
The results demonstrate significant efficiency gains. AI agents now generate more than 70% of code-change submissions at Uber. Engineers run over 30,000 AI agent tasks daily, and the number of employees using AI tools has more than quadrupled. Despite this expansion, token costs declined.
Uber CTO Praveen Neppalli pointed to experimentation with open-weight models as another factor driving down costs. "We continuously evaluate new models and deploy the best option for each use case," he noted on X.
From budget crisis to optimization
The turnaround marks a sharp reversal from April, when Neppalli told The Information that Uber had exhausted its entire 2026 AI budget in the first quarter. By August, he declared the company was moving past the "tokenmaxxing era" of unconstrained AI spending.
The shift illustrates how quickly enterprise AI economics are evolving as companies move from experimental deployments to production-scale implementations with cost discipline.
These details were reported by Madison Mills for Axios.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

