Here is the uncomfortable number: 79% of companies overshot their AI budget in the past year, and research puts 40-60% of typical token spend down as pure waste. Even Uber wasn't immune - it burned through its entire 2026 AI budget in four months before it put any limits in place. The good news is that the waste follows the same pattern almost everywhere, which means the fix does too.
Companies that get their bill down 40% or more do not use AI less. They make three moves.
Move 1: Put the right model on each job
This is the biggest lever. AI models come in sizes, and the price gap between sizes is huge - often 10-30x per token. Most teams run everything on a premium model because it was the default when they signed up.
But most everyday work - replies, summaries, drafts, formatting - comes out just as well on a lighter model. Industry benchmarks show routing routine work to right-sized models saves 30-50% on its own, with no visible quality drop.
Move 2: Stop paying for the same answer twice
Teams repeat themselves constantly: the same product questions, the same email formats, the same report structure, week after week. Every repeat is billed like it is brand new.
The fix is reuse. Shared prompts, saved answers, and caching mean a repeated question costs a fraction of the original. On true repeats, caching saves up to 90%. Across a whole company, this is usually worth another 10-15% off the bill.
Move 3: Retire the seats nobody uses
Seat licenses renew silently. The trial seats from last quarter, the person who left, the tool a team stopped using - all still billing. It is the least glamorous fix and the fastest one: find seats with no activity in 30+ days and cancel or reassign them. For most companies this is another 5-15%.
The order matters: measure first
Every successful cost cut starts the same way - with visibility. You cannot right-size models if you do not know which models are being used, and you cannot retire idle seats you cannot see.
- Week 1: connect your AI accounts, see actual usage per tool, per team, per person.
- Week 2: apply the obvious fixes - idle seats and model routing on one pilot team.
- Week 3-4: roll out to everyone, keep the monitoring on so the waste does not grow back.
What not to do
Do not ban tools or ration usage. AI genuinely saves your team hours; cutting usage to cut cost throws out the value with the waste. The goal is the same work at a smaller price, and that is exactly what the three moves deliver.
