Here is the uncomfortable number: 79% of companies overshot their AI budget in the past year, and research puts 30-40% of typical token spend down as pure waste. Even Uber wasn't immune - it burned through its entire 2026 AI budget in four months before it put any limits in place. The good news is that the waste follows the same pattern almost everywhere, which means the fix does too.
Companies that get their bill down 30-40% do not use AI less. They make three moves.
Move 1: Put the right model on each job
This is the biggest lever. AI models come in sizes, and the price gap between sizes is huge - often 10-30x per token. Most teams run everything on a premium model because it was the default when they signed up.
But most everyday work - replies, summaries, drafts, formatting - comes out just as well on a lighter model. Industry benchmarks show routing routine work to right-sized models saves 20-30% on its own, with no visible quality drop.
Move 2: Stop paying for the same answer twice
Teams repeat themselves constantly: the same product questions, the same email formats, the same report structure, week after week. Every repeat is billed like it is brand new.
The fix is reuse. Shared prompts, saved answers, and caching mean a repeated question costs a fraction of the original. On true repeats, caching saves up to 90%. Across a whole company, this is usually worth another 10-15% off the bill.
Reuse does not need engineering to start. A shared folder with the team's ten best prompts removes half the retyping on its own. The bigger half - caching identical requests behind the scenes - is what tooling handles, and it is why the repeat work your team does every Monday should not cost what it cost the first Monday.
Move 3: Retire the seats nobody uses
Seat licenses renew silently. The trial seats from last quarter, the person who left, the tool a team stopped using - all still billing. It is the least glamorous fix and the fastest one: find seats with no activity in 30+ days and cancel or reassign them. For most companies this is another 5-15%.
Seats hide in plain sight because each one looks small on its own. But seats only ever accumulate - they get added in thirty seconds on a busy day and removed never - and a handful of them across three tools is real money every single month.
The same math on a $3,200 bill
Percentages are easy to nod along to and hard to feel. So run the three moves on the bill at the top of this page - a small company paying $3,200 a month across its AI platforms.
- Right-sizing models (20-30%): worth $960-1,600 a month on this bill.
- Reuse and caching (10-15%): another $320-480 a month.
- Idle seats (5-15%): another $160-480 a month.
Add the ranges up naively and you get $1,440-2,560. In practice the moves overlap - an answer that is cached is not also billed on a cheaper model - so the honest total is the 30-40% from the callout above: $1,280-1,760 a month off. Take the very bottom of that range and the bill drops from $3,200 to $1,920, which is the January-to-March drop in the banner, give or take $20.
Now stretch it over a year. The conservative case hands back $15,360. The strong case clears $21,000. Nobody worked less, nobody lost a tool, nobody learned a new workflow. The spend changed and the work did not. And if your bill is $10,000 a month instead of $3,200, multiply everything by three - the percentages do not care about company size.
The order matters: measure first
Every successful cost cut starts the same way - with visibility. You cannot right-size models if you do not know which models are being used, and you cannot retire idle seats you cannot see.
- Week 1: connect your AI accounts, see actual usage per tool, per team, per person.
- Week 2: apply the obvious fixes - idle seats and model routing on one pilot team.
- Week 3-4: roll out to everyone, keep the monitoring on so the waste does not grow back.
Measuring first also settles arguments before they start. Ask three managers where the AI budget goes and you will get three confident, incompatible answers. The usage data ends the debate in one meeting: this team runs everything on premium, these nine seats are dead, this workflow repeats every week. Cuts that follow data do not feel like cuts. They feel like housekeeping.
How to check this in your own company this week
The week-by-week plan assumes you connect tooling. If you want a first answer before connecting anything, one person with admin access can rough it out in an afternoon:
- Pull three months of invoices from every AI vendor you pay and put the totals in one column. Most companies have never seen the combined number, and it is usually bigger than anyone guessed.
- Open each tool's settings and note the default model. If everything runs on the premium tier, assume the 20-30% right-sizing saving applies to you. Our guide to matching the right model to every task shows which work passes on a lighter model.
- Export the member list from each admin page, with last-active dates, and count the seats idle for 30+ days. The idle seats post walks through the routine, and the Copilot license utilization guide is a worked version for one common tool.
- Ask your three heaviest users what they asked the AI to do this week. Mark each task routine or complex. The routine share is the part a lighter model can carry.
- Apply the percentages from this post to your own totals. That one page of arithmetic is your business case.
"Isn't this IT's job?"
The most common objection, and in practice the reason the waste survives. IT keeps the tools running and the data safe. Finance sees one line per vendor on the statement. Neither owns the question "did that email need the premium model?" - so nobody asks it, and the bill quietly compounds. The 79% of companies that overshot their AI budget did not lack an IT team. They lacked an owner for the spend.
The fix is not a new hire or a committee. It is giving whoever owns the budget the same visibility IT has over uptime. That can be the manual checklist above, run monthly. Or it can be continuous: Optimize connects in about 15 minutes, is read-only, shows the first waste report within 48 hours, and most companies land their first cut inside two weeks. Either way, our AI cost management guide covers how to set up ownership so the numbers keep getting looked at.
What not to do
Do not ban tools or ration usage. AI genuinely saves your team hours; cutting usage to cut cost throws out the value with the waste. The goal is the same work at a smaller price, and that is exactly what the three moves deliver.
And do not wait for the annual budget review. AI spend moves monthly - new seats appear, defaults reset, a team adopts another tool - so waste grows back the same way it arrived. The companies that hold their 40% keep the monitoring on after the first cut. That is why the plan above ends with "keep the monitoring on", not with a victory lap.
