Blog  /  Playbook

How companies cut their AI bill 40% - without using AI any less

7 min read · May 2026 · Optimize team
January: $3,200  →  March: $1,900
Same team, same amount of AI work. The difference is three fixes.

Here is the uncomfortable number: 79% of companies overshot their AI budget in the past year, and research puts 30-40% of typical token spend down as pure waste. Even Uber wasn't immune - it burned through its entire 2026 AI budget in four months before it put any limits in place. The good news is that the waste follows the same pattern almost everywhere, which means the fix does too.

Companies that get their bill down 30-40% do not use AI less. They make three moves.

Move 1: Put the right model on each job

This is the biggest lever. AI models come in sizes, and the price gap between sizes is huge - often 10-30x per token. Most teams run everything on a premium model because it was the default when they signed up.

But most everyday work - replies, summaries, drafts, formatting - comes out just as well on a lighter model. Industry benchmarks show routing routine work to right-sized models saves 20-30% on its own, with no visible quality drop.

Right-sizing models: typically 20-30% savings, the single biggest fix.

Move 2: Stop paying for the same answer twice

Teams repeat themselves constantly: the same product questions, the same email formats, the same report structure, week after week. Every repeat is billed like it is brand new.

The fix is reuse. Shared prompts, saved answers, and caching mean a repeated question costs a fraction of the original. On true repeats, caching saves up to 90%. Across a whole company, this is usually worth another 10-15% off the bill.

Reuse does not need engineering to start. A shared folder with the team's ten best prompts removes half the retyping on its own. The bigger half - caching identical requests behind the scenes - is what tooling handles, and it is why the repeat work your team does every Monday should not cost what it cost the first Monday.

Move 3: Retire the seats nobody uses

Seat licenses renew silently. The trial seats from last quarter, the person who left, the tool a team stopped using - all still billing. It is the least glamorous fix and the fastest one: find seats with no activity in 30+ days and cancel or reassign them. For most companies this is another 5-15%.

Seats hide in plain sight because each one looks small on its own. But seats only ever accumulate - they get added in thirty seconds on a busy day and removed never - and a handful of them across three tools is real money every single month.

The three moves together: 20-30% + 10-15% + 5-15%, overlapping to a realistic 30-40% total reduction. For a normal company, 30-40% is the realistic target.

The same math on a $3,200 bill

Percentages are easy to nod along to and hard to feel. So run the three moves on the bill at the top of this page - a small company paying $3,200 a month across its AI platforms.

Add the ranges up naively and you get $1,440-2,560. In practice the moves overlap - an answer that is cached is not also billed on a cheaper model - so the honest total is the 30-40% from the callout above: $1,280-1,760 a month off. Take the very bottom of that range and the bill drops from $3,200 to $1,920, which is the January-to-March drop in the banner, give or take $20.

Now stretch it over a year. The conservative case hands back $15,360. The strong case clears $21,000. Nobody worked less, nobody lost a tool, nobody learned a new workflow. The spend changed and the work did not. And if your bill is $10,000 a month instead of $3,200, multiply everything by three - the percentages do not care about company size.

On a $3,200 monthly bill: $1,280-1,760 back per month, $15,360 or more per year.

The order matters: measure first

Every successful cost cut starts the same way - with visibility. You cannot right-size models if you do not know which models are being used, and you cannot retire idle seats you cannot see.

Measuring first also settles arguments before they start. Ask three managers where the AI budget goes and you will get three confident, incompatible answers. The usage data ends the debate in one meeting: this team runs everything on premium, these nine seats are dead, this workflow repeats every week. Cuts that follow data do not feel like cuts. They feel like housekeeping.

How to check this in your own company this week

The week-by-week plan assumes you connect tooling. If you want a first answer before connecting anything, one person with admin access can rough it out in an afternoon:

"Isn't this IT's job?"

The most common objection, and in practice the reason the waste survives. IT keeps the tools running and the data safe. Finance sees one line per vendor on the statement. Neither owns the question "did that email need the premium model?" - so nobody asks it, and the bill quietly compounds. The 79% of companies that overshot their AI budget did not lack an IT team. They lacked an owner for the spend.

The fix is not a new hire or a committee. It is giving whoever owns the budget the same visibility IT has over uptime. That can be the manual checklist above, run monthly. Or it can be continuous: Optimize connects in about 15 minutes, is read-only, shows the first waste report within 48 hours, and most companies land their first cut inside two weeks. Either way, our AI cost management guide covers how to set up ownership so the numbers keep getting looked at.

What not to do

Do not ban tools or ration usage. AI genuinely saves your team hours; cutting usage to cut cost throws out the value with the waste. The goal is the same work at a smaller price, and that is exactly what the three moves deliver.

And do not wait for the annual budget review. AI spend moves monthly - new seats appear, defaults reset, a team adopts another tool - so waste grows back the same way it arrived. The companies that hold their 40% keep the monitoring on after the first cut. That is why the plan above ends with "keep the monitoring on", not with a victory lap.

Sources: CloudAtler on AI cost overruns, Morph on the five cost levers, MLflow 2026 enterprise guide.
See what your company actually spends on AI
Book a free 15-minute demo and leave with your own waste estimate. No code, no commitment.
Book a Free Demo
Keep reading
Uber burned its entire 2026 AI budget in four months. Here's the lesson. Big model or small model? Matching the right AI to every task