How to budget tokens
About a week back, I was talking with a customer, and they started sharing about their AI initiatives. During the conversation, they mentioned that their token budget for the year was already gone. It lasted only four months. They did estimates and pilot runs, and still could not get it right.
And to my surprise, I found they are not alone. In April, Uber's CTO said the company had already used up its planned 2026 AI budget, only months into the year, after a surge in the use of AI coding tools.
Different industries, use cases, and even usage patterns. But the same problem: token consumption was way more than anyone had budgeted for.
This customer's number stayed with me because, like many other practitioners, I use monthly subscriptions, so I have never run into a financial constraint. The most I have had to do is wait 4-5 hours for the usage window to reset.
I got curious and started looking into how the industry is going about token consumption and estimates. I broke it down into four smaller questions:
- How do you estimate what a workload will spend before it runs?
- How do you keep it inside a limit once it does?
- How do you find out early that it is drifting?
- When the bill jumps, how do you find what caused it?
Here is what I found. Estimates and pilots miss because pilots use handpicked, skilled users, agents make far more model calls per task than chat ever did, and a few expensive runs pull the total up. So pilot with a mix of users, budget for the tail, and check against real spend every month. Vendors now offer hard spending caps, but none of them act instantly, and they stop a whole project at once. The caps that help sit closer to the spender: per team, per key, and per agent run. Alerts help only if they watch the monthly burn rate, not the yearly total. And finding the culprit works only if every call carries a team, app, or user label before the bill jumps.
How do you estimate what a workload will spend?
This was the question that bothered me most, because the customer had done the work. They estimated. They ran pilots. It still did not hold.
When I looked at how token usage grows, three reasons stood out. None of them is specific to that customer, but any one of them can break a careful estimate.
1. The pilot users are not your real users
A pilot group is usually handpicked. These are often the people who asked for the tool, already know how to use it well, and write clear prompts. A clear prompt gets the job done in a few rounds. A vague one sends the model back and forth, and every round costs tokens. After rollout, the tool reaches everyone, and each person tends to use it more as they start to trust it. An estimate built on pilot numbers is built on the most efficient version of the workload.
The remedy is to pilot with a mix of users. Include strong users, average users, and people who are new to these tools, and track their cost per task separately. The gap between the most and least efficient users tells you more about the real budget than the average does. It also shows where some training or a few shared prompt templates could bring the cost down before the full rollout.
2. Agents changed what one unit of work costs
When I use a chat assistant, one question is one model call. A coding agent working on one task reads files, runs commands, and calls the model again after every step. Every call sends the whole conversation so far, and that conversation grows as the task goes on.
To make that concrete, I put together a rough comparison. These numbers are my assumptions, not measurements or real pricing, and I left out output tokens and caching to keep the math simple.
| Per person, per working day | Chat assistant | Coding agent |
|---|---|---|
| Units of work | 20 questions | 10 tasks |
| Model calls | 20 | 300 (30 per task) |
| Tokens sent per call | 3,000 | 50,000 |
| Tokens per day | 60,000 | 15,000,000 |
| Cost at $5 per million input tokens | $0.30 | $75 |
Same person, same day, 250 times the tokens. If the budget was built on the left column and people move to the right one, it is gone early. Prompt caching can bring the real cost down a lot, which is one more reason this is hard to get right on paper.
The remedy is to pilot the same tool and workflow people will use in production, agent included, not a chat version of it. And count model calls per task, not only tokens per call, because the number of calls is what changes the most.
3. The expensive runs hide in the tail
Most tasks are cheap. A few go long, retry, or loop, and they pull the total up far more than an average suggests. A pilot report that shows only the average cost per task hides them.
The remedy is to report the 90th percentile and the single most expensive task next to the average, and to budget for that tail, not for the typical day.
Where that leaves the estimate
What the vendors offer here is limited. Both Anthropic and OpenAI have endpoints that count the tokens in a request before you send it. That helps size one call. It does not tell you how many calls a task will take, or how many people will be running tasks six months from now.
So my take is to treat the estimate as a range, not a number. Use the mixed pilot to get a typical cost per task and a worst case, multiply by the number of people you expect at the end of the year, not the start, and then compare it with real spend every month. The estimate is only a starting point. The monthly check is what keeps it honest.
How do you keep it inside a limit?
This is where my subscription habit made me look at things differently. On a subscription, the limit is built in. Claude's Max plan, for example, resets a session limit every five hours and adds a weekly limit on top. I used to see that window as an annoyance. Looking at it now, it is a hard cap that the vendor set for me. When I hit it, I stop, and the bill does not move.
Pay-per-token API usage does not work like that. By default, it keeps going and keeps billing.
While I was reading about how to put a cap on that, I came across Simon Willison's post, "We're going to need default hard budget caps on pretty much everything". He separates two kinds of limits. A hard cap says "after $X/month, cut this thing off and return errors." A soft cap only sends a warning email, and soft caps, he writes, "will not cut it."
That sent me to check what the vendors offer today.
| Vendor | Hard cap you can set | What happens at the limit | Is it instant? |
|---|---|---|---|
| OpenAI | Organization and project spend limits | Requests fail with a 429 error | No. "Enforcement is not instantaneous," so spend can slightly exceed the limit |
| Anthropic | Organization and workspace spend limits | Requests fail with an error saying you reached your usage limits | Not stated |
| Google Cloud | Spend caps, in preview, for the Gemini API, Agent Platform, and Cloud Run | New requests to that service in that project pause | No. Enforcement "isn't instantaneous," and you pay the overage |
| AWS | Per-project spend limits, rolling out gradually to new customers | The project pauses until you raise the limit | Not stated |
Two things surprised me. None of these caps promises to act instantly, because a provider can only stop spending it has already counted. And they are blunt. A cap on a project or workspace stops everything in it, including features your customers use. That is better than an open-ended bill, but if it is the only cap, a budget problem simply turns into an outage.
So I looked at what large companies are doing, and Microsoft gave a good example. This week, The Information reported that Microsoft had expected to spend at least $1 billion a year on internal use of Anthropic's models and has since cut that projection by more than a third. As part of that, individual monthly usage caps reportedly dropped from $100,000 to roughly $10,000 in most cases. That is a cap at the level of a person, not the whole company.
Putting the cap close to whoever spends is the pattern I would copy. I think about it at three levels.
Per team or person. Each team gets a monthly budget, roughly a twelfth of its share of the yearly budget. If your model calls go through a gateway, the one service that every call passes through, this is usually just a setting. LiteLLM's proxy, for example, supports a max_budget and a budget_duration on teams, users, and keys, and rejects requests once a budget runs out.
Per app or job. Each app or scheduled job gets its own key with its own budget, so one runaway job cannot spend everyone else's share.
Per agent run. This one lives in your own code, and it is the only cap that acts right away. Every model response already says how many tokens it used, so the agent can keep a running count and stop itself.
MAX_STEPS = 20
MAX_TOKENS = 200_000
steps = 0
tokens_used = 0
while True:
if steps == MAX_STEPS:
raise RuntimeError(f"Run stopped: step limit of {MAX_STEPS} reached")
response = call_model(messages)
steps += 1
tokens_used += response.usage.input_tokens + response.usage.output_tokens
if tokens_used > MAX_TOKENS:
raise RuntimeError(f"Run stopped at step {steps}: token limit reached")
if is_finished(response):
break
messages = run_tools(messages, response)
In plain words, the agent keeps two counters: how many steps it has taken and how many tokens it has used. Before each step, it checks the step count. After each step, it adds the tokens reported in the response. If either counter goes over its limit, the run stops with a clear message instead of carrying on. The step limit catches an agent going around in small circles, and the token limit catches one sending huge requests. The functions call_model, is_finished, and run_tools stand in for your own code, and the usage field names match the Anthropic Messages API and the OpenAI Responses API.
The provider cap still makes sense on top of all three, set above normal monthly spend, as a backstop for the day your own limits have a bug.
One trap I did not expect. When a spending cap trips, the error can look like an ordinary rate limit. OpenAI's spend limits return HTTP 429, the same code as "too many requests," and Anthropic uses 429 when an organization reaches the monthly cap for its usage tier. A gateway that switches to a backup provider on every 429 will send the same runaway work to a second provider and a second bill. I would treat spend-limit errors, such as project_spend_limit_exceeded and enforced_spend_limit_reached, as a stop. No retry. No failover.
How do you find out early that it is drifting?
Alerts are the soft cap Willison talks about. They do not stop anything. But they are still how a person finds out, so the real question is when they fire.
The math from the customer's story helped me here. If a yearly budget lasts four months and spending is even, a quarter of it is gone after the first month. An alert at 80 percent of the yearly budget would fire early in month four, when there is almost nothing left to save. A monthly view would have shown the problem after month one.
So I would set alerts against a monthly line, the yearly budget divided by twelve, not against the yearly total. A few vendor features help with that:
- Google Cloud and AWS both let a budget alert fire on forecasted spend, not just actual spend. A forecast notices a fast burn rate within weeks instead of months.
- Google added early anomalies in July. It flags daily AI spend that looks unusual for a project and names the top three SKUs behind the jump.
- Anthropic's Usage and Cost API typically shows usage within about five minutes, and Anthropic suggests it for building your own budget alerts.
The catch is delay. AWS Budgets updates up to three times a day and warns that you can pass a threshold before the notification reaches you. Google also notes a delay between usage and billing. For a monthly budget, a few hours does not matter much. For an agent stuck in a loop overnight, it does, and that is why the per-run cap in your own code matters.
I would also send the alert to the team that owns the spend, not only to finance. The people who can change a prompt, switch a model, or turn off a job are the ones who need to know.
When the bill jumps, how do you find what caused it?
A single invoice line does not say which team, app, or person spent the money. And labels cannot be added afterward. They have to be on the calls before the spike.
The simple version is separation. Give each team its own project or workspace, give each app or job its own key, and avoid shared keys. Then the vendor reports do most of the work:
- OpenAI's Usage API groups usage by project, user, API key, and model.
- Anthropic's Usage and Cost API groups usage by workspace, API key, and model, and its Claude Code Analytics API reports estimated cost per user per day.
- On AWS, Bedrock can tag costs through application inference profiles, and since April it can record which IAM user or role made each call.
- Google Cloud budgets can be scoped by project, service, and label.
There is a catch with gateways. AWS points out that when calls go through a gateway or proxy, the recorded caller is the gateway, not the person behind it. The same happens with any shared key. If you run a gateway, it has to do the tracking itself by passing a team or user ID with each request. LiteLLM, for example, tracks spend per key, user, team, and tag.
Once the labels exist, finding the cause is mostly sorting. I would group last month's spend by team, then by app, then by model, and compare it with the month before. Then I would look at three things: who started using it, which new workload showed up, and whether a model or context change made each call more expensive.
What this means for practitioners and architects
If you are having this conversation with your customers or within your organization, I would suggest four things:
- Turn the yearly budget into a monthly one. Divide it by twelve and compare it with last month's real spend. That one number shows whether there is a problem. For new workloads, estimate from a pilot with a mix of users and budget for the expensive tail, not the average.
- Give every team and app its own key. Split any shared keys into a project or workspace per team and a key per app or job, so the next spike has a name on it.
- Put caps close to whoever spends. Set monthly caps per team at the gateway, a step and token cap in every agent, and a provider cap above total spend as the backstop. Then trip each one once on purpose to make sure it stops the work instead of retrying or failing over.
- Alert on the forecast, and send it to the owners. Add a forecast alert and an anomaly alert that go to the team that owns the spend, not only to finance.
None of this makes a single call cheaper. That is the other half of the problem, and the Token and Cost Management chapter covers it through context trimming, prompt caching, and choosing the right model size. The AI Gateways chapter explains why the gateway is a good place for team budgets and labels.
Looking back at my own setup, that five-hour wait I used to grumble about was doing a job I never thought about. On a subscription, the vendor caps you whether you like it or not. On the API, nobody does unless you set it up.
Have you had this conversation with a customer or inside your own team? How did you manage token budgets in your project? Leave a comment below and share what worked for you, and what did not.
Comments
Comments are provided by Giscus and GitHub. Loading them connects your browser to GitHub, and GitHub's privacy terms apply.