It’s incredible that in 2026, AWS and GCP are only just now introducing this. It’s possibly one of the most obviously needed features for a cloud provider.
Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.
It's one part technical, one part a product decision. The technical part is that billing is not actually instant. As a most basic example, a VM reports its billing units every X period of time it is active. If there is some network blip but it's still running, then that billing data could be delayed.
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
I think there is a middle ground between deleting data and allowing 5000 VMs to be created to mine bitcoin. Obviously there are a lot of different scenarios to consider but the explosive costs seem to be constrained mostly to a couple of features which would be fairly safe to cap.
A very charitable take, in light of tech industry habits of exorbitant rent-seeking in scenarios of Platform Dominance (e.g. Google and Apple on the app store). We should remember AWS and Google companies are among the best in the world at A/B testing and extracting revenue from cloud services.
When you're one of only two real options out there, you can afford to demand users put up with things that on their surface seem ridiculous. Such as a billing system that (oops!) makes it difficult for customers to see where their costs are coming from, trim their largest sources of spend, notice meaningful changes in line item prices, or limit their spend. Wild how they can figure out a million different advanced services but gosh-darn-it can't figure out the hardtech of displaying line items.
Large enterprises can afford employees who are tasked full time with unwinding this capacity to mitigate the impact of these billing headaches. But I think this measure is introduced now because LLMs introduced a risk that these billing specialists could not control without caps.
If these cloud providers had needed a standard customer acquisition strategy to grow to their current size, hard caps and other “training wheels” features would already be in place to get people interested in and comfortable using the platform, with the hope of eventually getting a foothold into Enterprise like most SaaS startups have to do (“enjoy our product on a side project and then recommend us to your CTO!”). But AWS and GCP got to start as in-house providers for their own constellations of massive sites and back out from that to serving other hyper scale businesses first. The lack of friendly on-ramps and starter account features is a reflection of that origin more than anything.
I had a $.20/month recurring charge from AWS that I could only remove¹ by completely deleting my AWS account. That was enough to get me to give up on AWS for personal projects.
⸻
1. The key word was “I.” Maybe someone more skilled at navigating AWS’s menu structure than I could would have done it quickly and easily, but even though I knew what it was for, turning it off and not getting billed for it turned out to be a huge challenge. Thankfully it was only $.20, but if I were using the service for something that generated actual bills, that $.20 (and possibly more) would end up quietly siphoning money out of my pocket into Amazon’s).
Sounds familiar. I’m being billed £0.01/month for something in GCP, I don’t know what even after digging, but I’m too fearful to complain about it or disable the account lest it somehow gets my main Gmail account blacklisted somehow.
It's definitely technically difficult. You can't easily estimate how much an operation is going to cost before you kick off that operation, which means as soon as you get close to the limit you are at risk of tripping it.
Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.
Stop everything is pretty damaging any real business though. Things were better in the era of VPSs. You paid for a fixed amount of compute, if you ran a stupidly expensive operation than it just maxed out your system for a certain amount of time and things slowed down. But you didn’t kill the service entirely and you didn’t have unlimited potential price
Yes and then the choice is run it and forgive it, or, stop the process midway.
If you stop then you have to decide whether to charge for uncompleted work.
Interesting tradeoffs.
For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.
Because most enterprise users would much rather have overages in billing than outages. The opportunity costs on any serious service I deploy dwarfs usage pricing, at least at the level a generic cloud can determine.
So it’s a feature the best customers don’t want, that adds risk to those customers deployments, to appease the worst customers.
At least historically. Perhaps Simon is right that the calculation has changed.
It’s not a binary decision though. Any sensible enterprise has many AWS accounts. Often hundreds or thousands. It’s the only clear separation of privilege.
There’s no reason for the majority of them to have an infinite budget cap. Prod? Sure. UAT? Why not. The sandbox environment Johnny just spun up to test some new agentic workflow? Hell no.
It is a technical reason. Basically cloud billing is much more granular and across many more services / line items than most things that basically the pipelines that figure out how much you have spent take a long time to know how much you have consumed. I believe all cloud providers with granular usage based billing have this problem.
This is one of those features that customers think they want without having thought it through:
“Never let me spend more than $X” also means, “Shut down my business-critical app/service/solution at 2 am on a Sunday morning because Joel in IT forgot to plan for the new report runs.”
The product design work to let customers have the first thing without risk of major pain from the second thing is non-trivial.
Counter argument, this is the sort of thing that, especially for a smaller business or individual, can be the difference between a bad night and bankruptcy.
Sure it sucks that critical services blinked out at 2am. But what sucks even more is finding out the image on my ASG had a vulnerability that allowed someone to install a bunch of bitcoin miners which kept me fully scaled from midnight to 2am.
Or more likely, that a mistake in terraform 1000xed my spending.
Most people have predictable spending and could easily say "don't spend more than 10x what I normally spend". Or 1.5x, or 2x, 3x, etc. All depending on how they want to balance a runaway cloud expense.
The alternate conversation is "the new report run had a bug and cost us $1,000,000 over the weekend" and I think that one's usually worse.
But one could just have two categories of service - the default capped plan, and a special Enterprise one where you sign a contract making it clear you understand the consequences of not having a budget limit.
Also, if your average usage is $900, set your hard limit at $2,000, not $1,000. Then when the report runs $500 over expected, you get a soft limit email and still have your report. Even a "business critical" run is probably not actually worth more than double your average spend.
Big companies have thousands of budgets. An email is _worthless_. In fact, it would probably cause me to lose faith in a cloud that provided that as the control.
Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
These shouldn't even exist without a negotiated contract.
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
Monthly electricity bills are based on usage, and it works well but there’s a limit to how surprising a bill can be. The difference is the relative orders of magnitude you can be charged for these services you can go from 20$/month to 200k/month without warning.
I know AWS and similar sites have no such concepts, but other than those, this is an existing option of most kinds of service providers.
Realistically, most providers can't even allow you to consume $10k usage if they don't have the certainty that you can pay up. Prepaid is what gives them that certainty.
Hard caps are rare because companies find it more profitable to forgive sympathetic individuals' bills while raking in profits from corporations whose services have gone awry
Having a monthly summary or estimate of how your spending is going would be really useful, too.
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
My org has a leaderboard for AI spending each month, and I have found it interesting how fast the distribution decays, just within the top 10 users. I often think “what did these people do with all those tokens?” It’s interesting to think the answer to that question is “maybe not a lot?”
Right? If spending the most is lauded, why wouldn't I use the most expensive model, automate things that don't need automating, build things I don't need to build etc just to jack the spend up?
> In an ideal world, our agents could help with this.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on its OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
Counterpoint: if you can automate API calls on the client side, why can't you automate billing caps? If you want a machine that can run 24-7 and make money for you while you sleep (which let's face it is the motivation for a lot of AI takeup), isn't the onus on you to install cicuit-breakers?
Because many cloud services have incredibly complex or opaque pricing structures that make it difficult to impossible to determine how much something is going to cost you ahead of time, especially if it's usage-based a la network egress (and the usage statistics don't update frequently enough to make such circuit breakers possible to implement client-side).
They might not be able to predict your bill but how much time do they need to add up what you already spent to minimize your overage? And TBH how much time should be acceptable to exceed your cap before it's their fault for the lag in their software.
I would just not sign up for a service without price transparency, or pre-calculate my liability based on available information before pushing the (metaphorical) Deliver Now button.
Making incredibly complex and opaque pricing structures is not necessary for the providers to charge for and make a profit on their service. And being technically difficult is a lazy excuse. Cloud platforms have to solve many, much more difficult challenges to offer their services at all, they just don’t want to invest the time in more customer friendly billing because they expect it will result in reduced revenues.
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
I'm not sure where you got the impression that I'm making excuses for cloud providers. I'm just stating the way things are, not the way I think they should be.
I had an api key set to read only that somehow ran up a $400 bill, I contacted openai about it and never heard back. Not quite the same thing, but still, I find this very annoying.
We always did, the clouds convinced us that overages were the norm. You can blame credit ratings as another vector for big business to screw everyone over. Everything should have been pay in advance with an alternate billing method for overages if you want it.
The solution is to not give agents access to MCP servers.
The entire MCP ecosystem is ludicrous. You’re paying for inference for an agent to make the same decisions over and over again, and yet the actions they’re taking can be so easily written by those same agents into a bash script you can run again and again, deterministically and for free.
MCP is the problem. Having agents “use a product on your behalf” is the problem.
The pattern you’re looking for is that agents should write scripts that use products for us, and that doesn’t need a new protocol.
It'd be nice if more than AI spend worked this way, autoscaling is almost a mixed blessing because unpredictable pricing can be worse than the cost savings...
Ubicloud does not have hard budget caps, which I only realized this morning after moving all my CI over to them over the past few months. Fortunately I didn't learn the hard way.
I understand this is snark, but if you think about it, this is already implemented in electrical infrastructure. If I use too much power, the circuit breaker trips to protect me and protect the electrical grid. OP is about a billing breaker, but the parallels should be obvious.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
The peak throughput does result in an overall monthly limit though. For a house with a 200A main breaker, that effectively limits your electric bill to $7,000/month, which is very reasonable compared to the tens of thousands of dollars in a single day that a lot of cloud billing disasters end up costing.
Why do people think new laws are needed to solve every last problem in the world?
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
Yes and no. I suspect many of the hard limits were set arbitrarily, and we'll see a relaxation of limits as people get frustrated with the limited use they get out of them. And some services will genuinely need to be re written to support higher rps or risk losing customers
I'd support this provided we have the converse as well: if the customer doesn't pay their bill on time, the service gets shut down immediately. (Disclosure: I sell SaaS services to people who don't pay their bills on time).
One of the biggest benefits of not engaging with LLMs or any of this nonsense is you dont have to care about all these "self made" problems of the LLM-gliteratti.
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
Because it’s unlikely they’ll actually be able to collect that million dollars from a lot of those customers. Rephrased: why would your vendor want to make it harder to accidentally give you a million dollars of services in exchange for debt of dubious quality?
the premise seems a bit faulty to me. why should we be giving next token predictors access to spend our money? like what great benefit do we get from this that we should allow them unfettered access, but with safeguards in the form of hard budget caps?
I don't think Simon means you should hand off the spending to agents/LLM (which would also make me uneasy) but that if you're probing one for hosting/SaaS providers they should default to recommending ones with budget caps
This is about budget caps (“I don’t want to spend more than $100, cut me off once I spend that much”) not price caps (“no one is allowed to charge more than this price per token”).
Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
When you're one of only two real options out there, you can afford to demand users put up with things that on their surface seem ridiculous. Such as a billing system that (oops!) makes it difficult for customers to see where their costs are coming from, trim their largest sources of spend, notice meaningful changes in line item prices, or limit their spend. Wild how they can figure out a million different advanced services but gosh-darn-it can't figure out the hardtech of displaying line items.
Large enterprises can afford employees who are tasked full time with unwinding this capacity to mitigate the impact of these billing headaches. But I think this measure is introduced now because LLMs introduced a risk that these billing specialists could not control without caps.
⸻
1. The key word was “I.” Maybe someone more skilled at navigating AWS’s menu structure than I could would have done it quickly and easily, but even though I knew what it was for, turning it off and not getting billed for it turned out to be a huge challenge. Thankfully it was only $.20, but if I were using the service for something that generated actual bills, that $.20 (and possibly more) would end up quietly siphoning money out of my pocket into Amazon’s).
Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.
Green means go Orange means finish what you're doing but don't start anything new Red means stop everything
And probably a special rule to permit stable, critical spend through regardless, the same way we allow police and ambulance to run lights.
If you stop then you have to decide whether to charge for uncompleted work.
Interesting tradeoffs.
For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.
So it’s a feature the best customers don’t want, that adds risk to those customers deployments, to appease the worst customers.
At least historically. Perhaps Simon is right that the calculation has changed.
There’s no reason for the majority of them to have an infinite budget cap. Prod? Sure. UAT? Why not. The sandbox environment Johnny just spun up to test some new agentic workflow? Hell no.
“Never let me spend more than $X” also means, “Shut down my business-critical app/service/solution at 2 am on a Sunday morning because Joel in IT forgot to plan for the new report runs.”
The product design work to let customers have the first thing without risk of major pain from the second thing is non-trivial.
Sure it sucks that critical services blinked out at 2am. But what sucks even more is finding out the image on my ASG had a vulnerability that allowed someone to install a bunch of bitcoin miners which kept me fully scaled from midnight to 2am.
Or more likely, that a mistake in terraform 1000xed my spending.
Most people have predictable spending and could easily say "don't spend more than 10x what I normally spend". Or 1.5x, or 2x, 3x, etc. All depending on how they want to balance a runaway cloud expense.
But one could just have two categories of service - the default capped plan, and a special Enterprise one where you sign a contract making it clear you understand the consequences of not having a budget limit.
Also, if your average usage is $900, set your hard limit at $2,000, not $1,000. Then when the report runs $500 over expected, you get a soft limit email and still have your report. Even a "business critical" run is probably not actually worth more than double your average spend.
The price cap should be the number you’d be willing to spend to avoid an outage vs when you’d rather kill everything and work out what happened.
Sending an email when your budget gets low shouldn't be a big lift.
https://cloud.google.com/blog/topics/cost-management/new-ear...
Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
So my powerbills are predictable.
Whereas traffic spikes to websites are not.
This age of abusive AI crawlers and the non-revenue generating traffic has been a very real problem for me!
I think they explicitly said that.
I know AWS and similar sites have no such concepts, but other than those, this is an existing option of most kinds of service providers.
Realistically, most providers can't even allow you to consume $10k usage if they don't have the certainty that you can pay up. Prepaid is what gives them that certainty.
Plus, nobody wants to be the fired PM who said "I spent our eng. hours to achieve -20% revenue".
It works for some non-B2Bs.
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
Phone service, bank card, home internet etc.
If you don't pay your bill than they just cancel your membership and it works ok.
People in western countries are just getting shafted by companies for (mostly) no reason because an alternative balance is just inconceivable.
If you don’t have a mechanism for enforcing hard caps, you don’t get to send customers a bill for unlimited amounts.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on its OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
[1] https://github.com/tkgally/je-dict-1
[2] https://github.com/tkgally/eex-dict
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
Or if there was info that was after potentially horrible expensive operation.
That was blocking automated checking.
Seriously. You all asked for this.
It seems like you’re saying “hey, you asked for a product, so you deserve for it to have a user-hostile feature”
Like hey, you asked for trains? Well, then you have no right to complain about any aspect of a train.
I can’t think of anyone saying they would hate for AWS to support hard spending caps.
The entire MCP ecosystem is ludicrous. You’re paying for inference for an agent to make the same decisions over and over again, and yet the actions they’re taking can be so easily written by those same agents into a bash script you can run again and again, deterministically and for free.
MCP is the problem. Having agents “use a product on your behalf” is the problem.
The pattern you’re looking for is that agents should write scripts that use products for us, and that doesn’t need a new protocol.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
You can rack up an outrageous monthly electricity bill without tripping a breaker.
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
Cheers!
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
What on earth is Simon whittering on about?
This may happen also without vibecoding.
Article seemed clear on that?