Cloudflare天价帐单事件结案:
非常感谢在 @ashleypeacock 的帮助下,Cloudflare已经决定全额退还在Bug故障期间所产生的费用(等待到帐中),最后附上Cloudflare最后回复的工单全文,非常详细的解释了Bug的过程的原因。
这里回答一些评论问的比较多的问题:
1、你没有设置消费提醒吗?Cloudflare消费没有硬上限吗?
是的,我没有设置消费提醒,默认只有一个系统自带的10USD的提示,而我每个月的消费大概在100USD-300USD左右,甚至在此事之前我并不清楚设置消费提醒这个问题非常重要。
如果你在使用Cloudflare的服务,强烈建议您也设置一个到多个提醒配置:https://t.co/98KEoySvFh<你的帐户ID>/billing/billable-usage
(Ps: 在AI的提示下它要求我配置100、300、500、1000、3000五个档位+多个邮箱的提醒)
Cloudflare目前是后付费模式,也没有消费上限配置选项,所以一旦产生消费,是需要您支付的,但据官方透露未来不久会上线这个配置功能。
2、如果不付帐单会怎么样?直接跑路行不行?
理论上,当然可以,但我们小时候看古惑仔里里面有一句台词:出来混,有错就要认,挨打要立正。承认错误、承担责任,我认为是一个成年人的表现,当然,实在付不起就算了。
(好了装逼结束)除了上述原因,主要还是因为我的主力项目还在Cloudflare上解析和做负载均衡,另外在最近几个月时间内,大概快速Viber了几十个业余项目,所以哪怕迁移也没有这么快能完成,总的来说,还是要感谢Cloudflare这种基建设施,能够超级快速Ship各种新项目,能快速验证各种想法,如果AI是挥动魔法棒的那种能力,那Cloudfalre就是接住这些魔法产物的地方。
所以最终我还是付了,我的信用卡并没有这么多钱,是跟银行申请临时提额来实现的,现在知道可以退绝大部份的话,就和小时候在衣柜旧衣服里面翻出钱来的感觉差不多,完全是一种惊喜的感觉。
3、到底写了些什么项目?怎么能闯这么大祸?用Ai写代码不检查,你就是活该吧?
首先,我是看到Cloudflare更新了Sandbox的版本,于是我想测试一下,如果将某些开发任务分解为多个子任务,然后利用Sandbox的并发处理能力,能不能大幅度的提升开发效率?基于这个假设,我让AI迭代了几个不同的版本(包括一个主Agent来分解任务、多Agent自行沟通协调任务队列等不同的逻辑),但最终的结论是并没有明显的效率提升(可能是因为我能力问题)。
于是这个项目就在最后一次更新之后就搁置了,但它仍然部署在Cloudflare上,尽管无任何人使用,但在8.29日,也就是一个月前使用Codex的一次提交中,写了一个DO无限死循环的Bug,而这个Bug会在23天之后被唤醒,开始无限循环,产生消费。
我是一个大龄非科班程序员(自以为),没有在大厂工作过,是30岁之后完全自学的编程,学过Python基础和能看懂一点点TS,以前就是那种:我有一个想法,只差一个程序员的人,于是后来就自己学编程写自己想要的东西,写了几年,做了一些项目,赚了一点钱。
在彻底拥抱AI之后,就放飞自我了。所有代码都是不看不检查的,只描述需求和最终验收,并不是不想看,因为你真让我看我也看不懂,我只关心它能做到什么程度,能不能实现我要的效果,捅出篓子再说。
之前一直是使用Claude Code + Fable系列模型开发的,也从来没有出现过严重的问题,直到上个月底手上没有可用的帐号了,才不得不使用了Codex+GPT Sol 5.6模型开发,真就用了几天就出事了,这是一种惯性依赖,在配对什么智力程度的模型,就需要做对应程度的检查。
所以最终我买单也我也认的。
4、你说这是第二贵的教训,那人生第一贵的教训是什么?
评论中有人回复:结婚;
哈哈,当然不是,这是另一个故事,以后有机会再分享详情,主要是使用了非自己Kyc的币安帐户3年,收了一些钱在里面,突然被币安强制人脸认证,只能找原KYC的持有人协商找回,那一次损失的原本大概是3万刀,最终追回不到一半。
工单原文(技术细节可以看这里):
Hi **,
I want to start with an apology, and it's not a form-letter one. The experience you had — a five-figure bill, a support queue that didn't move fast enough on it, having to pay in full while waiting — that's a failure on our end, not just a billing edge case. I'm sorry it went that way.
Your refund has been processed in full: US$10,701.41, covering all the Durable Objects charges (rows read, rows written, requests and duration) from the period between September 23 and October 6. You should see it reflected on your account within a few billing days — reply here if you don't and I'll chase it down directly.
What actually happened, since the support loop clearly didn't explain it well
Your PiSessionDO alarm hit a nasty edge case in how Durable Objects schedules alarms. Once a checkpoint entered its refresh window, alarm() was rescheduling itself to Date.now() — which is effectively "fire again immediately." That's not a theoretical gotcha; it's a known sharp edge where the alarm fires at roughly 600 invocations per second with nothing meaningful happening except storage churn.
The cost driver wasn't CPU time or request duration — a DO that stays awake continuously costs a few dollars a month in duration. The real cost was SQLite storage operations. At hundreds of invocations per second, each touching storage, rows read and rows written compound fast. That's where the bill came from.
The fix you already shipped — a 15-minute minimum delay, a retry budget that stops rescheduling after N exhausted attempts, and a regression test — is exactly right. If those had been in place from the start, this incident would have cost dollars, not thousands.
A few things worth keeping in the codebase going forward
Since you're clearly building quickly, a few patterns that specifically protect against this class of issue:
On alarms: Never pass Date.now() or a past timestamp to alarm(). If the scheduled time has already passed, it fires immediately. The safe minimum is something like Date.now() + 60_000 (1 minute), but for a refresh/retry cycle, 15 minutes is a reasonable floor. Also: always store your retry counter in DO storage and stop rescheduling once the budget is exhausted. Without a hard stop, the loop survives restarts.
On spend controls: Under Notifications in your Cloudflare dashboard, you can configure billing usage alerts that trigger at specific spend thresholds. Setting these at 25%, 50%, and 75% of your expected monthly spend means something like this surfaces within hours, not two weeks. Worth setting up before you're back to building — it takes two minutes and catches a wide range of anomalies beyond just alarm loops.
On DO analytics: The Workers & Pages dashboard shows per-namespace invocation counts and storage metrics. A namespace that jumps from a few thousand invocations per day to hundreds of millions looks very different — it's the fastest early signal of a loop that's also worth checking periodically when you're iterating quickly on a DO-heavy service.
One general note on AI-assisted development and usage-based products: moving fast with AI tools is genuinely useful, but usage-based services like Durable Objects, D1, and R2 can have non-obvious cost profiles when code has bugs. The same speed that makes vibe coding productive can generate a billing surprise before you've had a chance to review what shipped. A brief manual review of anything that touches persistent storage or sets alarms before it hits production tends to catch the expensive class of mistakes early.
I hope you'll stay on the platform. What happened here was a combination of a real edge case in the product and a real failure in how we handled your case afterward. Both matter and we're not dismissing either.
If you have any follow-up questions about the credit, want to walk through your DO architecture before you re-enable the Workers, or want help auditing the rest of the codebase for similar patterns — I'm here.
Let us know if you have any other questions,
Daniel
Cloudflare Support
This reply was enhanced by Cloudflare Workers AI (Kimi K2.7 by Moonshot AI).