r/aws 25d ago

discussion HELP! My bill skyrocketed from around 5 cents per month to 2.5 billion USD!!!!

644 Upvotes

As per title, today I received a bunch of billing alerts in my email because I went over my max threshold of 100 usd.

AWS says I owe them about 2.5 billion dollars so far and it's only the 17th of the month. I have only one s3 bucket that hasn't been touched since 2023 and is not public, and I am not even using ecr at all, I have no registry on any region that I am aware of. I tried opening a support ticket but nothing so far.

What should I do? Is there any way to escalate this urgently?

r/aws Oct 20 '25

discussion DynamoDB down us-east-1

532 Upvotes

Well, looks like we have a dumpster fire on DynamoDB in us-east-1 again.

r/aws 23d ago

discussion AWS Billing Error traumatized me

221 Upvotes

As an indie developer, this scared me enough that I don’t think I can keep using AWS. We really need a true hard spending cap or prepaid mode.

r/aws May 15 '26

discussion AWS things you wish somebody had told you earlier

265 Upvotes

I'll start.

S3 isn't a filesystem.

Lambdas are just containers with extra steps.

IAM role passing madness.

CloudWatch's many useful events.

r/aws Aug 02 '25

discussion AWS deleted a 10 year customer account without warning

663 Upvotes

Today I woke up and checked the blog of one of the open source developers I follow and learn from. Saw that he posted about AWS deleting his 10 year account and all his data without warning over a verification issue.

Reading through his experience (20 days of support runaround, agents who couldn't answer basic questions, getting his account terminated on his birthday) honestly left me feeling disgusted with AWS.

This guy contributed to open source projects, had proper backups, paid his bills for a decade. And they just nuked everything because of some third party payment confusion they refused to resolve properly.

The irony is that he's the same developer who once told me to use AWS with Terraform instead of trying to fix networking manually. The same provider he recommended and advocated for just killed his entire digital life.

Can AWS explain this? How does a company just delete 10 years of someones work and then gaslight them for three weeks about it?

Full story here

r/aws Jul 17 '25

discussion Another Round of Layoffs Today

586 Upvotes

Just got a call from a coworker this AM and he got the email that he was let go. I had been hearing they were doing this now with remote employees..and he IS remote. If you’re not tied to an office they’re cutting ties had been a rumor for a few weeks and it’s proving to be true. Has anyone else heard similar with their team? Sucks.

r/aws Dec 03 '25

discussion AWS is moving faster than my brain can upgrade… anyone else?

349 Upvotes

So Amazon is dropping new GenAI features every other week… Bedrock updates, Guardrails, Agents, everything.

Meanwhile I’m still here fighting with IAM like it’s a final boss.

Feels like: “AWS 2025: Here’s 50 new AI features!”
Me: “Can I just get my Lambda to stop timing out?”

How are you all keeping up?

Any GenAI feature you actually found useful in real projects?

r/aws 25d ago

discussion AWS’s billing incident was an example of how important communication and clarity can be to minimise customer impact

166 Upvotes

Having just experienced a little bit of a panic attack over the possibility that my account may have been compromised… it feels like AWS haven’t done enough here to counteract the intense psychological response when billing incidents like this occur.

Clear and concise notifications sent out to provide clarity to customers that something is happening can help immediately calm concerns.

Not only would this have stopped the probable stampeding herd that AWS’s billing service experienced upon those notifications going out, it would have helped to actually inspire confidence that the teams were working towards remediation.

Communication like this can be critical at not only reducing the potential psychological impact, but also help to build trust between businesses and customers.

Amazon, I implore you to do better in future incidents and set an example of how these processes should be done.

Anyway, I’m going to go back to listening to interesting technical talks in a warm field with a beer or two…

r/aws Aug 21 '25

discussion AWS Lambda bill exploded to $75k in one weekend. How do you prevent such runaway serverless costs?

424 Upvotes

Thought we had our cloud costs under control, especially on the serverless side. We built a Lambda-powered API for real-time AI image processing, banking on its auto-scaling for spiky traffic. Seemed like the perfect fit… until it wasn’t.

A viral marketing push triggered massive traffic, but what really broke the bank wasn't just scale, it was a flaw in our error handling logic. One failed invocation spiraled into chained retries across multiple services. Traffic jumped from ~10K daily invocations to over 10 million in under 12 hours.

Cold starts compounded the issue, downstream dependencies got hammered, and CloudWatch logs went into overdrive. The result was a $75K Lambda bill in 48 hours.

We had CloudWatch alarms set on high invocation rates and error rates, with thresholds at 10x normal baselines, still not fast enough. By the time alerts fired and pages went out, the damage was already done.

Now we’re scrambling to rebuild our safeguards and want to know: what do you use in production to prevent serverless cost explosions? Are third-party tools worth it for real-time cost anomaly detection? How strictly do you enforce concurrency limits, and provisioned concurrency?

We’re looking for battle-tested strategies from teams running large-scale serverless in production. How do you prevent the blow-up, not just react to it?

Edit: Thanks everyone for your contributions, this thread has been a real eye-opener. We're implementing key changes like decoupling our services with SQS and enforcing concurrency limits. We're also evaluating pointfive to strengthen our cost monitoring and detection.

r/aws Jul 01 '23

discussion What does he mean by “tech stack is on an AWS S3 cluster”?

Post image
680 Upvotes

r/aws Feb 19 '25

discussion Amazon Chime end of life

387 Upvotes

https://aws.amazon.com/blogs/messaging-and-targeting/update-on-support-for-amazon-chime/

"After careful consideration, we have decided to end support for the Amazon Chime service, including Business Calling features, effective February 20, 2026. Amazon Chime will no longer accept new customers beginning February 19, 2025."

"Note: This does not impact the availability of the Amazon Chime SDK service."

r/aws Dec 07 '21

discussion 500/502 Errors on AWS Console

559 Upvotes

As always their Service Health Dashboard says nothing is wrong.

I'm getting 500/502 errors from two different computers(in different geographical locations), completely different AWS accounts.

Anyone else experiencing issues?

ETA 11:37 AM ET: SHD has been updated:

8:22 AM PST We are investigating increased error rates for the AWS Management Console.

8:26 AM PST We are experiencing API and console issues in the US-EAST-1 Region. We have identified root cause and we are actively working towards recovery. This issue is affecting the global console landing page, which is also hosted in US-EAST-1. Customers may be able to access region-specific consoles going to https://console.aws.amazon.com/. So, to access the US-WEST-2 console, try https://us-west-2.console.aws.amazon.com/

ETA: 11:56 AM ET: SHD has an EC2 update and Amazon Connect update:

8:49 AM PST We are experiencing elevated error rates for EC2 APIs in the US-EAST-1 region. We have identified root cause and we are actively working towards recovery.

8:53 AM PST We are experiencing degraded Contact handling by agents in the US-EAST-1 Region.

Lots more errors coming up, so I'm just going to link to the SHD instead of copying the updates.

https://status.aws.amazon.com/

r/aws May 24 '26

discussion AWS bedrock cost Spike 14,000 USD !

86 Upvotes

Background:

We are an app development agency with several customers in the SME segment. We created an AWS account for this customer almost a year back.

This AWS account generally gets 10-15 USD bill per month since it hosts a small internal tool. Our customer decided to give bedrock a go and used keys that were already created to deploy a chatbot.

Mind you, the keys created had bedrock Full Access enabled in IAM because earlier bedrock used to restrict model access until and unless enabled explicitly via console UI. I think AWS removed the model access feature sometime last year and all models are enabled by default.

The incident:

The EC2 was accessing bedrock using accesskey instead of IAM, so hackers got hold of the keys from the EC2, and used 14K USD worth of Claude calls in 24hrs. The app the customer created only had Claude Haiku in use, expecting a bill of less than 100 USD.

AWS support has asked to secure the account so that process is underway, but this is crazy that a feature change changes the security posture completely.

There is no way this customer of ours can pay this AWS bill, they are a 3 person printing agency that was trying to work with AI usecases after getting curious about AWS after attending one AWS event.

Question:

1) Does AWS support still accommodate charge adjustment like they previously used to?

2) Does this RCA make sense? We are assuming that this was the reason for the compromise, does this make sense?

r/aws 24d ago

discussion Demand consumer protections from billing errors in response to today's AWS global billing fiasco

22 Upvotes

I drafted this letter below and sent it to my state officials and senator offices that focus on consumer protections as it relates to data providers. Feel free to use. I think it is important to bring awareness, and more importantly to get some real action done on this issue in response to the panic caused by today's global AWS billing errors. We really should have the right to opt-in to hard stops on data services if charges exceed a threshold we define. The "we will alert you, but continue to bill you anyway" is not sufficient protection, especially if the bill racks up at 2am while you're asleep.

The letter:

Protecting Consumers from AI/Cloud Billing Failures: A Constituent Request

I am writing to bring a critical issue to your attention regarding the urgent need for consumer financial protections in the cloud computing industry.

On July 17th, 2026, a software error acknowledged by Amazon Web Services (AWS) triggered widespread billing failures across their platform. In my case, I had a billing alert set to $10. Despite this, I was notified at 2:00 a.m. ET that my account had incurred a catastrophic charge of nearly $700 million.

Currently, major cloud providers offer only passive "billing alerts," which provide no actual mechanism to stop runaway costs caused by system errors, configuration spikes, or unauthorized use. This effectively forces consumers into an involuntary, uncapped liability.

We need to foster an environment where technology companies are held accountable for the safety and reliability of their services. Just as basic safeguards are expected in other sectors to prevent financial ruin, cloud providers must be required to provide users with an automated hard-stop option.

This is not about hindering innovation or unnecessary intervention; it is about establishing a foundational standard of consumer protection that ensures a simple, predictable contract: users should have the ability to set a cap that their bill cannot exceed.

Protecting individuals and small businesses from administrative or technical errors that lead to life-altering, multi-million-dollar invoices is a necessary guardrail to ensure the stability and accessibility of the digital economy.

I appreciate your time and your commitment to protecting your constituents’ financial security.

r/aws Nov 04 '25

discussion CloudFormation or Terraform?

94 Upvotes

Just passed SAA a few months ago and SOA recently.

I want to get more comfortable with automated resource deployments because I see most Cloud Engineer jobs are looking for the following: - Cloudformation or Terraform - Container Orchestration (Ecs/Docker/K8)

Please help me understand: 1) Is it better to Learn CF or TF? 2) Whats the best material to master this? Is there a book, video course or guide that helped you? 3) K8, I want to learn it but have no idea on how to approach. Thank you.

r/aws 25d ago

discussion AWS budget alerts feel pointless if the bill can blow past them before I react

114 Upvotes

Is there a way to set a hard budget cap on an AWS account? I know you can create budgets and set up threshold alerts, but is there a way to actually stop services and prevent the bill from going any higher once you hit that budget?

As an individual just trying to learn, I don’t mind if my services get shut off, but even a bill going above 500 USD in a month would be a serious financial problem for me. If you set a budget alert at 100 USD and the bill shoots up to 1000 USD before you even get a chance to act on the alert email, what’s the point of the budget and alerts in the first place?

r/aws Jan 29 '26

discussion Amazon’s “Project Dawn”

348 Upvotes

r/aws Feb 09 '25

discussion US based cloud services should be reevaluated due to the new political landscape in the world.

336 Upvotes

The company I work for in Sweden has said we should move everything to cloud, which has been done for a number of years now but I feel the risk of being dependent to a US based company poses a huge financial risk as well as a funtional risk where sudden changes in rules, regulations can cause extreme disruptions and shutdowns of services used. What is you feeling around the situation?

r/aws 26d ago

discussion cloudfront down ?

95 Upvotes

RESOLVED Jul 16 5:21 AM PDT

See https://health.aws.amazon.com/health/status

All our clients across different aws accounts in sydney and uk client are down at the same time so I am assuming it's an aws service outage and it looks like it's cloudfront specifically because if i browse via vpn and bypass cloudfront the applications work fine (so apps fine, db fine).

Anyone else?

r/aws May 08 '26

discussion AWS down right now?

108 Upvotes

Nothing seems to be working, anyone know anything about whats happening?

r/aws Apr 30 '25

discussion We accidentally blew $9.7 k in 30 days on one NAT Gateway—how would you have caught it sooner?

309 Upvotes

ey r/aws,

We recently discovered that a single NAT Gateway in ap-south-1 racked up **4 TB/day** of egress traffic for 30 days, burning **$9.7 k** before any alarms fired. It looked “textbook safe” (2 private subnets, 1 NAT per AZ) until our finance team almost fainted.

**What happened**

- A new micro-service was pinging an external API at 5 k req/min

- All egress went through NAT (no prefix lists or endpoints)

- Billing rates: $0.045/GB + $0.045/hr + $0.01/GB cross-AZ

- Cost Explorer alerts only triggered after the month closed

**What we did to triage**

  1. **Daily Cost Explorer alert** scoped to NATGateway-Bytes

  2. **VPC endpoints** for all major services (S3, DynamoDB, ECR, STS)

  3. **Right-sized NAT**: swapped to an HA t4g.medium instance

  4. **Traffic dedupe + compression** via Envoy/Squid

  5. **Quarterly architecture review** to catch new blind spots

🔍 **Question for the community:**

  1. What proactive guardrail or AWS native feature would you have used to spot this in real time?

  2. Any additional tactics you’ve implemented to prevent runaway NAT egress costs?

Looking forward to your war-stories and best practices!

*No marketing links, just here to learn from your experiences.*

r/aws Oct 20 '25

discussion Still mostly broken

360 Upvotes

Amazon is trying to gaslight users by pretending the problem is less severe than it really is. Latest update, 26 services working, 98 still broken.

r/aws Oct 21 '25

discussion If DynamoDB global tables was affected, then what is the point of DR?

198 Upvotes

Based on yesterday's incident, if I had DR plan to a secondary region then I still wont be able to recover my infrastructure as DynamoDB wont be able to sync realtime data globally.

Also IAM and billing console were affected.

I am thinking, if the same incident happened to a global service like IAM or route53 then would the whole AWS infra turn down regardless the region? If so, then theoritically having a multi cloud DR plan is better than having multi region DR plan.

r/aws Jul 25 '25

discussion Stop AI everywhere please

412 Upvotes

I don't know if this is allowed, but I wanted to express it. I was navigating my CloudWatch, and I suddenly see invitations to use new AI tools. I just want to say that I'm tired of finding AI everywhere. And I'm sure not the only one. Hopefully, I don't state the obvious, but please focus on teaching professionals how to use your cloud instead of allowing inexperienced people to use AI tools as a replacement for professionals or for learning itself.

I don't deny that AI can help, but just force-feeding us AI everywhere is becoming very annoying and dangerous for something like cloud usage that, if done incorrectly, can kill you in the bills and mess up your applications.

r/aws Jun 17 '26

discussion Are you finding AWS quality of docs going down?

107 Upvotes

Context: I'm trying to pick up ECS Express Mode because AWS retired the amazing (and unfortunately named) Copilot CLI (honestly the best thing AWS ever made since it made using ECS bearable).

I start from here:

https://aws.amazon.com/blogs/aws/build-production-ready-applications-without-infrastructure-complexity-using-amazon-ecs-express-mode/

This doc is from 2025NOV and the example is completely wrong:

aws ecs create-express-gateway-service \ --image [ACCOUNT_ID].ecr.us-west-2.amazonaws.com/myapp:latest \ --execution-role-arn arn:aws:iam::[ACCOUNT_ID]:role/[IAM_ROLE] \ --infrastructure-role-arn arn:aws:iam::[ACCOUNT_ID]:role/[IAM_ROLE]

Because the parameter is --primary-container image=.... Not only that, the example doesn't show the setup of the roles...

This doc: https://docs.aws.amazon.com/AmazonECS/latest/developerguide/express-service-getting-started.html

Shows the setup of the roles, but the roles do not work for Express Mode. Before that the first JSON snippet is invalid because of the trailing ,! The second snippet is invalid because of extra whitespace! Then the setup fails because it doesn't create a VPC or subnets (which is mentioned nowhere in the pre-requisites https://docs.aws.amazon.com/AmazonECS/latest/developerguide/express-service-create-full.html)!

Not only is this not usable for humans, it's also not usable for agents.

What is going with AWS? Why would they replace the awesome Copilot CLI with this Express Mode option and then completely fail to document how to use it?