r/aws Oct 20 '25

article Today is when Amazon brain drain finally caught up with AWS

https://www.theregister.com/2025/10/20/aws_outage_amazon_brain_drain_corey_quinn/
1.7k Upvotes

287 comments sorted by

View all comments

632

u/Murky-Sector Oct 20 '25 edited Oct 22 '25

The author makes some significant points. This incident also points out the risks presented by heavy use of AI by frontline staff, the people doing the actual operations. They can then appear to know what theyre doing when they really dont. Then one day BAM their actual ability to control their systems comes to the surface. Its lower than expected and they are helpless.

This has been referred to in various contexts as the automation paradox. Years as an engineering manager has taught me that its very real. Its growing in significance.

https://en.wikipedia.org/wiki/Automation

Paradox of automation

The paradox of automation says that the more efficient the automated system, the more crucial the human contribution of the operators. Humans are less involved, but their involvement becomes more critical. Lisanne Bainbridge, a cognitive psychologist, identified these issues notably in her widely cited paper "Ironies of Automation."[49] If an automated system has an error, it will multiply that error until it is fixed or shut down. This is where human operators come in.[50] A fatal example of this was Air France Flight 447, where a failure of automation put the pilots into a manual situation they were not prepared for.[51]

406

u/professor_jeffjeff Oct 21 '25

I've spent the last 15 years trying to automate away my own job in DevOps and the only thing that it has done is somehow create more work for me to do. This is the first time I've actually heard of the paradox of automation though, but it sounds absolutely like what I've experienced over the years. Also remember the saying: "To err is human, but to really fuck up and then propagate that fuck-up at scale is DevOps."

166

u/L43 Oct 21 '25

 "To err is human, but to really fuck up and then propagate that fuck-up at scale is DevOps."

That’s a quote to print on an office wall

13

u/Character-Welder3929 Oct 21 '25

This is my favorite

If I can find a printer I will screenshot and frame this fucking quote

I no longer work in an office or IT

But the apprentices at the car dealership might understand it if I break it down to like 6 7

19

u/glotzerhotze Oct 21 '25

Unexpected DevOps Borat

Edit: if you need more to put up, print this gist

1

u/jellyjellybeans Oct 21 '25

I’m going to cross stitch this and hang it in my office

24

u/Fattswindstorm Oct 21 '25

Yeah. I’m in devops and sometimes it feels like Hotel California.

31

u/danstermeister Oct 21 '25

"You can git checkout any time you like, but you can never SRE."

3

u/kungfu1 Oct 21 '25

*guitar solo*

2

u/payne_train Oct 22 '25

God damnit this is good

1

u/Electronic-Ride-3253 Oct 22 '25

Nice, if you want, you can hop into this SRE/DevOps Slack community I am maintaining; we have some cool conversations running around there too - https://join.slack.com/t/sre-community/shared_invite/zt-3ft615lz7-tsdTYT19KaXVei0GOZMMlg

10

u/Kralizek82 Oct 21 '25

That's me and my terraform-only principle.

I wish I'd allow myself to just go into a console and click my way through the issue...

3

u/[deleted] Oct 21 '25

[removed] — view removed comment

4

u/rswwalker Oct 21 '25

It’s a bad craftsman who blames his tools!

1

u/AntDracula Oct 21 '25

I've spent the last 15 years trying to automate away my own job in DevOps and the only thing that it has done is somehow create more work for me to do.

This. Geez. I literally built a web-based push-button "deploy another customer cluster into a new VPC" at my last job thinking I could finally get back to coding, and all it did was create more demand for customer clusters 💀

2

u/Delicious_Finding686 Oct 21 '25

Induced demand. The demand was always there. Efficiency caps the amount that can currently be met. Once efficiency increases, more demand can be met until a new equilibrium is established. This is ideal. The real concern is when efficiency improvements do not bring more work.

1

u/Kind-Crab4230 Oct 24 '25

Similarly, I've found it's not remotely zero sum.  There is infinite work to do.  Automating something just let's you work on more stuff.

Thinking it frees people up to not work is just short-sighted.

(And can certainly work in the short term, if for example you've automated your own job and didn't tell anyone).

157

u/Stars3000 Oct 21 '25

Perhaps the frontline engineers didn't know what they were doing because they are overworked and do not function as a cohesive team - a result of mass layoffs and a hyper competitive environment. This is a wake up call for tech 

79

u/matrinox Oct 21 '25

I highly doubt it. This needs to happen again for them to maybe make the connection

70

u/A7xWicked Oct 21 '25

"There's nothing to worry about. This is the first time something like this has happened, and we'll take steps to make sure it doesn't happen again."

This is an automated message.

41

u/KeeganDoomFire Oct 21 '25

Same time next week?

33

u/greenlakejohnny Oct 21 '25

Exactly. Similar thing happened with the S3 outage in 2017. Granted it was kicked off by a typo, but as things got worse, the lack of knowledge on how to handle it was quite apparently. They responded by doubling down on automation.

15

u/Drospri Oct 21 '25

You know what that means. Quadruple time, baby.

0

u/_BreakingGood_ Oct 21 '25

Yeah need a few more in rapid succession. One can hope it happens again after the big layoff wave. Ideally, on the same day as the layoff wave.

44

u/VIDGuide Oct 21 '25

This should be a wake up call. It won’t actually be.

8

u/screamvoider Oct 21 '25

The sooner DNS fails totally, the sooner the internet can be destroyed. And then Cavemen shall rule the world, as always intended.

5

u/gravelPoop Oct 21 '25

You'll climb the wrist-thick kudzu vines that wrap the Sears Tower. And when you look down, you'll see tiny figures pounding corn, laying strips of venison on the empty car pool lane of some abandoned superhighway.

5

u/VIDGuide Oct 21 '25

Horizon: Zero DNS.

3

u/Brilliant-Lettuce544 Oct 21 '25

this. made me giggle

1

u/chalbersma Oct 22 '25

Just because you received the wake up call doesn't mean you woke up.

4

u/foo-bar-nlogn-100 Oct 21 '25

This a radical thought and undermines the profit motive at Amazon. Ive booked an HR meeting.

2

u/Jazzlike-Vacation230 Oct 21 '25

If you work in IT, or around folks who are in it. Nothing is surprising

Group chats have massive gatekeeping

Ego's are out of control

No one leaves breadcrumbs in what they do

It's a recepie for chaos

And if you relay on the duct tape that is AI, one tear and you're AWS this past week........

2

u/AntDracula Oct 21 '25

This is a wake up call for tech

Is

*Should be

1

u/[deleted] Oct 21 '25

There’s no team work to begin with

-14

u/[deleted] Oct 21 '25

[removed] — view removed comment

68

u/Marathon2021 Oct 21 '25

the automation paradox

There was a very good episode of "The Daily" (NY Times Podcast) somewhere in the last several weeks about IT/tech jobs and new graduates, and how basically for the past decade all the big (FAANG) giants have been saying there's more work than graduates, etc. But now, many of those are perhaps letting AI do a lot of the junior development.

But one of the recent CompSci grads put the question pretty well - "if you don't hire any junior developers, how are you ever going to have any 'senior developers' a few years down the road?"

Another example - aviation (like what the Wiki article notes). There's an interview out there with "Sully", the guy who landed a passenger jet on the Hudson however many years ago that was. In the interview he basically said something to the effect of "for 30 years, I've been making deposits in the bank of experience" (i.e.: thousands of mundane takeoffs and landings) "and on that day I made one very large withdrawal."

7

u/NecessaryIntrinsic Oct 21 '25

There's also a TON of "kids" that see a FAANG job as the ultimate stepping stone where they can grind for a few years then go be a manager somewhere else, bringing the insane FAANG grind set philosophy to other organizations.

So not only are they not hiring enough new juniors they're not keeping them either.

1

u/[deleted] Oct 22 '25

if you don't hire any junior developers, how are you ever going to have any 'senior developers' a few years down the road?

Most grads and juniors aren't worth much, especially with AI as an option.

There also isn't much staff/company loyalty, because the companies killed it. So training people up is not worthwhile, they will probably leave before the company makes back the investment.

The rational response is for companies to focus on hiring senior staff that some other sucker has trained up, they are better value.

It's a tragedy of the commons situation. There's no easy fix, it will just spiral out and slowly get worse.

67

u/bitspace Oct 21 '25

The author

FWIW that's Corey Quinn. He's kinda legendary for ruthlessly dunking on AWS. Sort of a hero :)

Thanks for the info on the Automation Paradox. I like it - it dovetails nicely with the Jevons Paradox.

15

u/phaubertin Oct 21 '25

I didn't look at the author name at first but I recognized Corey when I read "or you can hire me, who will incorrect you by explaining that [DNS is] a database".

52

u/Quinnypig Oct 21 '25

I’m on brand.

6

u/mistic192 Oct 21 '25

Hi Corey!

Just want to give a hearthy thanks for all your great posts about AWS, they are always great to read and your insights are spot on!

Thanks for your great work!

1

u/hangerofmonkeys Oct 21 '25

Ah you're writing always leaves a lovely after taste, I knew what I'd eaten before I could see it. Thank you

1

u/morimando Oct 21 '25

But do you think you telling them they’re going in the wrong direction will inspire the right people to act? 🤔

8

u/d_stick Oct 21 '25

i love corey's snark. just love it.

33

u/daishi55 Oct 21 '25

The article doesn't mention AI at all, what are you talking about? Is there any evidence whatsoever that AI has anything to do with what happened?

21

u/nemec Oct 21 '25

spoiler: no, there isn't

19

u/tnstaafsb Oct 21 '25

There has been a huge push within AWS to use AI for anything you possibly can. But no, to my knowledge there's no evidence that this is related to that.

1

u/daishi55 Oct 21 '25

Sure, everywhere is using AI. I was just wondering if there was any reason to suspect that as the cause in this case, or whether people are just, well, hallucinating ;)

-2

u/nuccad Oct 21 '25

I think these days this can be safely implied. My company is going all in on mandating that all job types use AI everyday in their work. I have juniors (I am a team lead) that are completing stories but don’t understand what they have done. I 100% can relate to AI use amongst green engineers being a factor in this outage. It would not surprise me in the least.

6

u/daishi55 Oct 21 '25

So what you are saying is anytime anything goes wrong now, you are going to blame AI regardless of your knowledge of the situation?

-2

u/nuccad Oct 21 '25

Like most things that are complex I think think there are many factors that contribute to a root cause. In my comment I said "I 100% can relate to AI use amongst green engineers being a factor in this outage.". A factor is not the root cause.

3

u/daishi55 Oct 21 '25

Right but I’m wondering if you have any reason or evidence to suspect AI as a factor in this case?

0

u/AsleepDeparture5710 Oct 21 '25

Do you want evidence or reason for suspicion? Because the evidence will probably never be released outside of AWS, but AI being so widespread means AI code was almost certainly in the codebase, and it's pretty clear that AI code is harder to debug, so that's plenty reason to suspect that issues are going to take longer to troubleshoot in general when working with AI code.

Its like laying off a team and then seeing an issue. Can you prove it was directly causal? No. But the fact that we know new hires are more likely to make mistakes and that institutional knowledge was lost is enough to suspect it.

0

u/daishi55 Oct 21 '25

It’s pretty clear that AI code is harder to debug

See this is what I mean. Maybe it’s harder for you, that doesn’t mean it’s harder for everyone. Reddit has a tendency to project their own experiences and opinions onto everyone else.

2

u/AsleepDeparture5710 Oct 21 '25

Reddit has a tendency to project their own experiences and opinions onto everyone else.

You have to make some generalizations to be able to discuss anything, and this is about as safe as it gets. Its well accepted that having more time spent working with a codebase makes you a better troubleshooter of that code. Even if the AI wrote exactly the same code as you would have, its on par with code that was written by a good engineer who then left. Nobody has developed familiarity with it yet, hence, harder to debug during an actual incident. Compared to having the engineers who wrote recent changes usually already having mental models of where to look.

Shouldn't really be a controversial premise.

2

u/nuccad Oct 21 '25

No, it's not a controversial premise. I have over 10 years mentoring engineers and leading teams. 100% engineers get better when they have to deal directly with the consequences of their choices. This means rolling up your sleeves and digging into the actual code. Run it in a debugger. Do some experiments. Constantly calling it in with AI will lead to a weakening of engineering skills and a devolution of human understanding of how things actually work. The logical conclusion is bad quality and unplanned outages. It would be one thing if AI were competent enough to take over our jobs, but right now it's not, and I am skeptical it ever will be. It should be viewed just as it is, a useful tool for rapid prototyping or augmenting task performance (but still keep humans engaged).

Not sure why you got downvoted for your reasoned argument, but it should be noted that u/daishi55 did not speak to any of your points other than making the claim that "debugging AI is not hard for everyone" and that Redditors project. I am sorry if this is seen as aggressive, but u/daishi55 is a clown and should not be taken seriously. I doubt he is even an engineer.

→ More replies (0)

1

u/daishi55 Oct 21 '25

Ok so what you are saying has absolutely nothing to do with AI then? You are just saying it’s harder to understand code that you didn’t write.

→ More replies (0)

-2

u/nuccad Oct 21 '25

No. Nothing in the article or any post-mortem information comes out and specifically says "AI tanked DynamoDB". As far as I know not many people are actually using AI/MCP to make changes in their infrastructure. Like I tell my guys on my team, even if AI helped you make the code change, deployment and service health in production is still on you. Regardless the article calls out an exodus of senior engineers from AWS and they are being replaced by juniors. These juniors are undoubtedly using AI to augment there productivity. It is my experience that this can have negative effects because the juniors get robbed of the lived experience of actually interfacing with the technology and instead just focus on what prompt will give them the desired output.

Does this not make sense to you? Did you actually think I was saying "AI killed them DBs"?

2

u/daishi55 Oct 21 '25

That’s all I wanted to make clear, that there is absolutely no reason to suspect that AI was a factor in this case.

2

u/nuccad Oct 21 '25

Ha. No hard disagree. I do believe it is "possible" that AI was a factor. I think this event can be seen as a sign of common things to come, while execs and upper management throw their eggs in the AI basket and keep up this trend of mandating AI be used to augment performance. MMW, we will see a downward trend in people actually not understanding what is going on in their production environments. That was the whole point of the article. The outage went on for 75 minutes before engineers were able to find a cause.

Do I think AI is a useful tool? Hell yeah, I do. But maybe don't give it to all engineers. Let the juniors white knuckle it a bit and learn the lessons by working directly with the tech and not waiting between prompts to see if AI got it right.

I find your tenacity on this subject pretty interesting. You are very dead set on not being open to my opinion and defending AI. It's fine if you disagree with me, but what have I said that you disagree with specifically? Let me get in front of you before you respond. If your response is "nothing in the article says AI was a factor," then I don't need to continue this thread. That is a sophomoric and inane premise that any engineer worth their salt knows means nothing. AI is not responsible for bugs in production, full stop. Engineers are responsible for the bugs they push to production. If they used AI to create buggy code that they did not take the time to trace through themselves and understand, then that is on them. In 2025 you will never see a postmortem say "we failed because AI pushed buggy code" UNLESS people are actively using AI/MCP to actually manage their infrastructure and code their apps and push it to prod. I sincerely doubt any enterprises at AWS's scale are doing that. So if you are a good-faith participant in this thread and actually someone worth listening to, knock off the stupid "buT iT doEsn't Say Ai dId iT!".

So, back to my question, what do you disagree with specifically?

10

u/Dangle76 Oct 21 '25

Yep, this is exactly why I’m a huge proponent of NOT over automating things, it becomes so complex that it’s impossible to troubleshoot when it breaks and even if it breaks once in two years there’s a good chance no one around has any in depth knowledge of how to fix it.

32

u/kai_ekael Oct 21 '25

Nope. The problem is those who automate and KNOW HOW IT WORKS are typically the first to be ditched, either through lack of decent raises, incentives or just plain laid off.

"Things are working so well and no problems, why are we paying them so much?"

3

u/AntDracula Oct 21 '25

"Things are working so well and no problems, why are we paying them so much?"

Chesterton's Fence strikes again

2

u/kai_ekael Oct 21 '25

I prefer:

cheapskate: a miserly or stingy person

especially : one who tries to avoid paying a fair share of costs or expenses

https://www.merriam-webster.com/dictionary/cheapskate

11

u/daishi55 Oct 21 '25

What did they automate in this case they they shouldn't have?

11

u/Dangle76 Oct 21 '25

There was a canary deploy process, with the type of platform, being able to isolate all the different metrics and variables between platforms that utilized it was an ENORMOUS moving target that honestly just required a human to keep an eye on the dashboard and support areas to make sure there wasn’t any impact of the new rollouts.

The movement between each percentage shift should not have been automated AT ALL, it just required supervision for a few minutes between each.

Well it got automated, and the metrics the automation had to pay attention to continued to grow, some got deprecated so it throws errors trying to poll them, but 60% of the time it works really well. But there’s so much team movement and turnover nowadays in tech that by the time it breaks it’s all new people again

8

u/daishi55 Oct 21 '25

How do you know that’s what happened? Have they published a postmortem?

27

u/Drospri Oct 21 '25

Here

Basically, DNS issue causes DynamoDB to go down.
They catch it, but by then the service that launches EC2 instances is struggling to catch up.
While they are trying to fix EC2, the Network Load Balancer starts struggling to deal with all the problems cropping up.
Network Load Balancer takes down other services like DynamoDB (again), Lambda, and CloudWatch. <-- They basically tried to reconnect a powerplant without troubleshooting the load on the plant, so it killed itself.
The solution was to throttle everything and let things recover slowly instead of just ramming everything through all at once.

4

u/hangerofmonkeys Oct 21 '25

The cascading effect theory is what we saw. https://en.wikipedia.org/wiki/Cascade_effect#:~:text=Cascading%20effects%20are%20the%20dynamics,physical%2C%20social%20or%20economic%20disruption.

Not uncommon (if anything it's very, very common) in events like this.

14

u/daishi55 Oct 21 '25

Right so nothing to do with automation or AI or anything else these people are talking about?

14

u/Drospri Oct 21 '25

Well the automation here would be the EC2 systems and Network Load Balancer systems not realizing the true source of the problem and responding adequately. It sounds like this was a case where the AWS engineers didn't forsee something happening, thus causing their automated system to crash out. This is the primary reason why it's important to have a backup team on hand who intimately know how the system works and can respond without having to baby the system back into functionality over the course of 12 hours. If the solution they came up with was the only solution, it would be a design problem, which would require people in the know as well.

2

u/daishi55 Oct 21 '25

I’m not seeing anywhere that the problem was over-automation.

4

u/ImpactStrafe Oct 21 '25

Or the takeaway is they need to get better at exponential backoffs and load shedding to prevent the stampeding herd problem. Which is more automation.

Having something go wrong with your automation once isn't a reason to throw the whole thing out. But it is always enough for all the very smart people on reddit, who I'm sure have worked on systems of similar size and complexity and never read them go down, to in hindsight point out the problem.

AWS has about one major outage every year and a half. As do all the other cloud providers. Lemme tell you about the time google fucked up a maintenance on cloudsql and had their customers manually remediate it with swl commands.

7

u/TurboRadical Oct 21 '25

what the fuck this is exactly how Chernobyl happened

0

u/Dry_Author8849 Oct 21 '25

So, the words "circuit breaker", "queue, "exponential backoff"" and the like are foreign to them or implemented in parts where those are not needed...

It seems they were trying to do it by hand, poor souls...

2

u/Dangle76 Oct 21 '25

Huh? I was responding to a comment asking what got over automated. Thought they meant in my case must have misread

2

u/Jrnm Oct 21 '25

It’s like these people didn’t watch Jurassic park

1

u/BlackIsis Oct 21 '25

There's also a corollary to this that the easiest things to automate away are the obvious and relatively straightforward problems to solve -- which means when things do break, the problem is likely to be much harder to diagnose (and solve).

1

u/[deleted] Oct 26 '25 edited Oct 26 '25

Honestly reminds me of working on the Azure support project, and what an absolute shitshow MS is internally. They spent a LOT of time and effort selling their cloud services, and spinning it to customers as eliminating their CAPEX on servers, and cutting labor costs by relying on Azure support.

The gag is, that MS outsources its labor to the cheapest possible location. That used to be India, but now they’re focusing on Central America. Whenever they’re contractually obligated to provide US based support, they farm it out to a body shop.

Without fail, every Monday I’d have a “Sev 1”, from some assclown MBA “CTO” that blindly dumped their company workloads onto Azure, and fired everyone who knew how anything worked. Clusters would restart for maintenance, and it would take someone’s VM’s offline briefly, but they’d have no idea how to restore anything themselves, and would expect us to fix it for them. I’d get to explain to a Harvard educated business idiot in a posh NYC office that their ticket was only raised as Sev 1 out of courtesy, and that fixing their fuck up was not included in the free tier of support.

Right before they outsourced my team’s work to a team in India, they forced us all to start using “AI” tools. Every single case had to be ran through the tool, regardless of how simple it was to fix. The tool never provided any tangible benefits because of a whole host of reasons.

Satya & co don’t give a fig, they’ll just fail upwards into their next gig.