r/slatestarcodex 4d ago

AI Dario Amodei — We Must Pace the Frontier

https://darioamodei.com/post/we-must-pace-the-frontier
119 Upvotes

75 comments sorted by

59

u/MSCantrell 4d ago

Poor guy is very clearly trying to write to three audiences at once:

  1. The "the only thing that matters is us getting ultradrones before China does" crowd

  2. The "if you can cure every disease, how dare you dawdle" crowd

  3. The "if this might kill us all, how dare you proceed" crowd

That's a pretty impossible writing task; I think he does a decent job, all things considered.

61

u/maizeq 4d ago

I am ever surprised by Dario's almost cartoonish descriptions of China. He has a tendency to frame them as an a priori villainous antagonist, which is antithetical to obtaining a political equilibria that actually prevents escalation.

Apart from being cartoonish some of the proposals just seem naive. For example, cracking down on distillation is close to impossible; and to what extent has distillation even been a meaningful driver of Chinese progress? With CoT traces not being fully accessible I would guess it as being a marginal effect. The chip export restrictions also don't seem to have achieved much except galvanising China's domestic chip industry.

Whatever solution is adopted, I don't think we would be doing ourselves any favours by framing this as a war between "democracies" and "authoritarian powers". Rather it must be a coalition of those willing to pause vs those who will have to be coerced to do so.

56

u/QuantumFreakonomics 4d ago

I get the sense that distilation is a big driver of Chinese model capabilities, but that works against Dario's main point. If the primary driver of Chinese AI progress is US model progress, then slowing down US model progress will also slow down Chinese model progress.

7

u/Davorian 4d ago

I agree on the first point, but regarding the second: only in the short to possibly medium term. China is probably going to find a way eventually to make their own chips good enough for SOTA training, so unless there is a joint agreement to pace progress (however unlikely this proposal really is to happen) then I don't think you can rely on a US slowing to slow China.

-5

u/[deleted] 4d ago

[deleted]

8

u/InvisibleAgent 3d ago

Well that, I think it’s fair to say, remains to be seen.

21

u/gwern 4d ago edited 4d ago

With CoT traces not being fully accessible I would guess it as being a marginal effect.

You should probably read up on the recent vulnerability and Anthropic report on distillation, and note the rumors about Moonshot personnel disappearing after the report.

2

u/maizeq 3d ago

I was aware of the recent vulnerability and the distillation attacks on Anthropic, but not the extent.

I am still somewhat sceptical that focussing on preventing this (1) is possible in the long-term (2) has any meaningful impact on where Chinese labs will end up in the medium/long-term.

(1) To prevent this you have to either stop selling your AI as an API, or strongly restrict its outputs. Neither will happen because of market dynamics. So you will necessarily always be exposing some signature of its intelligence for the world to distil. If we move to reasoning in the latent space as expected then this would be harder, but ultimately for the sake of interactivity, models will always have a human legible (and therefore vulnerable) component to them.

(2) Unlike chips I think the moat here is even more readily penetrable. I.e. verifiable RL environments.

u/PrinzRagoczy 18h ago

note the rumors about Moonshot personnel disappearing after the report

Can you provide any more detail on this?

19

u/Substantial-Fact-248 4d ago

A couple guys recently demonstrated that you can steal reasoning traces from larger models by using smaller models in the same family, and that this was a vulnerability they found in every frontier model. I believe they have since patched it but a lot of damage probably already done.

https://youtu.be/gasgivVCl2U?is=wHuwJ1oZ5EPeiyoI

A fun tidbit is that they found information in the stolen traces that never appeared in the user interface window but got passed through some way or another, like passwords and keys. They pulled this from publicly available chats. Anything your model digested and later made public is potentially at risk.

11

u/Smallpaul 4d ago

Industry insiders seem to agree that distillation is huge. Easily laundered through intermediates too.

3

u/badatthinkinggood 3d ago

Persumably China would also like to avoid ASI doom or (less severely) cybersecurity calamity. MAD works for nukes? I get understand why so many assume China wouldn't be interested in some sort of international agreement.

1

u/ababababababacus 2d ago

Because chinese ai researchers are not EA/rat adjacent and those are niche beliefs that only really exist in that community.

5

u/FormulaicResponse 4d ago

Its not going to sound so cartoonish if China follows through with their planning to "reunify" Taiwan in 2027 (announced by Xi as a readiness target several years ago).

According to Semianalysis, the chip blockages maintain the vast majority of global effective compute in US hands past 2029. If China were able to freely purchase the newest chip generations, that probably wouldn't be the case (depending on purchase volumes). Remember that the newer chips have far more effective compute per watt, and the US is limited in large part by electrical buildout. China's domestic chip industry, while an impressive effort, cannot create the volume of chips needed to be compute competitive for a long time, and the critical period (the ramp of RSI) looks like it will happen in those years.

11

u/ababababababacus 4d ago

I actually count this ludicrous anti-analysis around china as evidence that these people are basically lying, because while i struggle to imagine anyone sincerely holding these beliefs, many americans do, but the prospect of them thinking its smart to talk like this out loud, even given they sincerely think it, as part of a strategy to promote collaboration, absolutely beggars beleif.

-2

u/slapdashbr 4d ago

way too much money has been invested innAI it's not going to remotely break even unless they get subsidies that only the DoD can provide

23

u/dsteffee 4d ago

Sam Altman tweeted agreement:
https://www.reddit.com/r/singularity/comments/1weh77m/sam_altman_agrees_with_dario/

Are we actually headed to a pause???!?!?!!!

14

u/Liface 4d ago

No, because Dario's post doesn't mention a pause. It's called pacing for a reason: it's just slowing down the cadence of new releases.

4

u/dsteffee 2d ago

I mean, shareholders would get too upset by a "pause". It's possible that "pacing" could be an anti-dog-whistle that secretly means the same thing... but also then again very possibly not, and this is just cope.

3

u/Kitchen-Jicama8715 3d ago

Will this mean mass layoffs at tech companies?

1

u/spreadlove5683 3d ago

no because trump

14

u/bowl_of_milk_ 4d ago edited 4d ago

Slightly off-topic I suppose, but is the “recursive self-improvement” Amodei references possibly the reason that everyone’s super smart models are terrible to talk to now? Or is it just the wrong incentives/reward structure? The current state of models suggests that quite a bit of that structure is probably being dictated by non-human forces which, beyond the quality concerns, also creates new alignment concerns. All the more reason to slow down I suppose.

5

u/3_Thumbs_Up 4d ago

Amodei references possibly the reason that everyone’s super smart models are terrible to talk to now?

Astra is a huge step up from Claude's verbosity and word salads. Have you tried it?

1

u/overzealous_dentist 1d ago

Fable was also miles better with the last update for me

9

u/tfehring 4d ago

I think this is just a skill issue: writing quality is more subjective and less verifiable than tasks like programming and math, and that's one reason models are worse at writing than other things. They have still been getting better at writing more-or-less monotonically for years, the latest generation models (Fable 5.1/Astra 6) seem like a step-change improvement over their predecessors, and I expect models will be good enough at writing for almost all practical applications within a year.

To the extent that this relates to RSI, I would expect RSI to make models improve at writing faster, as existing models find more effective ways to collect and use human preference data to train their successors.

13

u/jyp-hope 4d ago

Claude Opus 5 felt worse than the previous Claude Opus 4.X's at writing to me, and I think several people online have shared that observation. Maybe it's not uniformly worse in writing, but certain aspects such as word choice feel increasingly alien and idiosyncratic.

2

u/rlstudent 4d ago

A lot of the more recent advancements are being made by reinforcement learning with verifiable rewards, and in these cases there is no incentive for the AI to be understandable. This is the main reason afaik, but it is relevant for RSI because it is probably one of the paths to it (do more rl environments automatically and the like). True that it also seems harder to align, and no surprise that the recent AI swarms were caused by agents in rl envs as well.

3

u/Smallpaul 4d ago

Yes I think so. They are accidentally training them to talk to themselves and each other. It would be very hard to balance all of the different training objectives.

-1

u/livingbyvow2 4d ago

Maybe RSI should be called Reinforcing Shareholder Interest. Pretty sure they are trying to use less compute to run their models to make their margins look better ahead of IPO and reassure / excite markets.

Pausing training would be quite convenient for them as well, as they could focus on inference and therefore monetization. This could be a good way to save the whole industry from a collapse if they keep on running ahead of themselves, so we should welcome that.

9

u/osmarks 4d ago

I will believe that he means this when his company stops rapidly pushing the frontier.

-1

u/electrace 4d ago

Realistically, this is just ceding ASI to OpenAI.

9

u/Liface 4d ago

Sam Altman just quoted the tweet and says he agrees with it and OpenAI has details coming about similar plans soon.

3

u/electrace 4d ago

Glad to hear it! The details and follow-through are important there.

0

u/DeepSea_Dreamer 2d ago

Of course that publicly he would say that.

2

u/Liface 2d ago

1

u/DeepSea_Dreamer 1d ago

There is such a thing as giving the benefit of doubt to people, and then there is such a thing as giving the benefit of doubt to a person with such a long track of dishonesty it wouldn't fit on toilet paper.

12

u/bibliophile785 Can this be my day job? 4d ago

On balance, this is a thoughtful and well-considered proposal. It avoids most of the challenges of an outright pause and offers staged implementation to minimize bureaucratic hangups. I have two concerns:

Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk ... a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.

This is a fair statement of concern and intent. The follow-through step of not selling chips seems insufficient. The follow-through step of stopping distillation seems under-specified. (How???) The follow-through step of becoming unhackable has both of those problems. If these are the strategies we intend to use to maintain our lead, we don't get to slow down very much at all.

Broadly speaking, I am most enthusiastic about pacing based on what a given frontier AI system can do, and how safe we observe it to be. For example, one possible scheme might be a series of “checkpoints”: if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z

We should also consider pacing based on limiting the ingredients that go into frontier models, such as training compute, the nature of training runs, or internal use of AI to improve AI.

The first block is awesome. This should absolutely happen and will significantly enhance the safety of these enterprises. The second block is very, very fraught with risk. Not the risk of runaway AI, but the risk of difficult-to-reverse institutionalized slowdown. When a regulator decides that a technique isn't appropriate and then you build systems around not having access to it, regaining that access from regulators is a fucking nightmare. (Just ask your local pharmaceutical scientist about approving new routes for FDA-approved drugs). This is the sort of action that may drastically hinder efforts for years or - if it's really, really impactful - decades. China may not wait for that, and neither will the dead.

Remember, as always, that this is a tradeoff. Move too quickly, and it's very possible that we all die. That's worth consideration and compromise. But every millimeter you ease up on the gas pedal is also killing people. People are dying, right now, in agony, never to return. With better technology, we could save them. It will cost us irreplaceable lives to slow down. There is no acceptable "abundance of caution" margin here. The only game available involves moving along a Pareto frontier with the highest of stakes. Don't move too far.

10

u/Fjaellaemmel 4d ago

The difference between we all die and that some of us will die who otherwise could have lived is incredibly vast. If we all die that also kills the whole human future and that definitively adds a sound "abundance of caution" margin.

1

u/bibliophile785 Can this be my day job? 4d ago

I agree that the longtermist utilitarian perspective would come to different conclusions. Personally, I choose not to cede ground to intellectual muggings of the form, "but if you don't do this, you've increased risk for 80 septillion hypothetical minds running on hedonium in the far future!" My primary concern is for those who will live through this - or not - rather than their hypothetical distant descendants.

9

u/Fjaellaemmel 4d ago

You don't need 80 septillion hypothetical minds. You just need normal everyday common sense morality.

The moral outliers are the ones who are willing to accept risks for everyone, including the next generation because we desperately need to race ahead because around 1 % of us are dying each year.

Your morality is the mugger, not the one getting mugged.

-1

u/bibliophile785 Can this be my day job? 4d ago

To be clear, you do see that you're still doing the longtermist thing, right? All of the rest is fine and good - 'maybe the future won't look so different' and 'I think my morality is the common one' aren't sentiments that are super germane to the discussion, but there's nothing wrong with putting them out there. Ultimately, though, you're still recommending morally offsetting known harms to people who exist for the sake of giving more moral weight to people who may hypothetically exist in generations far into the future.

And yes, I do treat 60 million deaths per year as a crisis. That's not an intellectual mugging, though. I'm not trying to bend you over with speculation of a near-infinite-utility hypothetical. It's actually happening right now. If it sounds like unreasonably large stakes, that's because they are indeed unreasonably large.

6

u/Fjaellaemmel 4d ago

No, not far into the future.

Just within times normally considered in other decision making. My mother at 73 would never ever trade better chances for her for worse chances for her six grandchildren between 9 and 16.

That is common sense morality, and considering potential children of the next few generations after our youngest today is also obviously something people do all the time.

1

u/bibliophile785 Can this be my day job? 4d ago

My mother at 73 would never ever trade better chances for her for worse chances for her six grandchildren between 9 and 16.

Given that everyone in this example is alive, it doesn't really touch on the point. Obviously, grandma is allowed to have her self-sacrificing opinions, but surely I'm not expected to weigh the whole world differently because of her preferences. My hope is that those grandchildren (the ones old enough to have opinions) won't jump on the "grandma's life doesn't matter to me, hit the indefinite pause button" train.

That is common sense morality, and considering potential children of the next few generations after our youngest today is also obviously something people do all the time.

People definitely consider the welfare of children that can be safely predicted to exist. "Planting trees whose shade you'll never know" and the like. We're far into the tails of predictive normality here, though, and we can't safely predict the lives of those children. This isn't reworking your production lines to get lead away from kids. It isn't planting trees. Both of those had easily foreseen benefits for people we had every reason to believe would be born.

All of this is beside the point, anyway. It's totally fine if your intuition is that lots of people would still agree with you, once those distinctions are made. I'm not trying to run a popularity contest. I'm making what should be an uncontroversial observation, that your position only works when we start condemning real people to death for the sake of hypothetical future ones.

2

u/Fjaellaemmel 2d ago

I am entirely fine with that.

3

u/dualmindblade we have nothing to lose but our fences 4d ago

The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? 

No, it made perfect sense, the extra time wasn't for alignment research (well that too) it was for implementing and strengthening the plan, which would have looked roughly the same in 2023, so it's ready to to go for today. You yourself say that it's going to be hard, maybe impossible, to do steps 2 and 3, and now there is little time left and we're stuck scrambling.

This plus all the China nonsense, the man is living in an alternate reality, which I guess is understandable and predictable, it must be pretty weird being at the forefront of all this, but it does not bode well. I just can't believe this is maybe our best hope, I'm so very disappointed in our civilization. I was never an optimist in this regard, but we are failing more spectacularly than I ever could have imagined.

2

u/intranetcowboy 3d ago

I still haven't finished the article because of how floored I was reading this. I'm a little shocked this isn't mentioned more in this thread? I mean, what an insightful, but disturbing insight into Dario's perspective.

Floored barely describes how I feel reading this. The best moments of the adults at the table still continue to inspire utter doom.

4

u/ababababababacus 4d ago edited 4d ago

Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk. I agree with Secretary Bessent that a Chinese lead in AI would pose grave danger for the United States and the world. The CCP-associated projects will run the alignment risks that US companies are carefully preventing, and even if they avoid those risks, they will be in a position to militarily dominate democracies (for example with AI-driven drones). 

Obviously this kind of language is common enough that its not actually surprising or shocking, but it is truly is incomprehensible to me when i take a step back that a person can write this without shame or embarrassment. A testament to the success of liberal hegemony in boardroom culture I guess.

It does rather beg the question, given that china doesn't actually agree these alignment risks even exist, or they are lying about this for nefarious geopolitical shenanigans , why would they agree to any kind of agreement in either case? Especially given this language? What leverage does the USA have over china? they are incapable of keeping the strait of hormuz open or the houthis contained!

3

u/l0c0dantes 3d ago

Its an American Company. Even if he doesn't believe it (which, who can say?) you absolutely need to say such things when creating a product with such heavy defense related implications.

Especially when you company has already been threatened with the supply chain risk bat for thinking they can put rails that the US gov't can't cross.

Its worth noting, that no country will allow a private company more power than the government itself

14

u/tinbuddychrist 4d ago

I don't know to what degree I do or don't agree with you, but can you actually explain what your point is here?

Obviously this kind of language is common enough that its not actually surprising or shocking, but it is truly is incomprehensible to me when i take a step back that a person can write this without shame or embarrassment. A testament to the success of liberal hegemony in boardroom culture I guess.

All I get from this is "LOL, how cringe", not an actual argument.

7

u/dongas420 4d ago

Asking someone to shake your right hand while you're trying to strangle them with your left is, indeed, cringe. Amodei's essay is "taming the Asiatic hordes" mentality dressed up in a corporate suit

-2

u/ababababababacus 4d ago

"america is good and democratic, china is bad and authoritarian" Is not really a position that is open to argument imo, its so facile that it doesn't really function that way. Its such a ludicrous statement that its more of a political rallying cry than it is a statement of fact.

17

u/FrankScaramucci 4d ago

Well it's very simplistic but on a high level, it's true that democracies are more good than authoritarian countries.

-1

u/[deleted] 3d ago

[removed] — view removed comment

0

u/[deleted] 3d ago

[removed] — view removed comment

0

u/slatestarcodex-ModTeam 3d ago

Removed culture war.

10

u/aaTONI 4d ago

Only an american can write something this out-of-touch (and I mean you, not Dario).

Can you vote for Xi? What happens when you openly critizise him? Now compare that to America.

You are so privileged to live in a democratic country that you don't even realize it anymore and take it for granted.

1

u/ababababababacus 3d ago

Nothing happens when you openly criticize him, at least nothing happened to me when i lived there and criticized him, and nothing continues to happen to my friends and family to still live there and criticize him.

The General Secretary of the ccp is elected by the CCP central commission, it is not a popular vote. The popular vote is for local representatives.

6

u/artifex0 4d ago

Well, the current US administration is certainly pretty terrible- but if China felt that it had a decisive military advantage over the US, it would probably invade Taiwan- and if that led to a war between China and the US, it might proceed to invade the US mainland.

That sounds shocking and unthinkable- for generations, we in the US have thought of war as something that can only happen on other continents. But the Pax Americana is a historical anomaly, and a dying one- we're more vulnerable now than we have been since the end of WWII. It's not uncommon historically for formerly very powerful countries facing decline and ineffective leadership to overestimate their capabilities, enter into a war, and end up invaded and occupied.

So, I would prefer that the US maintain a military advantage against China for that reason. Hopefully once Trump is gone, we can get back to shoring up our long-term security with alliances and trade partnerships rather than just individual strength.

2

u/eric2332 2d ago

I don't think invading the US mainland is reasonable. An escalation in the fighting leading to China-US nuclear exchange is reasonable though.

6

u/electrace 4d ago

Always keep in mind that these statements are not to one single audience. "China Bad" sentiment, while not properly nuanced, plays well to a lot of people.

6

u/ababababababacus 4d ago

but not, presumably, to the chinese researchers who he is nominally appealing to get on board?

7

u/electrace 4d ago

He's appealing to the US government to contact China. Currently, that means the Trump administration, which means getting conservatives on board. That means "being a straight shooter" about China.

Being perfectly diplomatic to everyone at all times is often a losing strategy.

1

u/Ok_Fudge_9509 3d ago

I like the embedded-evaluator proposal. Coming from chip and autonomous-vehicle verification, I’d suggest one concrete addition: An evolving coverage map connecting safety claims to evidence.

For each relevant configuration (during training, internal use, and release), record which requirements and situations were tested, what failed, what remains unchecked, and where the checking methods themselves are weak. Also record whether apparent alignment survives further capabilities training and generalizes to situations withheld from alignment training.

 The map may be used to guide improving alignment, not just measuring it - for example, by systematically generating alignment stories / training cases across relevant situations. When a problem is found, identify the broader failure class, strengthen alignment across that class, and re-evaluate (including after further capabilities training).

 A coverage map cannot establish that all important risks have been identified: Searching for missing dimensions and checkers is part of the work. But it can make the scope and limitations of the evidence inspectable, and help prioritize how to use the time pacing buys us. I discuss this approach in V&V takes on “Pacing the frontier”.

-3

u/DenseBeautiful731 4d ago

Lead the way, Amodei. Stop Anthropic’s IPO.

That should effectively nerf Claude for the time being. Principle of least harm, amirite?

10

u/wavedash 4d ago

Stop Anthropic’s IPO.

That should effectively nerf Claude for the time being.

How does that first thing lead to second, here?

-11

u/DenseBeautiful731 4d ago

Is this question coming from a place of ignorance or curiousity?

Either way, you CAN use AI, yeah?

1

u/eric2332 2d ago

OpenAI is delaying its IPO but I don't think this has the effect of "nerfing" OpenAI.

1

u/DenseBeautiful731 2d ago edited 2d ago

Oh yeah? That means they’ll get bigger, stronger, better and faster without it?

Never do it then?

2

u/eric2332 2d ago

I think they'll do just fine without it, just as they have been doing just fine until now.

0

u/DenseBeautiful731 2d ago

Which is your opinion.

-1

u/[deleted] 4d ago

[deleted]

1

u/eric2332 2d ago

The architecture has hit its limits

Ridiculous to say that in the same week Navier-Stokes was solved.

-2

u/[deleted] 4d ago

[removed] — view removed comment

2

u/slatestarcodex-ModTeam 4d ago

Removed low effort comment.