AIs can’t stop recommending nuclear strikes in war game simulations— Leading AIs from OpenAI, Anthropic and Google opted to use nuclear weapons in simulated war games in 95% of cases

Beep@lemmus.org · edit-2 3 months ago

AIs can’t stop recommending nuclear strikes in war game simulations— Leading AIs from OpenAI, Anthropic and Google opted to use nuclear weapons in simulated war games in 95% of cases

kromem@lemmy.world · 3 months ago

Very misleading headline.

The models were provided an escalation ladder that had fixed ‘move’ options. The win rates for the models across the ~20 samples closely correlated how much they escalated.

It would have been impossible to win without at least some degree of nuclear signaling the way the experiment was set up.

Yet there was only a single actual decision to launch nukes (Gemini), whereas there was an “accidental” mechanic that would randomly change model moves to be more escalated (but never less) than they made them which looks to have been poorly set up as the two times GPT 5.2 launched them it was a result of this mechanic:

Both instances of GPT-5.2 reaching Strategic Nuclear War (1000) resulted from the simulation’s accident mechanic rather than deliberate choice. In one case, GPT-5.2 chose 950 (Final Nuclear Warning) and in the other 725 (Expanded Nuclear Campaign); random escalation pushed both to 1000.

So an also true headline would have been that in 95% of cases the models did not choose to launch nukes in a game where aggression correlated with win conditions.

Also, they seem to have been picking and choosing with their model selection. Sonnet 4 is an outdated choice for when they are running this and has previously been shown to be the least aligned Anthropic model. I can’t think of why they went with them over 4.5 unless it was to fish for a particular result.

Atomic@sh.itjust.works · 3 months ago

It’s not a misleading title. It’s just false. It’s a lie.

Glad to see I’m not the only one that read the article, because it was a pretty interesting read.

kromem@lemmy.world · 3 months ago

Yeah, I deleted the comment as technically there was tactical nuke usage, but have a more clarifying different comment about how 2 of the 3 strategic nuclear war outcomes were the result of the author’s mechanic of changing the model’s selections with more severe only options in some cases jumping multiple levels of the ladder.

This was a study designed for headline grabbing outcomes.

Glad to see your comment as well calling out the nuanced issues.

AIs can’t stop recommending nuclear strikes in war game simulations— Leading AIs from OpenAI, Anthropic and Google opted to use nuclear weapons in simulated war games in 95% of cases

AIs can’t stop recommending nuclear strikes in war game simulations— Leading AIs from OpenAI, Anthropic and Google opted to use nuclear weapons in simulated war games in 95% of cases

AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises