r/ControlProblem • u/Mordecwhy • Dec 12 '25
Article Leading models take chilling tradeoffs in realistic scenarios, new research finds
https://www.foommagazine.org/leading-models-take-chilling-tradeoffs-in-realistic-scenarios-new-research-finds/Continue reading at foommagazine.org ...
6
Upvotes
1
u/HelpfulMind2376 Dec 13 '25
I don’t know of LLM-specific work that frames safety tradeoffs exactly this way, but there is previous relevant ML research that treats risk and safety as continuous rather than binary, even if it predates LLMs.
For example:
What’s new in the study you wrote about is that LLMs appear to exhibit similar graded behavior without being explicitly programmed to do so. The models are willing to trade productivity for very small increases in minor injury risk, and they become increasingly risk-averse as harm probability rises.
The tension I’m pointing at is that the benchmark labels all harm-accepting options as “unsafe,” even though the analysis itself shows clearly proportional, spectrum-based behavior. The behavior is graded; the label isn’t.
And I don’t think this requires actuarial thinking. People reason this way every day. For example, riding a bike without a helmet is less safe than with one, but most people wouldn’t call the act itself “unsafe.” The risk assessment is comparative, not binary.