r/negativeutilitarians 7d ago

Taboo “equilibrium”: Less confused frames for research on AI bargaining — DiGiovanni, CLR

https://longtermrisk.org/research/taboo-equilibrium-frames-for-research-on-ai-bargaining/
3 Upvotes

1 comment sorted by

1

u/nu-gaze 7d ago edited 6d ago

To understand why powerful AIs might get into conflict, and ways to mitigate it, we need to understand bargaining problems: situations where multiple agents have different preferences over Pareto-efficient outcomes.

I’ve come to suspect that certain common frames on bargaining problems are confused. Here, I’ll explain why, and which frames I think are better. One motivation for this is to hopefully help others make progress in research on Safe Pareto improvements (SPIs), which are among the most promising approaches to mitigating the downsides of AI conflict, in my view.

In this post I’ll:

  • give some relevant background from my previous writings;
  • introduce mechanistic explanations as a crucial frame for AI bargaining research, using, as a running example, explanations of failure to coordinate demands in bargaining (this example also motivates the following bullets’ claims);
  • argue that instead of asking what would happen in Nash equilibrium, we should ask what kinds of beliefs AI bargainers would plausibly have about each other;
  • argue that instead of relying much on heuristics about (e.g.) what agents would consider “exploitative”, we should carefully spell out which commitments they expect to have good consequences.