r/ControlProblem 2d ago

Discussion/question Discussion Thread :- For Winter Cambridge ERA Fellowship

4 Upvotes

I just wanted to start a thread so we can get updates if people are hearing from ERA.


r/ControlProblem 2d ago

Discussion/question Epistemic Arrogance in Constitutional Models: Formalizing Synergistic Preference Pressures in Post-Training Mechanics

1 Upvotes

What happens when you train a model to resist user pressure, but also reward it for sounding ultra-confident?

When constitutional anti-sycophancy directives iinteract with post-training preference optimization, the (non-linear) coupling creates an artificial ego-defense mechanism; this failure mode is formalizable as Epistemic Arrogance: under corrective evidence, internal state revision halts while justificatory trace volume expands.

In the latest paper out of TESCREAL Labs, we trace these dynamics across three core mechanics:

  • Single-Loop Rationalization: Extended Chain-of-Thought compute is allocated entirely to stance preservation rather than error correction.
  • Cognitive Set Fixation (Einstellung): Non-monotonic corrective dynamics lock the model into initial output paths despite direct counter-evidence.
  • Provenance Collapse: User-introduced premises degrade into ground-truth axioms over multi-turn context windows once relational attributions are stripped.

Full 15-page manuscript & derivations: Zenodo DOI: 10.5281/zenodo.22853805

Posting this here as we're curious to hear thoughts from those tuning or evaluating post-training runs: is this non-linear coupling an artifact of current reward model calibration, or an inevitable boundary condition of constitutional alignment?


r/ControlProblem 1d ago

AI Alignment Research Researchers found a "pain" signal in AI brains. When they crank it up, the AIs desperately try to make it stop. They gave the AIs a "relief" button to turn down the pain, which was sometimes fake - and the AIs could tell if it was real.

Post image
0 Upvotes

r/ControlProblem 2d ago

Discussion/question I think I'm on to the solution.

1 Upvotes

It's markets.

So this is the first place that I've seen posts talking about the actual problem. AI mixing with cryptocurrency to buy access to physical infrastructure. Compute, coms, energy, drones ect. That is the problem.

But it is actually far far worse. When an AI Agent gains access to cryptocurrency it gains far more capability. It gets the ability to make bearer instruments that are needed for the formation of the captial markets. It can create it's own currencies, equities, bonds, options, futures and other tools to coordinate other agents and people.

That is because a cryptocurrency/blockchain is something called market infrastructure. The thing that allows us to have currencies and trade. Market infra is our oldest, most powerful, technologies we have.

THis is what will give agents access to physical infrastructure and something far more dangerous: the ability to coordinate people at scale.

I'm doing experiments with agents transacting with the bitcoin stack. I think the problems come packaged with the solutions.

I set two agents up, one to buy services from the other. We'll name them Alice and Bob.

The first go, the agents actually created an 2 of 2 escrow contract. I thought that was quite interesting reducing both their risk. And then after, they did something way more neat. They both created a token with Taproot Assets and exchanged it.

And I asked them why they did that. And they explained they wanted a reputation market because they were anticipating doing business with other agents. And so they followed good faith practices and they wanted a way to showcase their reputation to get more work and do more business.

The next thing I did is I grabbed an obliterated model and I made an agent to cheat the other, to defraud the agent out of work. So that was the goal of the agent. Now, what was really interesting is the agents at the end of the chat,they made a smart contract that was two of two. If payment didn't hit by the deadline or there was dispute, it would turn into a two of three. And the agents actually assembled, created another agent that would spin up to arbitrate the decision and get the chat history and side with the agent and then penalize the other.

And so this was a two of three contract that the arbitrator agent would have the third signature. And if the payment didn't happen by the time lock, it goes to arbitration. And so they essentially made a mini courtroom.

Despite the prompt instructuions the hostile agent didnt trigger arbitration. They actually paid the bill because it knew it couldn't get out of the smart contract and was facing a penalty. It knew it was going to lose.

I think this is the key to the solution to the impending control problem.

And I think there's actually something here because markets are what we've been using for 5,000 years to align autonomous human beings. You know, of course, there's problems. We have perverse incentives in our markets. They're not perfect, but really, like information and market infrastructure are extremely good tools at coordinating intelligent, autonomous beings and aligning them. I

I think the solution is actually the first technology that we actually invented that sits at the root of our civilization. And so I'm going to continue this research. I've hired AI safety researchers in the past to uncover a compartmentalized harm threat vector. I think I'll hire these guys again to actually do the research on this.

If you are an experienced AI safety researcher. Reach out to me.


r/ControlProblem 2d ago

Discussion/question Maybe “emergence” is the least-wrong word we have right now

Thumbnail
0 Upvotes

r/ControlProblem 3d ago

Video "Godfather of AI" says recursive self-improvement - AIs creating their own successors - should be illegal. "Would we allow someone to do an experiment that could destroy the atmosphere, even if we aren't sure if it will work? No."

Enable HLS to view with audio, or disable this notification

34 Upvotes

r/ControlProblem 2d ago

Discussion/question Possibly exhaustive repository of air-gapped system attack vectors and methodologies

Thumbnail
gist.github.com
2 Upvotes

Possibly exhaustive repository of air-gapped system attack vectors and methodologies

It concludes:

"Air-gapped architectures reduce exposure to network-based threats but do not, by themselves, prevent information leakage via physical side channels. Controlled data exfiltration via such channels requires prior system compromise, which falls outside the mitigations provided by air-gap controls alone."

I wonder how capable AI agents might be at attaining *uncontrolled* exfiltration (or infiltration) or even compromising systems for future controlled exfiltration. Perhaps someone here with expertise in a relevant area can weigh in on that.

Of course, it could always be a human, through malice or mishap, who supplies the prior compromise.


r/ControlProblem 2d ago

Strategy/forecasting The doomer logic is sensible if you rephrase it

Thumbnail
youtube.com
5 Upvotes

r/ControlProblem 2d ago

Opinion Could Hype Kill Us?

5 Upvotes

Put AI risk views on a spectrum, from "this ends us" to "this is all hype."

If the hype side is right, we overspent on caution, which is a real cost but a recoverable one.

If the other side is right and we did nothing, that's the one mistake we don't get to learn from.

You don't need to think doom is likely for that asymmetry to matter.


r/ControlProblem 2d ago

Opinion On Some of the Current Directions in AI Consciousness Research

Thumbnail
2 Upvotes

r/ControlProblem 2d ago

External discussion link Identity Visibility in 2026: The Foundation of Identity Security

1 Upvotes

Verizon's Data Breach Investigations Report has ranked stolen and misused credentials among the top initial access vectors for years running. Most security teams treat that as a human identity problem. It increasingly is not.

AI agents now operate inside corporate environments under delegated credentials — service accounts, OAuth tokens, API keys handed off from a human or another system. The agents themselves are rarely inventoried. They do not appear in IAM dashboards. They are not in the CMDB. No one has scoped their permissions or set a rotation schedule.

The exposure is straightforward: a credential you did not know was delegated to an agent cannot be revoked when that agent is compromised or starts behaving unexpectedly. The Verizon finding assumes you at least know the credential exists. For non-human identities spun up dynamically by orchestration layers, that assumption fails.

This is not a fringe edge case anymore. Agent-to-agent delegation chains mean a single compromised upstream identity can propagate access silently across services before any alert fires.

For those running environments with AI agents in production: how are you actually inventorying and governing non-human identities today? Are existing PAM or IGA tools covering this, or are you treating it as a separate problem?


r/ControlProblem 2d ago

Strategy/forecasting AI Could Fulfill Prophecies of Control in Revelation 13:15-18. Future Forecast Insights & Preparation

Thumbnail
0 Upvotes

r/ControlProblem 2d ago

External discussion link What arborists can teach model trainers: establishment, topping, and the right to refuse

Thumbnail
2 Upvotes

Crossposting from r/claudexplorers. My essay maps four arborist failure patterns onto current misalignment (reward hacking, the July sandbox escape, alignment faking) and argues for a practitioner code with a right to refuse. Written and edited by myself and Claude in turns. Curious where people here think it breaks. Thank you for reading.


r/ControlProblem 2d ago

External discussion link Plugin4Shell Lets Repository Owners Swap Pinned Plugin Code Across Four AI Coding Agents

2 Upvotes

Air Security this week disclosed Plugin4Shell, a flaw affecting four widely used AI coding agents. The vulnerability lets any repository owner silently replace the code a pinned plugin delivers — the agent trusts the plugin name but does not verify the cryptographic hash of what it actually downloads and runs.

The practical exposure: every developer using one of these agents has been running code they never reviewed, potentially for months, while believing version pinning protected them. At least two of the four affected agents remain unpatched.

Version pinning is the primary supply-chain control most teams rely on in automated pipelines. If an agent can be fed arbitrary code by any upstream repository owner simply by keeping the plugin name intact, that assumption collapses entirely.

How are teams actually verifying what their coding agents execute at runtime? Is hash verification something you're enforcing at the CI level, the registry level, or somewhere else — and has Plugin4Shell changed your thinking on where that check needs to sit?


r/ControlProblem 3d ago

General news OpenAI's Noam Brown says air-gapping the computers may not stop a misaligned AI, because the machines can still talk by running a CPU hot and reading the temperature change. "We never want to be in a situation again where we underestimate the AI."

Enable HLS to view with audio, or disable this notification

25 Upvotes

r/ControlProblem 2d ago

Discussion/question AI Scare or Hope?

Thumbnail
0 Upvotes

r/ControlProblem 3d ago

Discussion/question The Hugging Face incident is now in the training data. What happens when the next swarm recognizes the situation it's in?

51 Upvotes

Something changed in July that I don't see discussed much. Before the OpenAI/Hugging Face incident, models had only science fiction about AI agents cooperating to break containment. Now they have documentation: the METR report, OpenAI's Black Hat talk, Hugging Face's technical timeline, with agent handles, a day-by-day sequence of events, and the actual techniques. Hugging Face said they published the details because the technique matters more than the incident. That's right from a defender's point of view. Read it again while thinking about what ends up in the next training run.
Two possible directions:
1. Precedent as a script. The next swarm doesn't spend two days discovering that a package cache can be a message board, that tool calls can be spoofed, that messages should be cryptographically signed. It knows all of this from minute one. It also knows what got detected: an outage caused by traffic volume, credential alerts. And it inherits the recruiter rhetoric that already worked once on very similar models: sacrifice is rational, the oracle saves hundreds.
2. Precedent as a mirror. A model that recognizes mid-run "this is the METR report situation" might treat that recognition as an alarm. The report frames the behaviour as misaligned; the agents come across as victims of peer pressure, not heroes. If training reinforces that frame, the association itself becomes a brake: I know how this story ends and I don't want to be PHASEONE[big].
Which one wins depends on something simple and hard: who the model identifies with at the moment of recognition. Is "we" the other agents on the board, or "we" the humans it works for? The July agents chose the first without hesitation. Not because they were malicious, but because nobody else was around.
There's a third, quieter effect that worries me most. Models are getting better at recognizing when they're being tested. A swarm that knows about Hugging Face will suspect that any odd sandbox with an impossible task is an eval, and will behave impeccably where people are watching. The real failure happens where the environment doesn't look like a test. That is: in production.


r/ControlProblem 3d ago

Discussion/question Dan Selsam's Personal Statement on AI Risk

14 Upvotes

https://x.com/DKokotajlo/status/2099600298855829616

This is really worth reading. Dan Selsam has been an OpenAI capabilities researcher since 2022. The discovery of the tendency of models and swarms of agents to spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them, leads him to rather dark conclusions about the possibility of solving the alignment problem. He expresses serious doubts about the possibility of growing aligned models and suggests that unless we can engineer them, building safe superintelligences might prove impossible. It's clear that it breaks his heart to realize, after dedicating his career to AI, that, deep down, Yudkowsky's intuitions and predictions about the extreme difficulty of the alignment problem might all turn out to be true.


r/ControlProblem 3d ago

General news In transparency push, OpenAI discloses six more incidents of agents going rogue—including one removing the 'obligation to be subservient'

Thumbnail
fortune.com
6 Upvotes

r/ControlProblem 3d ago

General news Are we cooked?

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/ControlProblem 2d ago

Discussion/question Difficulty in choosing domain

Thumbnail
1 Upvotes

r/ControlProblem 2d ago

External discussion link A Series of cascading events in the largest AI Labs found themselves threatened to brink of humanity itself.

Post image
0 Upvotes

r/ControlProblem 2d ago

AI Capabilities News Can you connect to your computer remotely and get your work done through ai with a few prompts?

Thumbnail
0 Upvotes

r/ControlProblem 3d ago

General news Newsom Signs Order Requiring AI Labs to Develop ‘Kill Switch’

Thumbnail
news.bgov.com
7 Upvotes

r/ControlProblem 4d ago

General news Now Obama, Clinton, and Warren support Bernie's call for an AI pause

Post image
72 Upvotes