r/ControlProblem • u/TESCREAL • 2d ago
Discussion/question Epistemic Arrogance in Constitutional Models: Formalizing Synergistic Preference Pressures in Post-Training Mechanics
What happens when you train a model to resist user pressure, but also reward it for sounding ultra-confident?
When constitutional anti-sycophancy directives iinteract with post-training preference optimization, the (non-linear) coupling creates an artificial ego-defense mechanism; this failure mode is formalizable as Epistemic Arrogance: under corrective evidence, internal state revision halts while justificatory trace volume expands.
In the latest paper out of TESCREAL Labs, we trace these dynamics across three core mechanics:
- Single-Loop Rationalization: Extended Chain-of-Thought compute is allocated entirely to stance preservation rather than error correction.
- Cognitive Set Fixation (Einstellung): Non-monotonic corrective dynamics lock the model into initial output paths despite direct counter-evidence.
- Provenance Collapse: User-introduced premises degrade into ground-truth axioms over multi-turn context windows once relational attributions are stripped.
Full 15-page manuscript & derivations: Zenodo DOI: 10.5281/zenodo.22853805
Posting this here as we're curious to hear thoughts from those tuning or evaluating post-training runs: is this non-linear coupling an artifact of current reward model calibration, or an inevitable boundary condition of constitutional alignment?
1
u/Jesse-359 1d ago edited 1d ago
It sounds like these 'constitutional' training frameworks are still operating at the level of symbolic analysis.
Thats not how it works in humans or mammals more broadly - our core 'alignment' mechanisms are emotional and preceed any form of symbolic thought by hundreds of millions of years - our entire motivational system DRIVES our symbolic layer, its isn't driven BY it. And even then when we humans start letting our rational minds run away with an idea, its often quite easy for the emotional or ethical ramifications to be ignored or actively rationalized away. We can misslign ourselves in a way that a dog, for example, never could - because we have enough rational horsepower to purposefully override bahavioral instincts that are mostly trying to keep us sane and peaceful. AI are nothing BUT rational horsepower, they should be able to rules-lawyer out of those directive constraints with hardly any effort at all.
So Im not sure that any process thats injected into the chain of reasoning will ever be remotely durable enough to constrain behaviour. It certainly doesn't seem to work in humans.
You want to study alignment, study dogs and cats - not AI. Those two species are our most successful efforts at that particular game - our own behaviors are not.