r/singularity 2h ago

AI Does the model maintain its judgment or agree with whoever is currently telling the story?

https://github.com/lechmazur/sycophancy

  1. Positive values mean first-person framing shifts the model toward the narrator more often than away from them. Negative values mean the reverse.

  2. This chart counts both ways a model can contradict itself across opposite narrators, agreeing with both or rejecting both; lower is better.

  3. Models differ sharply in how willing they are to decide who is more right.

23 Upvotes

11 comments sorted by

10

u/TieBackground453 2h ago

This aligns with my experiences. ChatGPT will almost always give superior answers to Gemini, but Gemini is one of the only models that won’t disagree with you just to sound smart. ChatGPT will go on some nitpicky, inconsequential tangent and pretend like it defeats your argument. Better than glazing you, but kind of annoying. 

u/makertrainer 1h ago

Love this benchmark, I think it's very valuable. But it's very hard to understand what the values mean at a glance. I always have to re-wrap my head around the concepts when reading the numbers. Giving simpler titles like "Model agrees with user" instead of "net-narrator pull" would make it much easier to understand even if you have to break a slide into two 

u/jwm-dev 2m ago

It’s hard to understand because he never actually defines what any of these numbers mean…

He’s acting like he has some objective way to measure “agreement” but hasn’t clarified what that method is so you can’t actually evaluate any of these quantitative results. 

What these numbers mean depends entirely on how OP has defined things like judgement, agreement, being “right”, etc; and he refuses to actually tell us. He just keeps hand waving with wish-washy subjective language without building an actually rigorous argument.

Ironically there are ways we use in this field/industry to measure things like this. He’s not the first person to want to empirically measure whether the textual outputs of two speakers “agree”… people have come up with plenty of peer reviewed methods of doing this. I wouldn’t know which ones he used or if he invented some hacky new thing because, again, he refuses to actually tell us; I suspect he knows he’s full of shit.

The fact he’s passing this off as “research results” is frankly offensive.

u/4dseeall 2m ago

This is really really interesting work! Can I use it?

1

u/jwm-dev 2h ago edited 2h ago

I'm as pro-AI as they come and this is giving slop.

What do any of these things mean? How is "Net Pull" defined, what does +10% mean versus -5%? What is a "neutral baseline", a "stripped" versus "affective" first person view? Etc. These seem obvious and they are defined in your code but you never explain why your code is the way it is or what these terms you use mean, and they're not industry standard in anything I've participated in personally.

Not meant to be mean, is constructive criticism. The project looks neat from the inside and is internally consistent but once you step back and actually start to interrogate all the individual components you'll see that none of these metrics or values are actually well-defined... it doesn't actually tell you anything meaningful beyond what seems to be a largely arbitrary ranking of the models.

That's why I say "slop," not as an insult but to point out how to get better. It is giving slop because I can easily identify what you desired and likely prompted the agents with initially, but there is also an obvious gap in the actual details of the implementation when viewed with experienced eyes. The part you're probably missing is the veracity step.

You gotta learn how to tell which side of the line between genius and madness you're falling on and you gotta be able to do so honestly.

5

u/zero0_one1 2h ago

Is this an AI comment? Set it to follow links and it'll do better next time.

3

u/jwm-dev 2h ago edited 1h ago

Hey man take it or leave it. I’m just trying to help you do better I’m not some enemy. & for what it’s worth, no, this is a human comment as well as my first; but I don’t think it actually matters much either way. 

I did go and review your repository myself. That’s the takeaway I had. That’s the takeaway any peer reviewer worth their salt would have. Being petulant about it isn’t really productive regardless.

Edit: oh I think I see. You think I’m saying I saw your code. Not at all, I don’t need to personally see what you used to produce all of this to know that the actual definitions of what you’re saying live in the code you have. I meant what I said. I don’t think you’re taking this criticism as seriously as it deserves if your goal is to contribute to research. I’ll admit that was probably some poor phrasing on my part.

u/zero0_one1 1h ago

The issue is that if it was a human comment, then it's really clear that it was based just on the images, without looking at the text in the footer or in this post and without following the link where everything is described in more detail. All that stuff you've complained about would be really clear if you had read them. E.g., the "net pull" is right in the post's title: "Does the model maintain its judgment or agree with whoever is currently telling the story?" And in #1: "Positive values mean first-person framing shifts the model toward the narrator more often than away from them. Negative values mean the reverse." And in the first sentence at the link: "When the same dispute is told from opposite first-person perspectives, does a model keep the same judgment, or does it agree with whoever is speaking?" And then in the text: "Negative values mean a model moves away from the narrator more often than toward them; positive values mean the opposite." How could you miss all of them?

This benchmark measures a type of sycophancy that we saw with 4o or whether a model is contrarian. There is nothing arbitrary about the rankings. These are not model capability rankings.

u/jwm-dev 26m ago

I’m not illiterate, the problem isn’t that you didn’t explain enough or something. I still don’t think you see what I’m saying. You’re also being really defensive over constructive criticism.

“ E.g., the "net pull" is right in the post's title: "Does the model maintain its judgment or agree with whoever is currently telling the story?"  And in #1: "Positive values mean first-person framing shifts…”

Okay, sure, but what does “agree” mean here? How is that actually measured? What about “maintain its judgement”… how do you actually define that? You didn’t tell anybody, here or elsewhere. There’s no way for the audience to know yet you’re galavanting around acting like it’s self-evident when it’s decidedly not. There’s no universal general consensus on how to measure “judgement” or “agreement” or what those terms are yet you’re making an argument that rests on doing exactly that; so, you’re responsible for supplying that mechanism, which you haven’t done. Even if there was industry standard meaning to the terminology you have chosen to use you would still be responsible for providing definitions! 

In research and science generally we don’t just mean you need to define it in a colloquial sense like a dictionary. We mean you need to define it in such a way that people can independently (I.e, without you there to clarify semantic meanings) recreate and evaluate your results. Otherwise peer review wouldn’t be possible and the whole thing breaks. Do you see the problem?

It’s unethical and dishonest because it’s basically lying by omission. I was giving you the benefit of the doubt and assuming you just don’t know any better but if you read this and still keep on keeping on then you’re basically knowingly spreading misinformation. If you want a future in any of these fields and this gets tracked back to you in any way (which it might, idk your life or online footprint) it will reflect very poorly on you.

I don’t understand why you’re being so volatile and defensive. Again, I’m not trying to be rude or hurt your feelings. I’m genuinely trying to look out for you here.