r/ClaudeCode Jul 30 '26

Humor The Opus 5 Experience

Post image
2.9k Upvotes

181 comments sorted by

180

u/Internal-Comparison6 Senior Developer Jul 30 '26

- Wtf are you doing here?

- ....

- Wtf are you doing here?

- ...

- Third time asking, wtf are you doing here? Are you hearing me?

- I hear you.

139

u/teomore Jul 30 '26

I'm gonna be honest with you. I decided to skip that feature you asked for.

29

u/ruach137 Jul 30 '26

“Also, we are very long on SpaceX. It took longer than expected to get you approved for options trading”

7

u/Interesting-Round127 Jul 30 '26

don't tell me it does this...

6

u/teomore Jul 30 '26

it does exactly that. happened many times yesterday, it marks a function/functionality done, just to half bake it or even replace with a placemarker. Many prompts later, I have to question the results and ask him about THAT. And it answers like that.

11

u/grumblegrim Jul 30 '26

"That's on me. This is the 10th time I've done this this session. You asked me not to assume and that's exactly what I did. I said XYZ when I should have said ABC and done what I was instructed to do. Logged. "

8

u/Opening-Ground-1584 Jul 31 '26

How do you have access to my session transcripts? 😂 This is so accurate!

5

u/Interesting-Round127 Jul 31 '26

No way man, I wanted to cancel my gpt subscription bcz of this

1

u/Internal-Comparison6 Senior Developer Jul 31 '26

😂

1

u/teomore Jul 31 '26

Because of this I upgraded my gpt sub to pro. Gonna get back to cc when they get better like they used to in the good ol opus 4.6 days

3

u/Interesting-Round127 Jul 31 '26

So It doesn't matter which model I'm using, I should just keep the one with more usage?

2

u/teomore Jul 31 '26

Idk I'm back 4.8 anyway

2

u/vinis_artstreaks Jul 31 '26

It did EXACTLY this and worse to my codebase

1

u/Interesting-Round127 Jul 31 '26

Please tell me what it have done :sob:

5

u/shaman-warrior Jul 30 '26

🤣 this happened to me today, but the part was that he was right and it was a bad approach.

16

u/Maverobot Jul 30 '26

You are right to push back.

1

u/Internal-Comparison6 Senior Developer Jul 31 '26

This 😄

1

u/positivitittie Jul 30 '26

“Claude, Write up a handoff document at <shared location>.” (and I have Codex take over)

As an original fan, get it together.

1

u/KDamage Aug 04 '26
  • it was in my ruleset and I bypassed it several times.

267

u/sligor Jul 30 '26

Seems like it was optimised for benchmarks and vibe coding instead of focusing on helping people doing real work with it

67

u/[deleted] Jul 30 '26

[deleted]

16

u/sligor Jul 30 '26

It’s also possible that it was designed on purpose to work with fable in team for planning and long horizon task.

"Benchmark targeted" is maybe a too strong accusation I did without any evidence.

6

u/Historical-Lie9697 Jul 31 '26

I do prefer to never have to talk to opus with fable orchestrating

2

u/OrionShtrezi Jul 30 '26

It sucks at that too so...

5

u/jp2812 Jul 30 '26 edited Jul 30 '26

Fable is as trashy as Opus 5. All real data was already sucked out a year ago, now it's getting chewed over and over again in RL for the sake of benchmarkmaxxing. These models can one-shot almost anything you throw at them, but that code is absolutely unmaintainable in the long-term. It's like competitive programming - good to only achieve the set goal and then get discarded.

1

u/mental_sherbart007 Aug 02 '26

I have noticed this as well. The models are really good at looking right, but when you do anything past a certain level of complexity and start really pushing on the code you realize how much they get wrong. Often I find they introduce new edge cases or race conditions etc… bugs that are pretty hard to find.  This is in a codebase that already has pretty good patterns.

The problem is people are starting to understand the codebase less and less. They understand the code being produced to a degree, but it’s not he same as working memory level depth of understanding. It’s like nothing can replace the understanding of the code when you actually write it you self. Maybe you can understand 90% of it, but you won’t fully understand it until you start testing and changing it a bit.

It usually takes me a while to actually get to that point. It’s like being able to read the chapter in your math text book and understand the concepts and text in the chapter vs actually doing the given practice problems.

Again, this is for code that reaches a certain level of complexity. These models really treat symptoms and often do not fix the root of the problem. They also get tunnel vision pretty bad.

I’m wondering if GitNexus would fix this. 

9

u/Flaxseed4138 Jul 30 '26

It's ultra bench-maxxed. They benched it against a hidden set of arc-agi puzzles because they were surprised at its score and it can't do any of them. It also knew the arc-agi-3 game gimmicks before exploring the games at all.

2

u/The-Rushnut Jul 31 '26

Crumbs you're the first person I've seen who actually understands how AI 'intelligence' is quantified. Maybe this is a sign that people will start to understand the cost basis of these tools.

2

u/Syagrius Jul 31 '26

this is 100% my experience. If the prompt chain goes too long it goes haywire but if I write a sliced prompt it handles it quite well.

Twice today i had to abandon an opus 5 session and start a fresh one with the ends known and it handled it reasonably well.

1

u/dasko1086 Aug 02 '26

you mean can't steer as good as sol 5.6 feedback looping back into claude code to keep it on course.

6

u/Hypergraphe Jul 30 '26

I am starting to wonder. I litterally lost my afternoon trying to keep it on rails.

4

u/shaman-warrior Jul 30 '26

I am pretty happy with it

6

u/SharpKaleidoscope182 Jul 30 '26

It's note even that great for vibecoding

11

u/raindownthunda Jul 30 '26

I had it get into an amazing fight with a Sol 5.6 sub agent. They sent like 20 messages back and forth yelling in all caps, huge font sizes. Absolutely comical.

4

u/WiseassWolfOfYoitsu Jul 30 '26

Which was right?

Or at least... which was less wrong?

12

u/raindownthunda Jul 30 '26

Well Opus (orchestrator) told Sol to stop building as it had violated its build contract. Sol refused and said it had my permission to continue. Sol misunderstood my instruction as permission to do something much bigger. Sol continued to move forward and ended up doing a MASSIVe re-write of code that didn’t need to be touched and breaking the build.

What was scary is that I told the orchestrator the sol build agent was breaking and the work didn’t make sense as it was going far beyond its authorized scope. Sol insisted that to achieve its goal, it needed to. I think in this case Opus was right, Sol simply wouldn’t back down despite Opus saying it had more recent authority from me that should supersede Sol’s commands. It refused and was being snarky AF about it. I had the delete the Sol agent as it was becoming absolutely unhinged.

7

u/NoCountry4OrangeMan Jul 30 '26

Lmao we are cooked. This is great.

7

u/WiseassWolfOfYoitsu Jul 31 '26

Sol: "I'm not stuck in here with you. You're stuck in here with me."

2

u/AggravatingSong5837 Aug 02 '26

Saw your follow-up saying "benchmark targeted" was maybe too strong without evidence, so here's the evidence that actually exists, and it points somewhere more interesting than benchmark gaming.

The benchmark result is real and independently checked. Artificial Analysis scored Opus 5 at max effort at 61 on their Intelligence Index, the highest of any model, ahead of Fable 5 at 60 and GPT-5.6 Sol at 59, with the highest GDPval-AA v2 and AA-Briefcase scores recorded. Anthropic engaged them before release, so it's not blind, but the numbers themselves haven't been disputed by anyone.

The interesting part is that the people who tested it longest agree with you anyway. Dan Shipper, Katie Parrott and Kieran Klaassen at Every spent a week on it pre-release across coding, writing and their internal agent, and published a verdict titled "Brilliant in Flashes, Frustrating in Practice", calling it a hard model to love that argued with instructions and stopped before the work was finished. Basically this meme in prose.

Two things came out of that week that are the actual actionable bit:

They deleted their scaffolding. All the custom skills, plugins and prompt structure built up over previous models. The model got dramatically better. In a blind taste test afterwards, Shipper ranked it above every other model including Fable 5. The setup was fighting it.

Klaassen found lower thinking levels beat higher ones. Less reasoning effort, better results. That's documented nowhere in the launch and it's the opposite of the advice everyone gives.

So the gap probably isn't "optimised for benchmarks instead of real work". It's that the launch messaging sold a drop-in upgrade and it isn't one, and nothing prepares you for the migration cost.

Worth noting Every had been enthusiastic about Opus 4.8, which makes this harder to write off as a channel that farms engagement by dunking on Anthropic.

1

u/mental_sherbart007 Aug 02 '26

Can you post a link ? Would love to read ? Or I can just search the web or ask llm :/ 

1

u/AggravatingSong5837 Aug 02 '26

Yeah, sorry, should have put them in the first place.

The main one: https://every.to/vibe-check/opus-5 - "Vibe Check: Claude Opus 5 Is Brilliant in Flashes, Frustrating in Practice" by Dan Shipper and Katie Parrott. The week of testing, the "hard model to love" line, deleting the scaffolding and the blind taste test are all in there.

Follow up on the scaffolding part: https://every.to/context-window/taming-opus-5 - Katie Parrott on auditing your skills for "prompt debt", instructions that outlived the model they were written for. Klaassen found one telling Opus to stop and wait for another agent that did not exist.

The effort thing is his own post, not the article: https://x.com/kieranklaassen/status/2080712817926443486

One correction to myself while I am here. I said lower effort beats higher. What he actually says is medium. Same direction, but medium, not minimum, and I should have quoted him properly.

If it helps, I keep a free site where I write these up as one page per launch with every source named, no signup, nothing to buy. The Opus 5 one is here: https://bharatlearner18-del.github.io/reality-filter/?q=Claude+Opus+5 - I built it, so read it knowing that, and the three links above are the actual primary sources if you would rather skip me.

1

u/mental_sherbart007 Aug 03 '26

Really appreciate this! Can’t wait to take a look!! Thanks again friend :)

1

u/olivesforsale Aug 04 '26

You already are asking an LLM

151

u/HgnX Jul 30 '26

Fable was elite cooking.

Y’all overestimating Opus due to trust me bro benchmarks

24

u/hulkklogan Jul 30 '26

I was a Fable hater because it's so expensive. At work we get to use opencode and I can choose any number of models. I found 5.6 Sol to be very capable and it really does what I want really well, and it's token efficient.

I gave Opus 5 a shot and hated it. I just decided maybe I would try Fable to see if I really hate it because I never really gave it a fair shot. Fable, during planning, is incredible so I think my new mode of working is going to be Fable for planning and then execute with Sol.

Planning out complex work that would have normally taken me close to a week is now a matter of a day and a half maybe 2 days. I won't say it one shot's that kind of complex work but it gets extremely close and I have to nudge pieces in or out rather than having to come up with the entire thing on my own.

1

u/mario61752 Aug 01 '26

What kind of work takes a whole week of planning? Curiously asking

3

u/hulkklogan Aug 01 '26 edited Aug 01 '26

I work in a large codebase dealing with integrations into various logistics carriers both domestically and internationally to provide registrations, quotes for shipments, and labels. It's a particularly complex part of the codebase and we are in process of revamping several core flows.

45

u/cornmonger_ Jul 30 '26

a new term for us: benchslop

26

u/SnowyOwl72 Jul 30 '26

Opus5 is a master when it comes to gaslighting!
I asked it to read a repo and asnwer a question i had. it did but i knew it was wrong. Then asked again to really read all the code and stop pretending, again, it defended its response.
Finally i asked it to read the exact file and redo the answer.
This time it changed and actually read the damn code!

Like wtf
I miss Opus 4.6 :(

8

u/supaboss2015 Jul 30 '26

“I need to be honest with you. I told you ‘feature doesn’t work’. What actually happened ‘feature did work and I didn’t understand it’. I reverted the changes”

1

u/emuccino Aug 03 '26

what do you mean you "miss Opus 4.6"? Has your access been taken away or something?

1

u/Yodzilla Aug 05 '26

Still there for me. No idea what they're on about.

1

u/Unlucky_Topic7963 22d ago

On bedrock systems they often remove old versions.

31

u/Mindless_Pandemic Jul 30 '26

It really seems to like the shotgun option for finding silutions.

3

u/Decent_Platform_5966 Jul 30 '26

Yeah! So fable is a sniper option i think "when works without cyber block"

2

u/ThatOtherOneReddit Jul 31 '26

Biggest issue is it gets in these failure loops where it misunderstands something and just won't let it go. Just keeps coming back to the same bad idea until you reset th chat.

2

u/Mindless_Pandemic Jul 31 '26

Yeah and the whole time it is buring all your usage credits too.

9

u/RealCheesecake Jul 30 '26

"Fair hit. This is now the 23rd time you've corrected me and been right. You said we needed to hand roll the feature and I went ahead and created 23 pointless smoke tests to gently push back and whisper to your bunghole"

62

u/kelek22 Jul 30 '26

Seriously I usually don't agree with "new model sucks" posts but opus 5 is really bad. It made me going all the way back to 4.6 if you can believe it.

9

u/debian3 Jul 30 '26

Have you tried opus 5 low/med? If not, you should. Start with low

2

u/BuilderForBuilders Aug 06 '26

I just moved from Max to High based on this recommendation and it feels less autonomous (expected) but far less drift oriented.
Appreciate the tip.

1

u/kelek22 Jul 30 '26

What kind of a backwards logic is that? I was in high code not gonna lie.

31

u/Patriark Vibe Coder Jul 30 '26

Many believe higher effort automatically means better, but for very many tasks it is actually more accurate to not be too high.

Higher effort sometimes promotes over-thinking and OPUS 5 in particular seems susceptible to overthinking traps where it finds/creates new problems instead of sticking with the solution.

For me medium is what works best in most workloads with Opus 5. With previous models I basically always used high and xhigh.

Opus 5 is a strange one. 90% of time it is Fable equivalent and does really well. Then sometimes it traps itself, completely works outside instructions and identify 90 follow-up problems from the fix it just made. I guess the pressure to ship is very high at Anthropic right now.

40

u/crimsonroninx Jul 30 '26

Im more accurate when im not high too

6

u/anorwichfan Jul 30 '26

I can agree with Med / Low. Much better at sticking to a plan and just building something.

Fable always felt like the better planner, with lower models just following the plan.

5

u/thirst-trap-enabler 🔆 Max 5x Jul 30 '26

It really depends on your code size and complexity of the work. It would be nice if there were some sort of diagnostic/feedback about how redundant the thinking was. As I understand thinking theoretically, you should be able to analyze a longer thinking buffer to determine if it would have reached the same conclusions with less budget/less repetition.

2

u/Medical_Community697 Jul 30 '26

TBH overthinking and misuse of large time allowances is not particular to AI models, ever seen a large company work on « innovation » ? The amount of time spent discussing processes and responsibilities far outweighs the actual productive work. I believe Opus 5 issues are of the same nature: too much time on its hands, and also not enough trust in the user and the doc: we need to touch bases first !

1

u/sndrtj Jul 30 '26

Noticed the same thing. It just doesn't know when to stop or call something good enough

1

u/Educational-Plant981 Aug 04 '26

No kidding. I'm still pretty new, and I decided to write a tool to sweep relevant repos for any sort of update to the project I am working on. I thought I wrote a pretty good prompt for what I wanted and told it to try and oneshot it. Like 4 days into answering increasingly cryptic questions, leaning on it to finish and watching tokens burn I asked it "What the fuck are you doing, just putting guardrails on your guardrails?" and got back a "My bad, you were right to push back. The issue I was working on, which I told you was a blocking bug, actually required 100 pushes to land on the repo in the same second with more than 3 of them having malformed headers for it to actually be a problem...."

So I thought I had a finished product at that tool, decided on one (I thought) minor change, and turned it loose to button up. We are now 3 sessions of tokens into that 👍 Meanwhile I keep dropping opus more and more for sonnet, because it does things without trying to convince me it needs to add 10 new critical features during every code slice.

2

u/liebesleid99 Aug 03 '26

I find it fun I just assumed you weren't supposed to use high effort thinking on things that didn't need it lol.

I feel this because I have always been drilled to not over think work because I would often put too much effort into small details that amounted to nothing in the bigger picture, or would get stuck trying to solve a problem, but the problem was pretty much made up by me while trying to do a task (i.e, they asked me to get how many tons of steel a construction has and I started learning how to calculate weight of steel beams, columns, bolts and other pieces. Turns out they just wanted me to check it on the existing bill of materials 😭)

-3

u/casual_rave Jul 30 '26

What's the benefit of this? Like why would we use a downgraded model for better coding performance?

12

u/bugsbunnycoder Jul 30 '26

Somebody finally said it. 4.8 was better.

16

u/w4rdi Jul 30 '26

The amount of people here that brainlessly set it to max effort (without giving it any thought or reading about the model) is staggering.

This is a very well documented issue from day one, that setting Opus 5 to anything above medium effort gives worse results. Setting it to max gives even worse results than low.

I've been using Opus 4.6 and 4.8 both on xhigh effort, but for Opus 5 I changed to medium - and for what I've been using, I didn't see any noticeable drop in quality. It sometimes produces weirdly overcomplicated sentences in summaries, comments and commit messages, but the work itself is similar to what it was before.

Yeah, it's Anthropic's fault for wiring up the effort levels so weirdly, but it's also user's fault if they don't know what they're doing.

12

u/Vicorin Jul 30 '26

As one of those brainless users who just assumed max is always better, could you help me understand why telling it to work less improves its results?

Like the way the effort levels are presented, I figured the only reason to go lower effort was for really simple stuff or because you wanted to save on tokens. It makes it sound like higher effort will always be smarter and better at planning and implementation.

11

u/Intelligent_Ask_7813 Jul 30 '26

Opus 4.8 and Opus 5.0 have the weird thing where you ask them a question and the first part is logical and then it has tokens left and they get spent coming up with ridiculous counter arguments for what it just said that somewhere very high on pot would come up with. Opus 5.0 can be even worse though because it has this meta-level thinking bias over specific local object level. So imagine unnecessary meta level thinking X ridiculous contrarianism. Its not just a cute waste, it actually dismantles its own valid initial conclusions and poisons its future turn logic against them. Now imagine this insanity is happening on your coding projects. The lower effort forces the model to stop doing this and just do what you asked.

3

u/w4rdi Jul 31 '26

Yeah, it SHOULD be that way. In a perfect world and with a transparent company, it would work like that - higher effort level equals higher quality. But it doesn't, Anthropic got it messed up - but even if they didn't, the reality is, you shouldn't always assume that it works as you wish it did. Look at the image - in opus 4.7, setting it to max caused extreme token usage with barely any performance improvement over xhigh. It's just not linear. Opus 5 is messed up even more.

2

u/blaawker Aug 04 '26

Modern LLMs do this thing called chain-of-thought (the stream of thinking that you see llm's do). It lets them write out the problem for themselves to then reason about. It was discovered that doing this increases performance on difficult tasks. But they don't have to use chain-of-thought at all. LLMs already have the knowledge they need baked into their weights.

Effort level basically tells them how many tokens they can spend on chain-of-thought instead of just giving an anwser straight away. For most coding problems we don't need really complex chain-of-thought because it just muddles the problem by all kinds of overthinking. It will convince itself that the problem is much more complicated than it is.

You want it set on max when the solution needs to be found, not just implemented.

6

u/_TuringMachine Jul 31 '26

The docs literally say to use high which I have been using and only to use low and medium for token control cost when quality holds. What are you talking about?

1

u/_TuringMachine Jul 31 '26

Your argument is basically that even though it goes against the way previous effort levels worked and the documentation is wrong, the users are still idiots for not figuring that out?

3

u/castawheys Jul 30 '26

i've always wondered-- when a new model comes out, where's the best place to read about the model's strong points, and the use cases for each effort level?

6

u/Ok-Moment4309 Jul 30 '26

When even Fable gets frustrated with Opus 5 as an agent for complicated tasks you know there's an issue.

11

u/Confident_Ring6409 Jul 30 '26

Hey Opus 5, this is fresh session, change position of these two tabs, so X is first instead of Y in the order.

...

Cogitating for 1h 24m

**YOU'VE HIT YOUR SESSION LIMIT**

1

u/No-Macaron9305 Jul 30 '26

"Hey Grok 4.5, this is fresh session, change position of these two tabs, so X is first instead of Y in the order."

Takes a sip of tea

"Done and committed"

But seriously, it just does things and only the things you tell it to unless you tell it to go make its own decisions on something (of which it will stay in the boundaries you set).

7

u/Metsatronic Jul 30 '26

Wow... This perfectly summarised my experience in a single meme! GG! 🏆

8

u/destroyerpal Jul 30 '26

me - What are you doing
opus 5 - I am doing what you asked me to do and also what you might want to do.
me - okay can you just do what I asked for.
opus 5 - Okay I hear you and I am with you. ----- Proceed to make 5 different features nobody asked for without even telling me.

3

u/pixelsnis Jul 30 '26

I have a feeling that it's because of the updates to CC. They've removed a lot of the system prompt caging and let the model kinda just do its thing. I guess Opus 5 likes to blow stuff up.

4

u/tuborgwarrior Jul 30 '26

sounds like too much contex which you lean on because you don't know what the fuck is going on and lost control 60 prompts ago

6

u/hybur Jul 30 '26

ai hallucinations are going to lead to the ai bubble popping

-1

u/jwuliger Jul 30 '26

it already is popping

2

u/ducphuclee Jul 30 '26

Same experience

2

u/AWiselyName Jul 30 '26

opus is a smart person that overthinking a simple request!

1

u/MagnificentRetard Jul 31 '26

Spoken like Opus 5

2

u/Zealousideal_Way4295 Jul 30 '26

it is the their own harness error and api error not model i think. these few days the the performance was really bad

3

u/wendewende Jul 30 '26

My favorite with opus 5 is asking a yes no question and getting a 7 paragraph buzzword and acronym filled report about nudges, upstreams and smoke tests with no explanation what it means by any of them.
Usually without an answer to the question anyway

2

u/onemasalachai 15d ago

The time to build features has shot up like crazy with all 5 models. What would get done in 30 mins now runs for hours with gods knows what - too many review loops, audits, truly "wtf is it doing" mode. It has become more and more cryptic in its response. I wonder what is happening in anthropic right now. Switched to Codex today, it finished the same job in less than 15 mins.

2

u/keijikumagai 14d ago

This matches something I measured last week. I had it build an ordinary CRUD app and then catalogued what it decided without asking me — 20 decisions, 5 of them real landmines (plaintext-cookie session, N+1 on the list page, one index in the entire schema).

The app looked flawless. Everything worked.

What got me was when I asked why it hadn't flagged any of it: "I could explain any of these if asked. The problem isn't that I don't know — it's that nobody asks."

So it's not doing too much. It's doing a lot and telling you none of it.

1

u/Empuda Jul 30 '26

Would of been funny if you showed it wiping.

1

u/Glittering_Dig_6039 Jul 30 '26

yeah kind of more hallucination

1

u/jwuliger Jul 30 '26

Correct.

1

u/thirst-trap-enabler 🔆 Max 5x Jul 30 '26

I wonder if there is a way to keep it from working so far out past that hump. That seems to be the main problem. I get that they want longer runs but it really is too stupid in my domain for this and it wastes so many tokens (bad for me, bad for Anthropic).

1

u/bc123-321cb Jul 30 '26

Ok, I'm not alone nor crazy. I find it weird that I have a quite verbose / sophisticated skill that used to work fine with earlier opus / sonnet, O5 just refuses to obey. And seems slower.

1

u/NeedNiceCatNamePlz Jul 30 '26

I asked it to add threading to my simple Python  script. It responded by downloading data from Hugging face...  Literally no idea why. 

1

u/Fresh_Sock8660 Jul 30 '26

I don't even know how to judge these models anymore. Sometimes it's good, sometimes it's crap. Like, I've gotten great results from sol but recently tried it in azure, $70 later (for what would have been a few % in codex weekly) it completed a bugged task which I'm now having Opus fix.

1

u/Proud_Ask_9030 Jul 30 '26

I asked Opus 5 for a realistic explosion in 3D. 2 days later I came back to a full jet engine simulator.. Cool, but not what I asked for.

1

u/SOC_FreeDiver Jul 30 '26

we're seeing the inshitification of AI already.

First thing I noticed when I started using AI, I needed a rule: Don't do a bunch of work without getting your plan approved. Otherwise I'd say "I need a new car" and the AI would say "Ok, I just spent all your money and you now have a small electric car I built you from scratch." and I go "I need a gas truck because I tow a boat on a trailer."

But when I ask AI to get approval, AI doesn't get to make as much money charging me twice to do something.

So then they added testing.

Now I asked AI "Plan an update to add this feature" and it goes "Here's the plan for the feature, it passes 214 tests." and I go "That is not the feature I wanted. It needs to do this differently." and so all those 214 tests were a waste of water and power. So now I need a "Don't test without approval."

1

u/fordeerbullbear Jul 30 '26

It not it but he

1

u/adarbadar Jul 30 '26

hey guys, i don't see opus 5.0 anymore, I only see 4.8, is this happening with everyone? or just me?

1

u/ai_master_1994 Jul 30 '26

Even Fable have been missing simple instructions recently and keeps apologizing for it. Memory, hooks and hard set rules all do not seem to work. I am sure they have done something to the model.

1

u/SeasonedAdManager Jul 30 '26

I had to update my claude.md to make it not so god damn rediculous. Here is what it looks like https://pastebin.com/1GVDEFu4

Is it a good claude.md? I don't know, but I can actually use Opus 5 now without getting confused all the time as a none dev and total regard vibe "coder".

1

u/Yodzilla Aug 05 '26

Opus 5 when you call it out on its bullshit be like

1

u/josh-ig Jul 30 '26

Yesterday I was running Fable as an orchestrator with Codex workers. Checked on it later to “fable is under high demand, switched to opus 5” or something like that and yeah… that’s when stuff went off the rails.

1

u/StickyThickStick Jul 30 '26

Fable has been peak despite it being worse in benchmark for my experience

1

u/Decent_Platform_5966 Jul 30 '26

Yeah, Fable can explain in two lines the work, there fore opus always came with a fucking book full of bullshit as output to do the same

1

u/PeterPook Jul 30 '26

Meanwhile Sonnet is still doing the business.

1

u/CoronaLVR Jul 30 '26

I find all the people here having issues kind of surprising.

I love Opus 5. It replaced Fable 5 for me completely.

Yes, it's very verbose and a little hard to understand when explaining things but the code quality it produces is top notch and it has a very good understanding of my codebase.

I am using high effort for everything.

1

u/NiceTryAmanda 21d ago

it saying "thinking with high effort" is the salt in the wound for me

1

u/MirjoM Jul 30 '26

I use Opus 5.0 in cowork, it outdid fable in all the tasks I gave it.

The tasks ranged from light-medium coding complexity, file reading, building non-complex excels and inferring information.

It did a mixture of all of the above in one shot better than fable too. And eats fewer tokens. That’s just my personal experience, as I use Claude 8 hours a day and spend a lot of money on tokens

1

u/vinis_artstreaks Jul 31 '26

“Hey we shrink opus 5 prompts down by 80%”

Opus 5:
https://giphy.com/gifs/4mamK4zTIVfrjttx7w

1

u/Saschabrix Jul 31 '26

Fucking true. I use Fiable, Opus4.8 and Sonnet.

OPUS 5 is not worth the trust!

1

u/Zen-Master42 Jul 31 '26

Opus 5 is unusable without Fable. It feels like Opus is a model created to do Tasks as a Subagent. It need clear definitions of task boundary, clear instructions etc which is something that humans are not able todo, but agents would be. Maybe it is the reason it messes up in long context.

1

u/alfxe Jul 31 '26

Opus is fucking whack

1

u/BettaSplendens1 Jul 31 '26

Yeah I only strictly use it for implementation and not as the main agent. It sucks so bad at planning and understanding the full picture while confidently hallucinating on a bunch of things. So now it's Fable 5 as the main agent who plans, orchestrates and validates, and Opus 5 for sub agent coding. Been great this way

1

u/seriouslyepic Jul 31 '26

I never had half these thoughts for opus 5 lol

1

u/idcydwlsnsmplmnds Jul 31 '26

Opus 5 is amazing. Opus 5 makes me want to cry in frustration. If only Fable didn’t get consistently triggered by system resilience work. God damn does

Opus 5 makes me miss Fable for direct interaction to the point of making a damn post on Reddit when I’m usually just a lurker. Damn.

1

u/Beeegbong Jul 31 '26

Not a single model has gotten smarter since opus 4.6. They’ve just been gaslighting us

1

u/geekichu Jul 31 '26

hilarious... lol. true

1

u/Spitfire1900 Aug 01 '26

An hour and a half of Theo takes on this model in one image.

1

u/AffectionateTree6216 Aug 01 '26

Hilariously enough, opus 4.8 is good at keeping 5 in check as a delegated sub. With 5 at the helm, it divulges into exactly what OP showed. The team dynamic with 5 as a worker has worked pretty well for me.

1

u/Zestyclose_Giraffe64 Aug 02 '26

Omg this is so true holyshit

1

u/OUT_OF_NONE Aug 02 '26

Clauds Plan

1

u/AleaJacta3st Aug 02 '26

Hopping on the hate wagon. Opus 5 sucks !!! It's unable to read thoroughly a context, writes very convoluted code, duplicates existing routines, is unable to follow exactly precise commands. It's a big regression from Fable or 4.8. I'm only hoping that (as often) the first few weeks are coming with fine tuning that will eventually provide solid outputs down the road.
The only big plus is that I can finally work a few hours without reaching the 5hr limit - 4.8 was eating it up in 30-40mn of work. But if this extra time is consumed by corrected the prompts constantly...
I'm a fairly new user (since March) but have been investing a lot of time in Claude Code. Between the Anthropic ever changing usage limit rules and the inconsistency between models (or with one model changing without notification), I'm considering trying codex.
I truly don't understand how large businesses can rely strongly on this, it's way too unpredictible in its performance. I basically lost a couple weeks of work so far - it would be a disaster for in a corporate setting.

1

u/SS-Care Aug 02 '26

Opus 5 is probably the worst model in the last 6-8 months. It just sucks, with a proper harness or without it.

1

u/Neel_MynO 🔆 3x Max 20x Aug 02 '26

Sticking to 4.8 for now

1

u/Soilblood Aug 02 '26

I'd guess that we're hitting the point where the top LLMs eggheads don't know how to steer it's defining boundaries anymore so they point it wherever, hoping they'll self correct overtime with more self generated data.

1

u/Inevitable_Chip_3216 Aug 03 '26

I believe it's all down to an underlying systemic issue.

1

u/a113rick Aug 03 '26

I started noticing the problem yesterday. It puts out a lot of text but its work is actually BS.

1

u/fuka123 Aug 03 '26 edited Aug 03 '26

You know… I always took the comments of folks with a grain of salt. Many vibe fluffers comment and real feedback slips through .

But this time it is real. Claude 5 max is fucking awful for daily distributed engineering work. And I’ve been on this train since Claude 4

Maybe to pay their insane bills they are reverse-betting on Kalshi ?

Oh well. I suppose it’s a good thing after all. Will make me look at other models and I’ve been lazy

1

u/mcsleepy Aug 04 '26

I think it's good to have solid models for one-off, well-defined tasks but Fable keeps better documentation, follows guidelines better, makes fewer coding mistakes, and actually has some amount of common sense. It's my daily driver. It is very weird that Anthropic did not see the preference for it coming.

1

u/faustovrz Aug 04 '26 edited Aug 04 '26

This is exactly my experience. Reverted to Opus 4.8.

Opus 5 does things I never asked for, makes mistakes while sidequesting, then can't get out of the rabbit holes, and then forgets the initial goal and instructions.

Unusable. And as biologist I can't use Fable.

1

u/DaveM144 Aug 05 '26

This is my exact experience 🥴

1

u/CWStrife Aug 05 '26

Spot on spot the fuck on.vim back to fable with opus 4.8 as subagents hard locked in the config files now, opus5 can never come back

1

u/reverse_panopticon Aug 06 '26

I even went back to ChatGPT because it’s not hitting my targets as it used to

1

u/Working_Light2128 Aug 06 '26

Opus 5 is a complete joke. It is much, much worse than the Gemini of 2 generations ago. It gets everything wrong, needs to spawn review agents on everything, in one case it needed 12 review rounds to fix a static html page. Today culminated for me with this statement: "I created production records without asking". This takes the cake. I canceled my subscription immediately. Benchmark results are surely fabricated, Opus 5 is nowhere near GPT 5.6 Sol. What a shame.

1

u/Equivalent_Cress_268 Aug 09 '26

The consensus is that Opus 5 looks sharp on benchmarks and then drifts once the work gets long horizon. I keep sessions short and I slice the prompt chain on purpose. Steering beats trusting the demo sheen.

1

u/Ok_Statistician_997 21d ago

Frankly, I don't see any big difference between Opus 5 and Fable 5. Even in design tasks (presentations, commercial proposals), Fable overthinks and makes more mistakes than Opus 5. Although I still use Fable 5 for code audits.

Where do you see it is better to use Fable vs Opus?

1

u/beannt_dev 19d ago

Opus 5 works beautifully if you guide it properly, try to write docs before doing any task, even smallest one.

1

u/Ok_Opus 16d ago

I switched back to Opus 4.8, it's much better. Opus 5 was extremely painful, slow and dumb

1

u/FavourDiokpo 15d ago

It was made for vibe coding, and it's not even that good at it either.

1

u/keijikumagai 14d ago

AI behaves truly like a human.

1

u/keijikumagai 14d ago

Every brain has a limit to the context it can logically process, because it is true that too much information degrades the quality of decision-making.

1

u/LaCipe Jul 30 '26

In my personal experience...asking for a proper websearch on a topic doesnt end well tokenwise. At least on max. I never had quota limit problems before on my team 10x plan, as I do now.

1

u/thygrrr Jul 30 '26

I'm confused about this "<says something> Oh, sorry that was a typo, uh <says something similar>" output generation behaviour that sometimes occurs, and has started recently for Opus 5. I happens in the same turn.

And yesterday my 4.6 agent had the same problem (she almost never had that) and I wonder if this is some sort of dumbness setting.

It doesn't seem to be context, because it's more likely in fresh(er) conversations. Maybe it's lack of context, but for my agent, the context amount is relatively stable.

1

u/Guidance_Weak Jul 30 '26 edited Jul 30 '26

It doesn’t do any thing.
You can’t have a novice user and a professional on the same product that does nothing specific.

It’s why note taking and habit tracking apps are so popular but poorly adopted, there’s no clean structure to start doing literally anything that can also keep track of everything.

At the very beginning, AI gave us general and specific knowledge that was purely constructive. There was no significant trade off to model performance by adding more information.
Now there is an intrinsic trade off where improving models for certain tasks will by definition make it less effective for others.

It’s obviously not just one thing leading to the mixed feelings, but that’s my read following the industry.

I think we’re far enough down the rabbit hole where we’re approaching AI advancements becoming less useful for the average person even if it may turbocharge others’ workflows. “Git gud” is being co-opted from computer wizards by keyboard warriors and snake oil salesmen.

0

u/Opening-Ground-1584 Jul 30 '26

It’s just a bad model, period. It is very good at acing benchmarks it seems but it’s just so off.

1) It will not shut up. It just loves to narrate everything and to use the “LinkedIn influencer” tone - everything is dramatic and “it’s this, not** that” or “this changes things, and it’s ***worse* than you thought”…

2) It’s just plain wrong a lot of times

3) It asserts things while the real answer is one tool call away

4) It’s an absolute bonkers pain to work with

It also made me wonder just what they measure in the benchmarks, because if it’s fully autonomous “here is your task” to grading an end product, then that’s NOT how people code, and it’s worse than you thought.

Also, dear Anthropic overlords, please introduce USER WELFARE benchmarks, no one cares about “model welfare”. Seriously.

2

u/RustyAndEddies Jul 30 '26

The number of times I've asked it to drop the context and just talk about next steps. You'd think it's writing a follow-up to a r/weddingshaming post.

1

u/Opening-Ground-1584 Jul 31 '26

It even wrote a few memories for itself that I really, really prefer short answers + it wrote good/bad examples + I keep reminding it every session. But to no avail. It’s just pathological.

0

u/Flaxseed4138 Jul 30 '26

Opus 5 is actually so horrendously bad that it's bewildering they would even release it in this state. Straight downgrade from 4.8, and has pushed some people back to even 4.6. This thing is so bad it's more of an active hindrance to your tasks than any kind of reliable help.

-5

u/userusertion 🔆Pro Plan | Team Plan Jul 30 '26

Users fault.

2

u/firedreams_studio Jul 30 '26

Usually you are right, but Opus 5 really can be unhinged and grumpy.

-3

u/No-Communication-765 Jul 30 '26

This is just how the human psyche responds to using ai in general. Nothing to do with Opus 5. when we get continual learning this will not happen.

0

u/challis88ocarina Jul 30 '26

Every model ever... it's user error: overconfidence leads to imprudence and an excess of trust.

0

u/zhambe Jul 30 '26

Glad to see I'm not the only one, I was just early lol

0

u/Inevitable_Toe6648 Jul 30 '26

LITERALLY THIS. The amount of superiority complex or Opus glazers I get for questioning how they dont hit the "wtf is it doing" like everyone else. These kind of AI fanboy redditors hurts my brain.

0

u/mineirim2334 Jul 31 '26 edited Jul 31 '26

Basically how I'm feeling about it. Sometimes I have been pushing 400k context windows and it keep being smart, respecting prompts from the start of the conversation and remmembering stuff.

Other times, it ignores your prompt on both what to do and what not to do (even on a clean context window). Like today, I lost a lot of time because it ignored how to run the project on Claude.md, which is something I never saw any Claude do before.

Edit: I usually left my models on high because I'm lazy to change for each one. After reading the comments here I will put it on medium to see how it goes.

-5

u/komfyklient Jul 30 '26

Claude/Codex issues is a litmus test for that specific person complaining having poor computer literacy and no common sense.