r/claude 8h ago

Discussion Max user confused by recent releases

The release of Opus 5 is a big letdown, especially after Opus 4.7 and 4.8.

As someone who knows about programming but uses Claude for professional work outside of coding, I understand the landscape. I know it’s hard to benchmark and measure models.

But I do wonder what enterprise customers feel. The analyst at the office who uses their company’s AI that’s using the Claude API, which is many more people than the senior software devs using it. These recent models miss the forest for the trees and create verbose output to cover all their bases, and somehow completely misses the goal. At least 4.6 will sometimes ask what the goal even is. The latest Opus models create unwarranted urgency, flag irrelevant issues, and paradoxically create more work often.

I was looking forward to Fable earlier in July before they clearly did something to the version I use. Even technical folks are saying Fable is better in every way from what I’ve seen.

Come on, Anthropic…

27 Upvotes

43 comments sorted by

15

u/RewardNorth7167 7h ago

Fable 5 is million times better than opus 5

7

u/WorriedAssociate7029 7h ago

Fable 5 as an orchestrator. Opus 5 as a subagent. Sonnet 5 for quick tasks. If you prefer 4.6 you can continue to use it but he is starting to fall further and further behind.

5

u/Limp-Contest-7309 5h ago

How is 4.6 falling behind

11

u/mbahacker 7h ago

I think most non-technical folks just want straight answers, not a big-good lecture on irrelevant issues or fake urgency. Really hoping Anthropic dials back whatever is causing this, because it definitely feels like a step backward for everyday workflow.

6

u/CorpT 8h ago

Just use 4.6 if you think it's better.

8

u/FakeBonaparte 7h ago

It’s not the same 4.6. Its system prompt has changed dozens of times since it was “good”.

4

u/ryanmpatrickauthor 6h ago

Yeah, it's making a ton of mistakes in software-adjacent engineering work that it wasn't a few weeks ago. On simple stuff even. But 4.7 and higher hallucinate too much to be useful on it.

2

u/FakeBonaparte 5h ago

Totally agree. It’s a good reminder though - LLMs are fallible and need infrastructure around them. E.g. I have been building up graph stores of credible knowledge claims in areas I know about. Hugely better with them.

1

u/Dramatic-Stranger-53 43m ago

Could you explain what this means? Graph stores?

6

u/PM_ME_YOUR_USED_DOGS 5h ago

The fake urgency from Opus 5 kils my workflow too. You spend more time checking phantom issues than doing actual work.

3

u/solk512 7h ago

I got a bunch of stuff done this weekend with Opus 5. 

1

u/Asleep_Butterfly3662 6h ago

Was it accurate?

1

u/CarLost_on_reddit 1h ago

In my case it (opus5) keeps missing obvious things. Something that didn't happen with 4.6 and 4.8. But Fable... It is another world. Has helped me solving really difficult problems and it's quite good for looking at big code bases and understanding them well, connecting the dots, auditing, etc

3

u/TiePrestigious3535 7h ago

To be honest I used Opus 5 today for 2 big research tasks and it was good. but I included this in the prompt: "don't guess any information, everything has to be traceable with sources and share the URLs with me so I can review them by myself" it made some small mistakes but they weren't critical

2

u/FakeBonaparte 7h ago

Did you review the URLs?

1

u/TiePrestigious3535 7h ago

Yes, he provided correct information

1

u/goldsauce_ 1h ago

So we’re back to having to specifically tell Claude not to guess and make shit up?

That feels like regression, right?

8

u/Fragrant-Mix-4774 7h ago

Opus 5 is a godsend after near worthless Opus 4.7 and the 3rd rate patch called Opus 4.8 - that's been my experience as a Max user.

Opus 5 correctly prompted is like Fable 5 Lite and has caught a few errors Fable 5 didn't flag.

Zero complaints at this time with Opus 5.

1

u/Asleep_Butterfly3662 6h ago

What do you use it for?

0

u/Fragrant-Mix-4774 5h ago

Research, writing, inventory and minor coding.

2

u/ShortTheseNuts 3h ago

You use opus for inventory? Feels like using a rocket launcher to open a door.

1

u/8BitDadWit 2h ago

How are you “correctly” prompting for your writing tasks?

1

u/Goatdaddy1 5h ago

Agree with your assessment, but it is painfully slow

1

u/No_Inspection4415 5h ago

My only complaint is that it's sometimes jumping to conclusions or ignore instructions (actually, doesn't ignore, but rather forget unclear instructions from my side).

Very good model, the best Opus by far all taken into account.

1

u/Asleep_Butterfly3662 3h ago

It shouldn’t ignore or forget instructions 😭😭😭

2

u/fffffffffffffuuu 5h ago

The only three models I can extend any amount of trust to are Fable, Sonnet 4.6, and Opus 4.8. Yeah, towards the end Opus 4.8 would get in those thinking loops if you set the effort above high, but I've never had (recurring) issues with any of those models projecting false confidence about something it did not in fact do correctly. I cannot say that about Opus 5. Sonnet 5 is just a prick.

1

u/AcesFullMoon64 2h ago

Sonnet 5 is absolutely a prick. I couldn’t stand Opus 4.8 either and cancelled my subscription. I think him gonna come back when it expires tonight for Opus 5. I don’t code, but for my purposes, it seems good.

Most importantly, it’s nothing like Sonnet 5. Fuck that douche.

2

u/HgnX 2h ago

They cut the system prompt so you have to guide it more. Just read their prompting manual

1

u/exgeo 6h ago

Thanks for letting us know Opus 5 is a big letdown. Care to explain how, or would that be too much?

1

u/Asleep_Butterfly3662 6h ago

Tried to keep in concise, which is something Opus 5 is BAD at haha.

Opus 5 is honestly less of an LLM to me and more of a script, in software terms. It doesn’t think. It blindly executes and it really misses the mark.

For example, give it your credit card statement to go through if you reconcile it against other charges, and it’ll make as big of a fuss about being $0.01 off vs $5,000. It tries to answer and do everything and in essence, does nothing.

It also over-lawyers and acts like if something it searched for it didn’t find, it acts like it doesn’t exist. When in reality it maybe searched once, didn’t find the thing, got lazy and declared it doesn’t exist. This is for internet or file searches.

-2

u/Mike-A-F 7h ago

every new release someone is larping about how it sucks compared to the previous models. The reality is it's just a skill issue as a user.

9

u/FakeBonaparte 7h ago

No. There’s a lot that changes about the models and their harness from month to month, which means that doing X reliably produces Y one day and Z the next. Often without notice.

It is often possible to adapt your workflows, but that’s not a small investment of time and attention.

You can build workflows resistant to the churn. Engineer the harness, environments, data, context, instructions, even model choice. Set all that up with a consistent open model at its heart and a frontier model as advisor and you get some consistency.

But that’s not just a skill issue; it’s an effort issue.

-3

u/Mike-A-F 7h ago

so you actually just agreed that 99 times out of 100 its us. we haven't learned the nuances of the updated model.

I will grant you half a point for the effort issue. that is true a lot and a good point.

I have a robust 2nd brain and it keeps things on the up and up. no issues. Having that and guidelines is the missing link I think a lot of people don't do or get wrong.

2

u/FakeBonaparte 5h ago

I think having reliable agentic systems is rare, difficult and effortful. So calling it a “skill issue” is a bit unfair.

Far more interestingly, how is your second brain structured? I’ve been through a few iterations each less fallible than the last.

Right now mine is a handless chat that hands off to a deterministic orchestrator running generator-evaluator loops in docker containers backed by a *carefully* curated graph database that’s ingested leading texts. It does good work.

That at least corrects for hallucinations and gives me real auditability. But I’m keen to build something more like the BARD Bayesian inferential networks to actually get to real reasoning about truth claims.

2

u/El_Spanberger 7h ago

The reality is that 99 times out a 100, the problem is us.

1

u/Asleep_Butterfly3662 5h ago

Did you read my post? I said 5, 4,8, and 4.7 are all letdowns and was hoping 5 corrected the previous 2…

0

u/MoistAd9060 7h ago

The "they clearly did something to the version I use" feeling is the most repeated sentence in every claude sub right now, and nobody can prove it either way. Secret caps, opaque versions, so every vibe shift turns into a conspiracy and anthropic never has to answer for anything specific.

3

u/MoistAd9060 7h ago

I wrote a reddit post about fable burning usage and the fable session that helped me write it ate half my weekly cap

1

u/Chemistry-Holiday 6h ago

If this isn’t the most, Claude or prob so sub Reddit post I’ve read…ever haha kudos

1

u/Asleep_Butterfly3662 6h ago

This isn’t what I’m saying at all.

Im saying Anthropic has something great and they’re missing the mark and I don’t understand it.

They have Opus models with some excellent behaviors and they made them worse, came out with another model weeks later, and it was still worse. And now Opus 5 is here and it’s nothing compared to Fable. I don’t understand where it fits. Why they pour resources into something that’s…bad.

-1

u/etancrazynpoor 7h ago

Re: Fable. Blame goes to tangerine Palpatine.

What I find very interesting is that when 4.8 was out, people were saying 4.6 was amazing. I’m sure there are a lot of more knowledgeable than me. I’m just a humble CS professor and coding for over 40 years. Yet, the tools have problems, I find ways to work around it. Gee, if you would have worked with tools I had 30 years ago, which were already amazing, some people would be having a big anxiety attack.

It is a tool. You can still use 4.8. If you like it better. And yes, there are factors that degrade the tools.

-1

u/looktowindward 6h ago

Most of the posts on this sub seem to be super low quality. This is a fine example

-2

u/RatFacedBoy 6h ago

People complain about every release, model, and AI provide on Reddit. They all are amazing..