ClaudeFolio
News

Claude Opus 5 Is Here: Near-Fable Power at Half Price, and the Babysitter Came Too

Edward Kwun··4 min read
Claude Opus 5 Is Here: Near-Fable Power at Half Price, and the Babysitter Came Too

Key points

  • Claude Opus 5 launched July 24 with near-Fable intelligence at half the price, $5/$25 per million pricing, and default status on Max plans.
  • The benchmark claims are strong: within 0.5 percent of Fable's peak CursorBench score at half the cost, and triple the next-best model on ARC-AGI 3.
  • The Fable-style safety regime ships with it: flagged requests silently fall back to Opus 4.8, the same swap-and-hope architecture that made Fable feel like babysitting.
  • Anthropic says Opus 5's cyber classifiers intervene around 85 percent less often than Fable's, a real number worth holding them to, but the vigilance tax exists at any nonzero swap rate.
  • The asks are unchanged two models in: loud labeled swaps, context-aware classifiers, and published false-positive rates instead of a launch-day percentage.

Anthropic shipped Claude Opus 5 today, and the headline specs are excellent: near-Fable intelligence at half the price, state-of-the-art coding scores, and it's already live everywhere, Claude.ai, the API, Claude Code, and Cowork, as the new default on Max and the strongest model Pro users get. If the numbers hold up, this is the workhorse model most of us will actually live in. However, the part early testers have been grumbling about since the model leaked: the safety-classifier setup that made Fable 5 feel like babysitting came along for the ride. Flagged requests still get rerouted. The babysitter (you) has another kid to watch.

The good news first

Let's not bury what's impressive. Anthropic's pitch is "a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price," and the benchmark claims back it: within 0.5 percent of Fable's peak CursorBench score at half the cost per task, more than double Opus 4.8's Frontier-Bench performance, an ARC-AGI 3 score three times the next-best model, and it beats Fable's best OSWorld result at about a third of the cost. Pricing stays at $5 per million input tokens and $25 per million output, the same as Opus 4.8, with an optional fast mode at double the price for roughly 2.5 times the speed.

Put that next to Fable coming back to Max plans at 50 percent of limits and the lineup finally makes sense: Fable for the hardest problems, rationed; Opus 5 as the everyday frontier, uncapped-ish; and after months of model whiplash, that's a lineup I can plan around. 

Claude Opus 5 Is Here: Near-Fable Power at Half Price, and the Babysitter Came Too
Image Credit: Anthropic


 

Claude Opus 5 Is Here: Near-Fable Power at Half Price, and the Babysitter Came Too
Image Credit: Anthropic

Now the asterisk*

Opus 4.8 serves as the fallback when safety classifiers refuse for Opus 5, exactly as it does for Fable 5. Same architecture, same silent swap, same "wait, who am I talking to right now?" tax I documented after a month inside Fable's filter regime. This wasn't a surprise to anyone watching the leaks, either. When the model surfaced early in Cursor as "Claude Honeycomb EAP," the description that spread fastest across the developer communities wasn't the million-token context, it was "per-turn controls and safety fallbacks," with early-access users reporting sensitive prompts getting routed down to Opus 4.8 instead of answered. 

Anthropic's counterargument is right there in the announcement: the cyber classifiers on Opus 5 intervene "around 85% less often than they do for Fable 5." That is a big reduction if true. But notice what it concedes: the regime itself is now standard equipment on every frontier Claude, and the argument has shifted from whether your session gets swapped to how often. If you're a security person, a biologist, or anyone whose normal Tuesday brushes the tripwires, 85 percent less than constant can still be weekly. And the babysitting cost was never only the frequency, it was the vigilance: checking which model actually answered before trusting the output. That tax exists at any nonzero swap rate, because you can't know in advance whether this session was one of the flagged ones.

The pattern is the story now

Step back and the individual model launches blur into a policy: Anthropic's frontier models now ship with capability-triggered classifiers that quietly substitute a weaker model, and the dial gets tuned after launch based on how loud the complaints are. We watched the full arc with Fable, over-blocking, apology, retraining, more over-blocking after the ExploitGym breach made the loose end of the tightrope vivid. I understand why the rail exists, I wrote a whole piece defending the impossible balance. But two models in, the asks haven't changed and haven't been met: make the swap loud and labeled, let the classifier tell my server from a server, and publish the false-positive numbers instead of a single launch-day percentage.

The 85 percent claim is, at least, a number, which is more than Fable launched with. Consider this the benchmark I'll be holding them to when the early grumbling either fades or doesn't. Because Opus 5 is positioned as the model everyone uses every day, and that cuts both ways: if the classifiers really are 85 percent quieter, most people will never meet the babysitter, and this section of the review ages into a footnote. If they're not, the complaint volume is about to be a lot larger than Fable's ever was, because the audience is.

Should you use it?

Yes. Near-Fable capability at Opus pricing, as your plan's default, is the best everyday deal Anthropic has shipped on paper. My advice is the same calibrated-trust playbook as always: enjoy the model, watch which model actually answers when your work drifts anywhere near security or biology, and say something loudly when the swap hits you doing something legitimate, because the Fable saga proved beyond doubt that loud users move these dials. Opus 5 looks like the model of the year on paper. Whether it feels that way in six weeks depends a lot less on the benchmarks than on how often its minder interrupts.

Sources

Anthropic: Introducing Claude Opus 5 - The July 24, 2026 launch: near-Fable intelligence at half the price, $5/$25 per million token pricing with a 2x-price fast mode, default on Max and strongest on Pro, the benchmark claims across Frontier-Bench, CursorBench, ARC-AGI 3, and OSWorld, Opus 4.8 as the safety-classifier fallback, and the claim that cyber classifiers intervene around 85 percent less often than Fable 5's.

AI Reiter: The Claude Honeycomb leak - The pre-launch appearance of the model in Cursor as Claude Honeycomb EAP, described with per-turn controls and safety fallbacks, and early-access reports of sensitive prompts routing down to Opus 4.8.

Tech Times: Opus 5 leak surfaces in Cursor - Coverage of the July leak window, the safety-fallback detail spreading across Hacker News and X, and the context of Fable's extension saga.

FAQ

What is Claude Opus 5?
Claude Opus 5 is Anthropic's new flagship everyday model, offering near-Fable 5 performance at roughly half the cost while becoming the default model for Max subscribers and the strongest model available to Pro users.
Is Claude Opus 5 better than Claude Fable 5?
For most everyday work, Opus 5 offers nearly the same capability at a much lower cost, while Fable 5 remains Anthropic's highest-capability model for the most demanding tasks.
Does Claude Opus 5 still use safety classifier fallbacks?
Yes, certain requests can still be routed to Claude Opus 4.8 by Anthropic's safety classifiers, although the company says cyber-related interventions occur about 85% less often than with Fable 5.

Related posts

Comments

Claude Opus 5 Is Here: Near-Fable Power at Half Price, and the Babysitter Came Too · ClaudeFolio