Introducing Claude Fable 5.1 and Claude Mythos 5.1 Anthropic

Introducing Claude Fable 5.1 and Claude Mythos 5.1  Anthropic
Terminal-Bench-Science 0.1Accuracy vs Cost
  • Fable 5.1
  • Fable 5

Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise

Terminal-Bench 4.0Accuracy vs Cost
  • Mythos 5.1
  • Fable 5.1
  • Mythos 5

Terminal-Bench 4.0 scores by cost (log scale), at each effort level. Claude Fable 5.1 and Claude Mythos 5.1 are the same underlying model; the gap between them reflects the tasks on which our earlier, less precise cyber safeguards intervened. With the improvements we’re making to these safeguards today, we expect the difference between the models to be much smaller

Humanity’s Last ExamAccuracy vs Cost
  • Fable 5.1 (with tools)
  • Fable 5.1 (no tools)
  • Fable 5 (with tools)
  • Fable 5 (no tools)

Humanity’s Last Exam scores by cost (log scale), at each effort level. CursorBench 3.2.0 scores by cost (log scale), at each effort level

CursorBench 3.2.0Accuracy vs Cost
  • Fable 5.1
  • Fable 5

CursorBench 3.2.0 by cost (log scale), at each effort level

Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol
Agentic scientific researchTerminal-Bench-Science 0.1 [1] 52.6% 24.7% 29.0% 22.4%
Agentic codingTerminal-Bench 4.0 55.8%60.9% (Mythos 5.1) 42.0% 52.3% 37.3%
Knowledge workGDPval-AA v2 1853 1723 1824 1711
Computer useOSWorld 2.0 [2] 77.9%partial 72.9%partial 75.4%partial —partial
Computer useOSWorld 2.0 41.7%strict 36.1%strict 39.6%strict —strict
Multidisciplinary reasoningHumanity’s Last Exam 60.9%no tools 57.8%no tools 56.6%no tools —no tools
65.0%with tools 63.8%with tools 63.6%with tools —with tools
Business workflowsAutomationBench 31.4% 17.1% 26.9% 19.6%
Agentic codingCursorBench 3.2.0 73.4% 70.5% 70.0% 67.2%
Fable 5.1 was evaluated with its production safeguards enabled. On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench. In all other interventions from our safeguards, cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5. This likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks.
Quote

“In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.”

CompanyJane Street Capital
AuthorCraig Falls, Head of Quantitative Research
01 / 22

Claude-designed protein binders (orange) for each of 12 targets (grey). Every design in the video was confirmed to bind in the lab. Structures shown are ESMFold2 predictions.
Radar image (Magellan): bright cone, radian lava flows
Magellan radar
Altimetry 10-20km footprint
New DEM (300m) a volcano 15km across
A small shield volcano on Venus
Inference speedup

Inference speedup for seven open-

Estimated cost savings on genome-wide analyses
  • Original implementation
  • Optimized

Estimated GPU cost of three genome-wide analyses before and after optimization, at cloud list price. Evo 2 40B saves more on a whole job (2.3x) than per forward (1.4x) because some of its optimizations only pay off across many sequences

Indexed cost of Fable usage
  • Cache reads
  • All other tokens

Indexed cost of running the same workloads on Fable 5 and Fable 5.1, at usage-based pricing measured at default effort over four weeks of actual usage in August 2026. Typical workload covers Fable usage across Claude Enterprise, Claude Code, and the API. Highly agentic workload covers context-heavy, tool-heavy work, where cache reads make up most of the cost

Try ClaudeStart building

1Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise

2OSWorld 2.0: Scores are on the benchmark authors’ August 2026 task release; Fable 5 and Opus 5 were re-run under the same conditions. Because the task files differ from earlier releases, these numbers aren’t directly comparable to previously published OSWorld 2.0 results, which is why no competitor score is shown

3These three targets are (EGFR, Nipah G, 15-PGDH) and come from Adaptyv Bio’s protein design competitions. The Nipah G comparison is against de novo designs targeting the receptor-binding site on the G head (best: ~8–12 nM, N1032). A de novo entry from Nick Boyd/Escalante Bio that targets a different region (the stalk) reached ~1.4 nM (design_7), comparable to our best binder

Related:

Digital Automation Training Benin: 5 Winning Skills Employers Demand in 2026

<a href="https://yoursite.com/automation-africa/" title="WhatsApp Marketing Automation Africa: 6 Dangerous Mistakes Brands Make in Nigeria”>
WhatsApp Marketing Automation Africa: 6 Dangerous Mistakes Brands Make in Nigeria

Want to learn this practically?

Join Justfine Infotech and build real digital skills in AI, automation, web development, digital marketing, office productivity, e-commerce, freelancing and cybersecurity.

Available Programmes:
6 Weeks Certificate • 3 Months Professional Certificate • 6 Months Diploma • Full Professional Diploma

WhatsApp:
+229 01 57 57 99 15
+229 01 66 68 11 60

Enroll Now

Source: www.anthropic.com

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top