Introducing Claude Fable 5.1 and Claude Mythos 5.1 Anthropic
Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise
Introducing Claude Fable 5.1 and Claude Mythos 5.1 Anthropic Read Post »


















