Claude Fable 5 Reclaims Top Position: Latest Update on Its Recent No.1 Ranking Achievement

Claude Fable 5 Reclaims Top Position: Latest Update on Its Recent No.1 Ranking Achievement

Claude Fable 5 takes the top spot on the MirrorCode AI programming leaderboard, boasting a massive lead that leaves all other competitors far behind.

This time, Claude has completely broken the AI programming leaderboard with a massive leading gap!

On the just-updated MirrorCode leaderboard, Claude Fable 5 once again topped the list with an absolute success rate of 64%

GPT-5.6 Sol, which closely follows behind, only scores one third of that figure!

GPT-5.5, ranked fourth, performs even worse. It not only only achieves a 10% success rate, but is also completely outperformed by its own predecessor GPT-5.4

What is more surprising is that when using reate reaches 64%; when switched to the niche language Ada, its performance still stays as high as 61%

It is worth noting that in the open f Ada

But for Fable 5, the performance only drops by 3 percentage points

This is really interesting

If the model mainly relies on memorizing the syntax and common writing patterns of popular languages, its performance should drop significantly after switching to Ada

Yet the current result points to another possibility —

The most powerful models have got rid of the constraint of specific corpus, and begun to learn how to build a complete software project from scratch

Stuck at the 100% Clear Line, the 10-Billion-Token Limit Test

Specifically, the complete MirrorCode benchmark includes 25 target programs, covering fields such as Unix utilities, interpreters, data query, bioinformatics, cryptography, and compression tools

In the test, the model is placed in an isolated environment with no internet access, no access to the third-party dependencies. All it can get is high-level documentation, a part of visible tests, and the original program that can be called repeatedly

After that, it needs to continuously input data into the original program, observe the outputs to guess the internal logic, and gradually write a new program with completely consistent behaviors

The latest leaderboard selects 15 Medium and Large targets from the full set. Each target is implemented in two different languages, and each language runs three times

Moreover, the completion rate of both visible tests and hidden tests must reach 100% to be considered a pass — even 99.9% does not count

If one single edge case is missed, the entire run is still counted as a failure

To force the model to fill the last few gaps, MirrorCode pushes the single-run budget to 10 billion tokens, and allows a maximum continuous runtime of 7 days

The most costly task in the paper is even more extreme: the model ran continuously for 19 days, with a single run costing 2600 US dollars

During these 19 days, the model will repeatedly run the original program, compare results, add missing features, and then re-run all tests. When an error occurs, it checks the cause; when the output does not match, it switches to a new hypothesis; after the local part passes, it moves on to tackle the next gap

The whole process is more like a multi-day debugging session, rather than a single generation operation

The gotree example is the most intuitive demonstration

This bioinformatics tool originally has about 16,000 lines of Go code and more than 40 commands

Claude Opus 4.7 spent 14 hours and 251 US dollars, passing 2000 out of 2001 tests, reaching a 99.95% completion rate

Although it missed a rare edge case that processes date comments and was blocked by MirrorCode’s 100% clear line, it has compressed the engineering workload that used to take weeks to complete into just more than ten hours —

Epoch estimates that without the help of AI, human engineers need at least 2 to 17 weeks to complete the same task

In the Corpus Desert, Fable 5 Only Loses 3 Points

The MirrorCode paper once used StarCoder’s public training mix as a reference

Python accounts for about 8% of it, while Ada only accounts for 0.034%, meaning the former is about 230 times the size of the latter

Of course, we cannot know exactly how much Ada code the closed-nce, it is already intuitive enough how scarce Ada is

This language is mainly used in aerospace, national defense and other safety-critical systems. Its community size, number of tutorials and open

And Fable 5 is obviously not translating Go code into Ada line by line

It is more like figuring out exactly how the original program works first, then switching to another language to rebuild the exact same behavior

In contrast, other models do not perform so stably

GPT-5.6 Sol drops from 24% to 19%, GPT-5.4 drops from 21% to 12%, and GPT-5.5 even plummets from 17% to 5%. The moment the language is switched, the performance gap is immediately amplified

Handing Over the Entire Project to AI

Nowadays, Cursor has enabled hundreds of Agents to collaborate for nearly a week, writing more than 1 million lines of browser code distributed across 1000 files from scratch

Anthropic arranged 16 Claude Agents to run in parallel for nearly 2000 sessions, and finally built a 100,000-line C compiler that can compile the Linux 6.9 kernel

In the sample of individual Codex users disclosed by OpenAI as of May, 70.2% have submitted at least one task estimated to take human engineers more than one hour of workload; 25.6% have submitted tasks that take more than 8 hours

The unit of work people hand over to AI is changing from a single piece of code or a single bug, to an afternoon, a whole week, or even the entire project

The 64% score on MirrorCode is exactly a quantitative measurement of this kind of “project-level delegation”

It proves that as long as the goal is clear enough and the results can be automatically verified, cutting-edge models can already independently complete some medium and large software projects

For everyone who is handing over work to AI, this change is right around the corner —

In the past, you needed to supervise it to write every piece of code, but in the near future, you will most likely only check at key nodes to make sure it does not go off track

Code will become cheaper and cheaper

And those who can clearly describe the problem and clearly verify the results will become more and more valuable

References: https://epoch.ai/MirrorCode

This article is from the WeChat Official Account “AI Era”, author: ASI Revelation; editor: Moses, published with authorization from 36Kr

Related:

Digital Automation Training Benin: 5 Winning Skills Employers Demand in 2026

<a href="https://yoursite.com/automation-africa/" title="WhatsApp Marketing Automation Africa: 6 Dangerous Mistakes Brands Make in Nigeria”>
WhatsApp Marketing Automation Africa: 6 Dangerous Mistakes Brands Make in Nigeria

Want to learn this practically?

Join Justfine Infotech and build real digital skills in AI, automation, web development, digital marketing, office productivity, e-commerce, freelancing and cybersecurity.

Available Programmes:
6 Weeks Certificate • 3 Months Professional Certificate • 6 Months Diploma • Full Professional Diploma

WhatsApp:
+229 01 57 57 99 15
+229 01 66 68 11 60

Enroll Now

Source: eu.36kr.com

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top