GPT-5.6 vs Claude Fable 5: what the benchmarks really say
One leaderboard puts Anthropic ahead, another puts OpenAI ahead. We line up the numbers published after the GPT-5.6 launch and show which of them can actually be verified.
On SWE-Bench Pro, Claude Fable 5 led GPT-5.5 by more than twenty percentage points. After GPT-5.6 reached general availability on 9 July 2026, that gap narrowed noticeably, but it did not close. On the Artificial Analysis Coding Agent Index, however, OpenAI's new model took the lead, so the answer to "GPT-5.6 vs Fable 5" depends on which leaderboard you decide to treat as yours.
In short
- Anthropic reported 80.3% on SWE-Bench Pro for Fable 5, against 58.6% for GPT-5.5. That is a vendor-reported score, not an independent measurement.
- Independent aggregators list Fable 5 at 80.0% and the strongest GPT-5.6 variant, Sol, at 64.6%.
- On that test the gap shrank from more than twenty percentage points to about fifteen.
- On the Artificial Analysis Coding Agent Index the order flips: GPT-5.6 Sol scores 80 points, 2.8 ahead of Fable 5, while using less than half the output tokens.
- The 1.5-million-token context window remains unconfirmed. It appears neither in OpenAI's own communication nor in launch-day coverage.
Where Fable 5's lead came from
Claude Fable 5 arrived on 9 June 2026 as the first Mythos-class model released broadly. Anthropic described it as state of the art on almost every benchmark it tested, but the announcement itself gives no percentages for individual tests.
The figure quoted most often came from SWE-Bench Pro, a set of real issues taken from code repositories. The model has to find the right file on its own, write a fix and pass the tests. In Anthropic's comparison, reproduced by Vellum among others, Fable 5 scored 80.3% against 58.6% for GPT-5.5 and 69.2% for the earlier Opus 4.8. A gap of more than twenty points on a test like that is not measurement noise.
Who measured that score, and why it matters
The 80.3% figure is the vendor's own. It comes from Anthropic's comparison, not from a measurement that someone outside the company repeated, and none of the sources that quote it describe the conditions under which the model worked through the tasks.
Independent aggregators put Fable 5 slightly lower, at 80.0%. The difference is cosmetic; the principle is not. Until someone outside the company reproduces a measurement, a figure from a launch announcement is marketing material, not evidence.
What GPT-5.6 actually delivered
The launch slipped. Instead of a June release, OpenAI opened a limited preview to a small group of partners on 26 June 2026, and general availability followed on 9 July 2026. Before that, as TechCrunch reported, the US administration had sought to restrict the model's rollout over fears of misuse. Anthropic's models had already faced similar restrictions, which we covered in our piece on the US government block on Fable 5 and Mythos 5 (in Polish).
The family comes in three variants: Luna (the cheapest), Terra (mid-tier) and Sol (the most capable). After launch, Sol scores 64.6% on SWE-Bench Pro, Terra 63.4% and Luna 62.7%. On the Coding Agent Index run by Artificial Analysis, Sol set a new record of 80 points, 2.8 above Fable 5, and did so using less than half the output tokens. OpenAI also claims 54% higher token efficiency on coding tasks.
Closing the gap, not erasing it
Side by side, the published figures look like this:
- SWE-Bench Pro, vendor-reported: Fable 5 at 80.3%, Opus 4.8 at 69.2%, GPT-5.5 at 58.6% (Anthropic's comparison, as reproduced by Vellum).
- SWE-Bench Pro, after the GPT-5.6 launch: Fable 5 at 80.0% and GPT-5.6 Sol at 64.6% in independent aggregators' tables; Terra scores 63.4% and Luna 62.7%.
- Coding Agent Index (Artificial Analysis): GPT-5.6 Sol at 80 points, 2.8 ahead of Fable 5, on less than half the output tokens.
- Intelligence Index (Artificial Analysis, early September 2026): Fable 5.1 at 66 points, GPT-5.6 Sol at 61.
Read honestly, this is progress without a knockout. On SWE-Bench Pro the distance between the leader and OpenAI's strongest model shrank from more than twenty percentage points (80.3% against 58.6%) to about fifteen (80.0% against 64.6%). That is real progress for a single release, and still a loss on this particular test. On the Coding Agent Index the order is reversed and Sol is in front. On the Intelligence Index from the same Artificial Analysis, measured at the start of September, Anthropic's newer Fable 5.1 is ahead again.
The conclusion is awkward for headlines: there is no single winner, only a winner per benchmark. "Which model is better?" asked without qualification has no correct answer. The useful version of the question is: better at what, on what budget and in what configuration.
The other side of the coin: what we do not know
The 1.5-million-token context window that featured in many pre-launch previews appears neither in OpenAI's materials nor in launch-day coverage. Until it shows up in the documentation, it is an unconfirmed number and should be treated as one. Window size also tells you less than it seems to, because models can lose details from the middle of a very long document, a problem we covered in how AI gets lost in long text (in Polish). How retrieval systems cope with limited context by cutting pages into smaller pieces is explained in chunking: how AI splits your page into usable pieces.
A bigger context window helps a model see more at once. It says nothing about what the model will do with that context.
The second unknown is the quality of the task set itself. BenchLM notes that an OpenAI audit in July 2026 judged around 30% of the benchmark's public tasks to be flawed and withdrew its earlier recommendation to use it. A score on this test therefore needs its configuration and task selection checked before it becomes the basis for a purchasing decision.
The third point is the simplest. All of these tests measure engineering work in open-source repositories. Your own tasks are most likely something else entirely.
What this means if you use AI rather than build on it
Unless you are building your own product on top of one of these models, none of the figures above should decide your choice of tool on its own. There are three practical consequences.
Cost. What matters is not the price per million tokens but how many tokens the model spends on your task. A model that is dearer on the price list but finishes the job in half the steps can come out cheaper on the monthly bill. That is why Sol's result on less than half the output tokens deserves as much attention as the headline score.
Use case. Models diverge far more on agentic work (multi-step tasks with tools and long sessions) than on straightforward writing. If your use is product descriptions, summaries and email replies, the difference between the top of the leaderboard and a model one tier down will be invisible to you, while the difference on the bill will be very visible.
Visibility. Which model your team uses internally has no bearing on whether AI assistants recommend your company to customers. That is separate work: earning a place in the sources assistants quote, covered in where AI citations come from, and making your own pages easy to retrieve and cite, which is the job of Generative Engine Optimization.
How to test it on your own work in a week
Instead of trusting other people's tables, run your own micro-test. A week and a spreadsheet are enough.
- List five tasks you genuinely hand to AI at work. The repetitive ones, not the ones from a demo.
- Run each one on two models from different families, with an identically worded prompt.
- For every run, record three things: whether the output is usable without edits, how many minutes the edits took and what the call cost.
- After five tasks you have your own ranking. It will differ from every public one, and that is the point, because it measures your case.
- Repeat after every major release. The top of the leaderboard reshuffles every few weeks: Fable 5.1 and ChatGPT Astra entered the race before this particular contest could be settled. What those two releases change for search is covered in Fable 5.1 and Astra: what changes for SEO and GEO.
Whichever model wins your test, being found by the assistants your customers use is a different question with different evidence. The groundwork is laid out in SEO vs AEO vs GEO.
Common questions
Which model is better at coding, GPT-5.6 or Claude Fable 5?
It depends on the benchmark. On SWE-Bench Pro, Claude Fable 5 sits at around 80% against 64.6% for GPT-5.6 Sol. On the Coding Agent Index run by Artificial Analysis the order is reversed: Sol has 80 points, 2.8 ahead of Fable 5.
What does GPT-5.6 score on SWE-Bench Pro?
After launch, the Sol variant scores 64.6%, Terra 63.4% and Luna 62.7%. Independent aggregators list Claude Fable 5 at 80.0% on the same test.
Has GPT-5.6 caught up with Claude Fable 5?
Partly. On SWE-Bench Pro the gap shrank from more than twenty percentage points to about fifteen, and on the Coding Agent Index GPT-5.6 Sol took the lead. It is not ahead on every test at once.
Does GPT-5.6 have a 1.5-million-token context window?
That figure comes from pre-launch previews and has not been confirmed, either in OpenAI's own communication or in coverage from the launch day of 9 July 2026. Until it is confirmed, treat it as uncertain.
How should I choose a model for my own tasks?
Not by the leaderboard. List five tasks you genuinely hand to AI, run them on two models with the same prompt and record three things: whether the output is usable without edits, how long the edits took and what the call cost.
Sources
- Anthropic, launch of Claude Fable 5 and Mythos 5 (9 June 2026)
- Vellum, Fable 5 benchmark breakdown (80.3% and 58.6% on SWE-Bench Pro, 69.2% for Opus 4.8)
- llm-stats, SWE-Bench Pro table (80.0% Fable 5, 64.6% GPT-5.6 Sol, 58.6% GPT-5.5)
- BenchLM, SWE-Bench Pro table and caveat on task quality (8 September 2026)
- TechCrunch, launch of GPT-5.6 Sol, Terra and Luna and the score of 80 on the Coding Agent Index (9 July 2026)
- heise online, Intelligence Index of 66 for Fable 5.1 against 61 for GPT-5.6 Sol (2 September 2026)
Read next
Find out whether AI recommends your company.
Start with the free SEO and GEO audit, delivered in 5 working days. We check how the models describe your brand and hand back a prioritised list of changes.