Tue, 1 Sep

Anthropic Releases Claude Fable 5.1: Doubles Scientific Performance, Cuts Cached Context Costs by 4x, and Expands Context to 1M Tokens

Max Ivanov · 01.09.2026 21:29 · 4 min read

Anthropic has unveiled its new flagship AI models: Claude Fable 5.1 and a specialized version, Claude Mythos 5.1. Designed for long-running autonomous tasks, large-scale code refactoring, and scientific research, the model more than doubles the performance of its predecessor in terminal environments and becomes significantly more affordable thanks to a fourfold reduction in cached context costs.

The model, identified as claude-fable-5-1, is now available in the Claude web version, Claude Code, the official API, and cloud services including AWS Bedrock, Google Cloud Vertex AI, and Microsoft Azure.

Breakthrough in Terminal Benchmarks and Comparison with Competitors

The biggest technological leap for Fable 5.1 is in complex agentic environments, where the AI must independently write console commands, work with the file system, and analyze compilation errors.

According to the official Anthropic announcement, results on key standardized tests are as follows:

BenchmarkClaude Fable 5.1Claude Fable 5Claude Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.152.6%24.7%29.0%22.4%
AutomationBench31.4%17.1%26.9%19.6%
Terminal-Bench 4.055.8%42.0%46.2%
CursorBench 3.2.073.4%70.5%70.0%67.2%
GDPval-AA v2 (Knowledge Work)1853172318241711
Humanity’s Last Exam (with tools)65.0%63.8%63.6%
OSWorld 2.0 (Strict / Computer Use)41.7%36.1%39.6%

The largest gap appears in the scientific test Terminal-Bench-Science, where Fable 5.1 outperforms Fable 5 by 2.13x (52.6% vs. 24.7%) and more than doubles the flagship GPT-5.6 Sol (22.4%). In classic code autocomplete tasks on CursorBench, the improvement is more modest (+4.1% over the previous generation).

Experts will be able to form a complete picture of the competitive landscape after independent measurements are published on the Artificial Analysis analytics portal.

Autonomy Up to 38 Hours: Real-World Company Cases

Unlike standard chatbots, Fable 5.1 is optimized for scenarios where the agent works for hours without human intervention:

  • Millennium case: The model located and fixed a floating system error (occurring once per million runs) that human engineers had been unable to track down for five years. The AI correlated memory dumps (core dump) with the source code of external libraries;
  • MongoDB case: The agent autonomously explored the company’s repositories and within days built a working prototype of a complex service with continuous self-testing;
  • Ramp case: The model ran a continuous 38-hour debugging cycle on a machine learning pipeline, independently launching and analyzing six parallel experiments.

Fable 5.1 Economics: Caching 4x Cheaper

According to the Claude API pricing page, the base generation cost remains at $10 per million input tokens and $50 per million output tokens.

However, the cost of reading cached context (cache reads) has dropped fourfold, from $1.00 to $0.25 per million tokens. For programming and working with massive repositories, where the agent constantly re-reads the same codebases, this reduces overall costs by 25–45% compared to Fable 5.

According to the Claude platform documentation, the model features:

  • A context window of 1,000,000 tokens;
  • A maximum output of up to 128,000 tokens;
  • A knowledge cutoff date of June 2026;
  • An adaptive reasoning mode (Adaptive Reasoning), set by default to High in Claude Code and Medium in the web interface.

Differences Between Fable 5.1 and Mythos 5.1

Alongside Fable 5.1, Anthropic introduced a closed variant, Claude Mythos 5.1. It shares the same base architecture but with expanded allowances in cybersecurity and biology for verified corporate and government researchers.

In the consumer-facing Fable 5.1, engineers reworked safety filters: the number of false blocks on legitimate code during vulnerability searches in Claude Code dropped by 60%, though creating working exploits and decompiling binaries remains blocked.

Scientific Discoveries: Mapping Venus and Bioengineering

The new architecture’s capabilities were tested on real scientific tasks:

  • Astronomy: Processing archival radar data from NASA’s Magellan mission, Fable 5.1 independently trained a neural network and produced a detailed elevation map of a third of Venus’s surface with a resolution of 2–3 km (up from 10–20 km), improving terrain accuracy by 25%. The map has been released under a Creative Commons license to support NASA’s VERITAS and ESA’s EnVision missions;
  • Bioengineering: The Mythos 5.1 variant designed synthetic protein structures with tenfold higher binding affinity than the leaders of the Adaptyv Bio competition, and optimized seven open genomic neural networks, speeding up their computations by up to 2.5x.

Enjoy VseZavislo?

Add us to your preferred Google sources to see our news more often.

Add us to your Google

Share

Leave a comment