Co-Founder & Lead Programmer of AcceleratedLogic AI
Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1 in September 2026. This guide separates the vendor's benchmark reports, pricing changes, and safeguard notes, then explains what a buyer can—and cannot—conclude from them. It summarizes published material; AcceleratedLogic did not independently rerun these evaluations.
Fable 5.1 and Mythos 5.1 are not separate base models
Anthropic describes Fable 5.1 as generally available and Mythos 5.1 as the same underlying model with different safeguards and more restricted access. Mythos access is limited to trusted programs for approved cybersecurity and life-sciences work. When comparing their benchmark results, keep that safeguard difference in view: an allowed response and a blocked or redirected response can change a task score even when the base model is shared.
Anthropic's release announcement and system-card index are the primary references for capabilities, safety testing, and access conditions. Check the current product terms for your account; access and retention conditions can vary by plan and eligibility.
What the benchmark numbers show
Anthropic reports 52.6% on Terminal-Bench-Science 0.1 for Fable 5.1, compared with 24.7% for Fable 5 in its evaluation setup. The release gives a standard error of roughly 3.5–4.5 points per model and notes that results on the public leaderboard differ from its own reproduction. That uncertainty does not erase the reported gap, but it is a reason to report the benchmark version, harness, and variation with the headline score.
On Terminal-Bench 4.0, Anthropic reports 55.8% for Fable 5.1 and 60.9% for Mythos 5.1. Anthropic says the models share the same underlying system and attributes the difference to safeguards that intervened on some tasks. This is a result from Anthropic's evaluation, not evidence that Mythos is a generally better choice for ordinary coding work; access restrictions and intended use differ.
The release also reports scores for GDPval-AA v2, AutomationBench, CursorBench, OSWorld, and Humanity's Last Exam. These remain vendor-reported measurements. For example, Anthropic notes that some computer-use results are affected by safeguards, and that its OSWorld task release is not directly comparable to earlier versions. Read the benchmark-specific footnotes before comparing numbers across companies or releases.
The pricing change and its practical limits
Anthropic lists Fable 5.1 at $10 per million input tokens and $50 per million output tokens, with cache reads at $0.25 per million tokens. The announcement says this makes cache reads 75% cheaper than Fable 5, and estimates that this reduces total cost by about 25% on typical workloads and up to about 45% on highly agentic workloads in Anthropic's measured usage. Those are workload-dependent estimates, not guaranteed savings for every application.
A simple way to apply the change is to inspect your own usage breakdown. If an agent repeatedly reuses a large conversation or tool history, cached tokens may form a large share of input. If most requests are short and uncached, the cache-read reduction will have less effect. Compare total cost per accepted task, including retries and tool calls, rather than comparing the cache price in isolation.
Effort settings also change the cost-quality trade-off. Anthropic reports that Low or Medium effort can achieve results similar to or better than Fable 5 in some tested cases at lower cost. Treat that as a configuration to evaluate on your task set, not as a general guarantee.
A practical evaluation checklist
1. Select representative tasks and define what counts as a correct, complete result before testing.
2. Record model name, effort setting, safeguards, prompt, tools, benchmark version, and number of runs.
3. Track task success, error severity, latency, token usage, retries, and human corrections.
4. Compare cost per accepted result across effort settings and against a simpler model or deterministic baseline.
5. Keep security and data-handling review separate from capability scores; model access and terms are part of the deployment decision.
Fable 5.1's launch is notable for its reported agentic-science gains and lower cache-read pricing. The strongest conclusion for a buyer comes from combining those vendor disclosures with a small, reproducible evaluation on the buyer's own workload.