RankingClaude Opus 5
Claude Opus 5
65 published benchmark measures · 21 benchmark families contribute across 6 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 6 effort levels across 68 benchmark/harness combinations →
Reported effort · Mixed settings
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Auto thinking: 2 observations
- Max: 19 observations
- Not specified: 85 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
21 contributing families across 6 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Claude Opus
- Released
- —
- Context
- 1,000,000 tokens
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/about-claude/model-deprecations
- Default Capability family coverage
/badge/claude-opus-5.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA-Briefcase | Supporting evidence | 1685Reported settings & sourceArtificial Analysis long-horizon knowledge projects; rubric and panel pairwise judging; Claude max effort. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase | Supporting evidence | 57.2%Reported settings & sourceArtificial Analysis; max effort; component rubric pass. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase | Supporting evidence | 1980 ratingReported settings & sourceArtificial Analysis; max effort; component analytical quality. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| AA-Briefcase | Supporting evidence | 1572 ratingReported settings & sourceArtificial Analysis; max effort; component presentation. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.4 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Agentic IF Index · internal | Supporting evidence | 59.1 indexReported settings & sourceInternal composite instruction-following evaluations; no fixed task count; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Agents' Last Exam · not specified | Supporting evidence | 55.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / Agents' Last Exam / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI · 1 | Hard reasoning | 97.5%Reported settings & sourceARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI · 2 | Hard reasoning | 90.42%Reported settings & sourceARC Prize semi-private validation; verified Fable5.1 max effort; comparator figures as reported in summary. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.16 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| ARC-AGI-1 · 1 | Hard reasoning | 97.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-1 / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | 90.4% | ARC Prize verifiedcontributes to capability | official board | 2026-07-24 |
| ARC-AGI-2 · 2 | Hard reasoning | 90.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-2 / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-3 · 3 | Hard reasoning | 30.2%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Abstract Reasoning table / ARC-AGI-3 / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Coding Agent Index v1.4 · v1.4 | Supporting evidence | 68.1 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Artificial Analysis Coding Agent Index v1.4 / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Artificial Analysis Intelligence Index | Supporting evidence | 63 | AA Intelligence Index | official board | 2026-08-12 |
| Artificial Analysis Intelligence Index v4.1.1 · v4.1.1 | Supporting evidence | 63.1 index scoreReported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / Artificial Analysis Intelligence Index v4.1.1 / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AutomationBench | Agentic | 26.9%Reported settings & sourcePrivate held-out board; simulated business-workflow app APIs; all assertions must pass; max effort stated for Fable5.1/Opus5. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.6 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| AutomationBench · not specified | Agentic | 26.9%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / AutomationBench / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AutomationBench · public v3 | Agentic | 50.3%Reported settings & source600public workflow tasks; deterministic end-state assertions;pass@1; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BenchCAD · not specified | Supporting evidence | 82.1%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; tools enabled; three Anthropic evaluation modifications Provider-published result; comparator measurements are not automatically independently reproduced. Astra launch footnote5: Claude scores use three modifications described in the Fable5.1 system card; not same controlled setting. GPT-6 Astra: A new generation of intelligence · Professional table / BenchCAD / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BenchCAD · Vision2Code 1000-file subset | Supporting evidence | 0.366 voxel IoUReported settings & sourceRandom1000 of17900files; five runs; adaptive thinking max; no tools; corrected camera prompt, raw shapes accepted, last code fence parsed. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BenchCAD · Vision2Code 1000-file subset | Supporting evidence | 0.821 voxel IoUReported settings & sourceRandom1000 of17900files; five runs; adaptive thinking max; with tools; corrected camera prompt, raw shapes accepted, last code fence parsed. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BioMysteryBench · Human Difficult | Supporting evidence | 51.8%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BioMysteryBench · Human Difficult | Supporting evidence | 49.4%Reported settings & sourceLinuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BioMysteryBench · Human Solvable | Supporting evidence | 91.4%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| BioMysteryBench · Human Solvable | Supporting evidence | 90.1%Reported settings & sourceLinuxterminal,bioinfotools,Python,R; allowlistednetwork; Gemini/GPTselfcomputed,Claudeproviderreports Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp | Agentic | 90.8% | BrowseComp reported | lab self-report | 2026-07-24 |
| BrowseComp · not specified | Agentic | 90.8%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Professional table / BrowseComp / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Chartography | Multimodal | 29.6%Reported settings & source100tasks; adaptive thinking max; five runs; no tools; tools condition has container, standard libraries and crop tool. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Chartography | Multimodal | 83%Reported settings & source100tasks; adaptive thinking max; five runs; with tools; tools condition has container, standard libraries and crop tool. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| CharXiv Reasoning | Multimodal | 83.7%Reported settings & sourceNo tools; Gemini/GPT/Opus selfcomputed; Sonnet selfreported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| CursorBench · 3.2.0 | Supporting evidence | 70%Reported settings & sourceCursor production agent harness; independently measured by Cursor; max effort. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.8 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| DeepSearchQA | Agentic | 90.4 percent F1Reported settings & source900questions; common search backend/browser harness; answer-set F1; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 | Coding | 73.6% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DeepSWE v1.1 · v1.1 | Coding | 73.7%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / DeepSWE v1.1 / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ExploitBench · not specified | Supporting evidence | 70%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitBench / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| ExploitGym · not specified | Supporting evidence | 22%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; launch-reported configuration Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / ExploitGym / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Finance Agent v2 · source release snapshot; version not specified | Supporting evidence | 58.6%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Extended (score) · 1.1 | Supporting evidence | 63.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Extended (score) / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierCode 1.1 Main (score) · 1.1 | Supporting evidence | 53.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / FrontierCode 1.1 Main (score) / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| FrontierMath Tier 4 (v2) · v2 | Hard reasoning | 73.2%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / FrontierMath Tier 4 (v2) / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| FrontierSWE · 2 | Supporting evidence | 0.52 fractionReported settings & sourceProximal agent harness; max effort; 34 tasks, five trials/task; mean score on 0..1 scale. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.5 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| GDP.PDF | Supporting evidence | 37%Reported settings & sourceAll-pass rate; allmodels selfcomputed byGoogle Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA | Agentic | 1735 | Artificial Analysis GDPval-AAcontributes to capability | official board | 2026-09-12 |
| GDPval-AA · 2 | Agentic | 1824Reported settings & sourceArtificial Analysis independent agentic shell/web evaluation;220GDPvalgoldtasks; blind pairwise Elo; max effort for Claude. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.15.3 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GDPval-AA · 2 | Agentic | 1824Reported settings & sourceArtificial Analysis publicboard snapshot; effort as reported Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA · 2 | Agentic | 1824Reported settings & sourceArtificial Analysis Stirrup shell/web harness;220tasks; blind pairwise comparisons; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GMMLU | Knowledge | 92.5%Reported settings & sourceMean accuracy42languages; adaptive max; one trial; no tools/custom system prompts. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| GPQA Diamond | Hard reasoning | 93.232% | GPQA Diamond reportedcontributes to capability | independent repro | 2026-09-11 |
| GPQA Diamond · not specified | Hard reasoning | 93.7%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / GPQA Diamond / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Harvey Legal Agent Benchmark · source release snapshot; version not specified | Supporting evidence | 6.7%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HealthBench | Supporting evidence | 67.1%Reported settings & sourceRaw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Professional | Supporting evidence | 73.4%Reported settings & sourceRaw rubric score; adaptive max; five trials; no tools/custom system prompt; Opus4.8grader; Fable5.1 safety fallback toOpus5. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.17 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Professional | Supporting evidence | 59.8%Reported settings & sourceLength-adjusted score; HealthBench Professional paper method; no tools; Opus4.8grader; five trials. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.1.A;8.17.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| HealthBench Professional (length-adjusted) · not specified | Supporting evidence | 56.4%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; official paper scoring; length-adjusted; OpenAI reproduction; GPT-5.4 grader; unclipped Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Science And Health table / HealthBench Professional (length-adjusted) / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| HLE-Verified · source release snapshot; version not specified | Hard reasoning | 54.4%Reported settings & sourceGoogle launch chart; benchmark methodology linked on page; reported comparator settings vary. First-party reported result; comparator results retain the source evaluation setup. Gemini 3.8 Flash launch performance · Performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Humanity's Last Exam | Hard reasoning | 56.6% | HLE no tools | lab self-report | 2026-09-01 |
| Humanity's Last Exam | Hard reasoning | 56.3% | HLE no tools | lab self-report | 2026-07-24 |
| Humanity's Last Exam | Hard reasoning | 63.6% | HLE with tools | lab self-report | 2026-09-01 |
| Humanity's Last Exam | Hard reasoning | 64.7% | HLE with tools | lab self-report | 2026-07-24 |
| Humanity’s Last Exam | Hard reasoning | 56.6%Reported settings & sourceFull2500questions; no tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Humanity’s Last Exam | Hard reasoning | 63.6%Reported settings & sourceFull2500questions; with tools; auto thinking;1Mtotal token cap; no compaction; Opus4.6grader; restricted fetch and contamination review for tools. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A;8.12.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Humanity's Last Exam (w/ tools) · not specified | Hard reasoning | 63.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; tools enabled Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / Humanity's Last Exam (w/ tools) / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| JobBench | Supporting evidence | 65.7%Reported settings & source65tasks; mean rubric score; official OpenCode harness and file-aware grader; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LABBench · 2 | Supporting evidence | 84.2%Reported settings & sourceSelfcomputed; Linuxterminal,bioinfotools,Python,R,network; macroaverage11subtasks Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LMArena Text Arena | Human pref | 1493 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| LVBench · static | Multimodal | 75.4%Reported settings & sourceNo tools;1024frames Gemini/GPT,300frames Claude dueAPIlimit; model-specific frame budget: 300 Frame budgets differ; table labels Gemini3.8static explicitly. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MILU | Knowledge | 92.1%Reported settings & sourceMean accuracy11languages; adaptive max; five trials; no tools/custom system prompts. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.18.2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OfficeQA | Agentic | 78.1%Reported settings & sourceExtracted-text Treasury corpus in sandbox; code execution; production Messages API with safeguards/fallback;128koutput cap. Pro133questions. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OfficeQA Pro | Agentic | 66.9%Reported settings & sourceExtracted-text Treasury corpus in sandbox; code execution; production Messages API with safeguards/fallback;128koutput cap. Pro133questions. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.15.1 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Organic Chemistry · 2 revised | Supporting evidence | 65.7%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| OSWorld · 2.0 August2026 task release | Agentic | 75.4%Reported settings & sourcepartial pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero. Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OSWorld · 2.0 August2026 task release | Agentic | 39.6%Reported settings & sourcestrict pass@1;108tasks; five runs;1080p;500steps; max effort; Opus4.8grader; task fixes; Fable safety interventions score zero. Same-condition reruns; supersedes earlier OSWorld2results and incompatible with previous task files. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.14.3 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| OSWorld 2.0 (v2026.08.08, offline set, partial score) · 2.0 / v2026.08.08 offline | Agentic | 70.2%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; offline set; partial credit; v2026.08.08; official task/grading settings Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Computer Use table / OSWorld 2.0 (v2026.08.08, offline set, partial score) / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld binary · 2.0 08.08 | Agentic | 31.4%Reported settings & source108tasks; common internal GUI framework; execution-based checkers; binary metric; reasoning max Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld partial · 2.0 08.08 | Agentic | 68.3%Reported settings & source108tasks; common internal GUI framework; execution-based checkers; partial metric; reasoning max Muse1.2 uses older06.24tasks; all others08.08. Do not compare these versions as identical. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld partial · 2.0 August2026 fixed tasks | Agentic | 75.4%Reported settings & sourcePartialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OSWorld-Verified | Agentic | 83.39% | OSWorld-Verified reportedcontributes to capability | official board | 2026-08-01 |
| ProgramBench · 166 golden-task subset | Supporting evidence | 85.4%Reported settings & sourcemini-swe-agent without six-hour timeout; excludes34flaky-reference tasks; tests restricted to reference-passing tests; up to1Mcontext. Hidden-test pass rate, not fraction of completely solved programs. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.11.1 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Protein Design · Library Ranking | Supporting evidence | 48%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Protein Design · Sequence Generation revised grader | Supporting evidence | 42.4%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| ProteinGym · Hard | Supporting evidence | 47.7 percent rank correlationReported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Protocols · Troubleshooting | Supporting evidence | 61.1%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Protocols · Understanding network-restricted | Supporting evidence | 80%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SingleCellBench | Supporting evidence | 60.6%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SpatialBench · Verified | Supporting evidence | 72.5%Reported settings & sourceAnthropic life-sciences evaluation; bash/editor/packages except protein design no tools; protocols use bash/editor/search; Understanding revised network restriction. Internal evaluation; not necessarily publicly released; compare only same task/grader revision. Percent rank correlation is not accuracy. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.19 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SRE-Bench · 262 binaries / 19 programs | Supporting evidence | 12.5%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API; reduced or absent production safeguards; pass@1; all six objectives required Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Cybersecurity table / SRE-Bench / Claude Opus 5 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-Atlas Codebase QnA | Supporting evidence | 52.7%Reported settings & source124tasks/11repos; publicQnA mini-swe-agent rubric pass@1; GPT/Opus owncards; reasoning max Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Multilingual | Coding | 89.5%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Multimodal | Supporting evidence | 59.4%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Pro | Coding | 79.2% | SWE-bench Pro reported | lab self-report | 2026-07-24 |
| SWE-bench Pro | Coding | 79.2%Reported settings & sourceAnthropic reported configuration; adaptive thinking at max effort, default sampling, five-trial mean unless section specifies otherwise; context at most 1M. Comparator configuration follows cited prior card or board. Agent identity is not established as mini-swe-agent. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table 8.1.A; section 8.2 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| SWE-bench Verified | Coding | 96% | SWE-bench Verified official | lab self-report | 2026-07-24 |
| Terminal-Bench · 2.1 | Coding | 89.1%Reported settings & sourceTerminus2only; Gemini selfcomputed, othermodels officialboard/ArtificialAnalysis Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 2.1 | Coding | 86.7%Reported settings & source89tasks; each model nativecoding harness in internalframework/cloudsandbox; officialverifier; GPT owncard; reasoning max; native harness family: Anthropic (exact harness revision not specified) Native harnesses differ; not a Terminus2-only comparison. Muse Spark1.3 evaluation methodology · PDF page4 performance table · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench · 4.0 | Coding | 52.3%Reported settings & sourceClaude Code --bare max effort, 15 trials/task over 66 tasks; Anthropic internal reruns. Uses more precise section values rather than rounded headline table. Distinct from2.1; timeouts/resources changed. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.6 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Terminal-Bench · 4.0 | Coding | 51.8%Reported settings & sourceOfficialpublicboard highest scoring thinking level; nativeagents may differ Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 4.0 · 4.0 | Coding | 52.6%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Coding table / Terminal-Bench 4.0 / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench Science 0.1 · not specified | Hard reasoning | 30%Reported settings & sourceMaximum reported across reasoning efforts; OpenAI research environment or API Provider-published result; comparator measurements are not automatically independently reproduced. GPT-6 Astra: A new generation of intelligence · Academic table / Terminal-Bench Science 0.1 / Claude Opus 5 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench-Science · 0.1 | Coding | 29%Reported settings & source70tasks; Claude Code --bare max; Fable10trials/task, Opus12; GPT Codex CLI max from public board. Claude Fable 5.1 and Claude Mythos 5.1 System Card · 8.7 · reviewed 2026-09-12 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · Verified June2026 | Agentic | 80.6%Reported settings & sourcePass@1;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · Verified June2026 | Agentic | 87%Reported settings & sourcePass@3;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · Verified June2026 | Agentic | 73.1%Reported settings & sourcePass³;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| Toolathlon · Verified June2026 | Agentic | 23.5 turnsReported settings & sourceaverage turns;108tasks/three trials; internal harness; max effort; Fable safeguards+Opus4.8fallback; Opus safeguards/fallback disabled; pinned containers/data. Claude Fable 5.1 and Claude Mythos 5.1 System Card · Table8.15.5.A · reviewed 2026-09-12 | Published configuration | lab self-report | Reviewed 2026-09-12 |
| VulcanBench v3 | Supporting evidence | 78.3% | VulcanBench v3 bare-bones API · Report 10 · high | official board | 2026-07-26 |
| VulcanBench v3 | Supporting evidence | 87% | VulcanBench v3 bare-bones API · Report 10 · low | official board | 2026-07-26 |
| VulcanBench v3 | Supporting evidence | 82.6% | VulcanBench v3 bare-bones API · Report 10 · medium | official board | 2026-07-26 |
| LiveCodeBench | Coding | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |