RankingClaude Sonnet 4.6
Claude Sonnet 4.6
40 published benchmark measures · 16 benchmark families contribute across 7 task areas. 5 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
Choose a configuration
A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.
Compare measured configurations →Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Compare 7 effort levels across 68 benchmark/harness combinations →
Reported effort · Maximum reasoning + unspecified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Maximum reasoning: 11 observations
- Not specified: 33 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
16 contributing families across 7 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Anthropic
- Catalog status
- active
- Availability
- Public provider catalog; account and region restrictions may apply
- Family
- Claude Sonnet
- Released
- —
- Context
- —
- License
- proprietary
- Model card
- https://platform.claude.com/docs/en/about-claude/model-deprecations
- Default Capability family coverage
/badge/claude-sonnet-4-6.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AdvancedIF rubric-level · source release snapshot; version not specified | Supporting evidence | 86%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| AIME 2025 · source release snapshot; version not specified | Hard reasoning | 95.6%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AIME 2025 avg@16 · source release snapshot; version not specified | Hard reasoning | 86.9%Reported settings & sourceMaximum reasoning; Sonnet4.6 external API truncation caveat. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| AllenAI IFBench · source release snapshot; version not specified | Supporting evidence | 57.1%Reported settings & sourceMaximum reasoning; Sonnet4.6 external API truncation caveat. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| APEX-Agents · source release snapshot; version not specified | Agentic | 26.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| ARC-AGI-2 | Hard reasoning | 60.4% | ARC Prize verifiedcontributes to capability | official board | 2026-02-17 |
| BeyondAIME avg@16 · source release snapshot; version not specified | Supporting evidence | 47.3%Reported settings & sourceMaximum reasoning; Sonnet4.6 external API truncation caveat. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BFCL v3 · source release snapshot; version not specified | Supporting evidence | 76%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · source release snapshot; version not specified | Agentic | 74.7%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| BrowseComp · source release snapshot; version not specified | Agentic | 74.7%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| Claw-Eval · source release snapshot; version not specified | Supporting evidence | 68.3%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Collie · source release snapshot; version not specified | Supporting evidence | 67.7%Reported settings & sourceMaximum reasoning; Sonnet4.6 external API truncation caveat. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image1.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| CorpusQA · source release snapshot; version not specified | Supporting evidence | 79%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| DeepSWE v1.1 | Coding | 29.9% | DeepSWE v1.1 reportedcontributes to capability | official board | 2026-09-03 |
| DRACO · source release snapshot; version not specified | Supporting evidence | 75.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval rubrics · source release snapshot; version not specified | Agentic | 75.7%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GPQA Diamond · source release snapshot; version not specified | Hard reasoning | 89.9%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| GraphWalks <=128k · source release snapshot; version not specified | Long context | 96%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| HealthBench Professional · source release snapshot; version not specified | Supporting evidence | 38%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| IF Bench · source release snapshot; version not specified | Supporting evidence | 50%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| LMArena Text Arena | Human pref | 1473 | LMArena Textcontributes to capability | official board | 2026-09-11 |
| LongBenchV2 · source release snapshot; version not specified | Long context | 66%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MCPAtlas · source release snapshot; version not specified | Agentic | 61.3%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MedXpertQA · source release snapshot; version not specified | Supporting evidence | 49%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| MMLU Pro · source release snapshot; version not specified | Knowledge | 87%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MMMU-Pro · source release snapshot; version not specified | Multimodal | 74.5%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| MultiChallenge · source release snapshot; version not specified | Supporting evidence | 57%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OmniDocBench · source release snapshot; version not specified | Supporting evidence | 86.9%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| OSWorld Verified · source release snapshot; version not specified | Agentic | 72.5%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| OSWorld-Verified | Agentic | 72.11% | OSWorld-Verified reportedcontributes to capability | official board | 2026-03-08 |
| SimpleQA Verified · source release snapshot; version not specified | Knowledge | 29%Reported settings & sourceMicrosoft own evaluation suite; Sonnet max reasoning; Tables 12/19. LongBenchV2 capped 256k,408 questions; CorpusQA GPT-5.4 high judge; 4 runs. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Tables 12 and 19 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SVG-Bench · source release snapshot; version not specified | Supporting evidence | 64.1%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Verified · source release snapshot; version not specified | Coding | 79.6%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Verified · source release snapshot; version not specified | Coding | 79.6%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWE-bench Verified · source release snapshot; version not specified | Coding | 79.6%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configurationcontributes to capability | lab self-report | Reviewed 2026-09-06 |
| SWEAtlas-QnA · source release snapshot; version not specified | Supporting evidence | 31.2%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| SWEAtlas-TestWriting · source release snapshot; version not specified | Supporting evidence | 31.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau3 Airline · source release snapshot; version not specified | Supporting evidence | 83%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau3 Banking · source release snapshot; version not specified | Supporting evidence | 28.4%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau3 Retail · source release snapshot; version not specified | Supporting evidence | 75.9%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| tau3 Telecom · source release snapshot; version not specified | Supporting evidence | 70.4%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| Terminal-Bench 2.0 · 2.0 | Coding | 59.1%Reported settings & sourceMAI avg 4 runs, temperature 1, top-p .97; simple ReAct bash/string-replace; Terminal-Bench ignores predefined timeouts. Comparator configurations from cited reports. First-party reported result; comparator results retain the source evaluation setup. MAI-Thinking-1 technical report · Table 11, PDF page 53 · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| VIBE-V2 · source release snapshot; version not specified | Supporting evidence | 42.8%Reported settings & sourceMiniMax launch figure methodology. SWE internal ClaudeCode 4 runs; Terminal 8CPU16GB2h128k Terminus2; Paper/PostTrain Ralph-loop12h; GDPval internal rubric grader; BrowseComp discard64k; OSWorld361tasks. Comparator provider/leaderboard scores where footnoted. First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| YC-Bench · source release snapshot; version not specified | Supporting evidence | 100000 USDReported settings & sourceLaunch figure; agent final assets First-party reported result; comparator results retain the source evaluation setup. MiniMax M3 model card benchmark figure · figures/benchmark.jpeg · reviewed 2026-09-06 | Published configuration | lab self-report | Reviewed 2026-09-06 |
| GDPval-AA | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |
| MMLU-Pro | Knowledge | — | — | — | — |
| SWE-bench Pro | Agentic | — | — | — | — |
| SWE-bench Verified | Agentic | — | — | — | — |
| Terminal-Bench 2.1 | Agentic | — | — | — | — |