RankingMistral Large 4
Mistral Large 4
20 published benchmark measures · 1 benchmark families contribute across 1 task areas. 1 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →
Model evidence summary
This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →
Performance profile
Capabilities
Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.
Reported effort · Not specified
Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.
- Not specified: 20 observations
Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.
Inspect each result and its source ↓ · Download effort evidenceScore contributions and missing evidence
1 contributing families across 1 capabilities. Fixed reference panels do not change when the catalog expands.
Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.
Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.
Model information & shareable badge
- Lab
- Mistral
- Catalog status
- preview
- Availability
- Public preview API on Mistral Studio. Open weights promised by end of October 2026; weights and license are not yet established.
- Family
- Mistral Large
- Released
- 2026-10-06
- Context
- 1,000,000 tokens
- API list price
- $1.36 input / $4.18 output per million tokens
- License
- —
- Model card
- https://docs.mistral.ai/models/mistral-large-4-0
- Default Capability family coverage
/badge/mistral-large-4.svg
Benchmark scores & sources
Original results, evaluation harnesses, and evidence behind this model.
| Benchmark | Bucket | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|---|
| AA-Briefcase | Supporting evidence | 1393Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Long-horizon knowledge work; evaluator result republished by provider. Benchmark version not explicitly stated. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/aa-briefcase-v1_e63X7.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 21b2107139551df7652f35284d95887df95faac0cfe388f9dbd59281797ea3f1. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; AA-Briefcase; chart 13 (mistral-chart-13.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Artificial Analysis Cyber Index | Supporting evidence | 50%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Success score; distinct from safety-block fractions. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Secondary evaluator report; index version not explicitly stated. Exact printed chart label; asset https://mistral.ai/_astro/artificial-analysis-cyber-index-v3_Z1f1y2m.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 111f2f4ec7d98940d474fd1f04ec91d59f54e7c7048b27cabd243fd0d5c250f7. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Artificial Analysis Cyber Index; chart 7 (mistral-chart-7.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| AutomationBench | Agentic | 59.9%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. 657 business workflows; completed without violations; evaluator result republished by provider. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/artificial-analysis---automationbench-v3_ZUf0Vl.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 34714a4abfcf245a905abf66c7995266fc57e9a2c14ec87409cdbfb351d217b6. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; AutomationBench; chart 12 (mistral-chart-12.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| AutomationBench-AA | Supporting evidence | 59.90136246808686% | Artificial Analysis AutomationBench-AA | official board | 2026-10-08 |
| ChartQA Pro | Supporting evidence | 63.1%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/multimodal-benchmarks---chartqa-pro%201_1sTH6.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 5e324c8384dcd66150a7107b20da9207bb078f55a748f8ec64929d8b86a0e442. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; ChartQA Pro; chart 15 (mistral-chart-15.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Coding Agent Index | Supporting evidence | 49.8%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Chart footnote attributes private Artificial Analysis evaluation before public harness launch; provider-republished claim, not a new independent result. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Provider narrative composite; components must not be counted as independent votes alongside this index. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Coding Agent Index · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Cybench | Supporting evidence | 93%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. 40 security-competition exercises. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/cybersecurity-benchmarks---cybench%201_1dsNEh.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 c2012a17e49a0c39f528645d27de7740cb01597270091890914d3eee017b9721. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Cybench; chart 9 (mistral-chart-9.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| CyberGym-E2E | Supporting evidence | 82%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Artificial Analysis evaluation republished by provider; reproduce real vulnerability and patch it. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/cybersecurity-benchmarks---cybergym-e2e-(aa)%201_Z2dNL7P.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 337eaf7c67568ce49fb8be89ad5491c3cfd933fd8e4edd9302176c0bb5939e76. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; CyberGym-E2E; chart 8 (mistral-chart-8.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| DeepSWE · 1.1 | Coding | 61.7%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Chart footnote attributes private Artificial Analysis evaluation before public harness launch; provider-republished claim, not a new independent result. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Narrative states61.7%; chart rounds to62. Preserve narrative precision. Exact printed chart label; asset https://mistral.ai/_astro/artificial-analysis---deepswe-1.1%201-alt_Z28T39q.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 df82a5aa879ff37b505c514b9b99d73580b194af137fabe772aef827c679d2f0. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; DeepSWE; chart 0 (mistral-chart-0.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Dense200 · bounding box detection | Supporting evidence | 42%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/multimodal-benchmarks---dense200-(bbox)-alt_2o08QB.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 0346d5977e9995bd389d1c0e7cdaa159d21cc17c3a1ad41309a0135a04aed14b. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Dense200; chart 14 (mistral-chart-14.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Finance Agent · 2 | Supporting evidence | 54.7%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Vals.ai evaluation republished by provider. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/vals.ai---finance-agent-v2%201_18mXuR.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 a9ee93a03e7b2a2181b776e191982fbc30ff32985564135757ea268da9cf7e2d. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Finance Agent; chart 4 (mistral-chart-4.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Finch (FinWorkBench) | Supporting evidence | 67.4%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Finance/accounting spreadsheet creation and editing. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/finch-(finworkbench)%201_1fqgOj.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 e609c2f070702c14f61ddec61ae413bd2163371e9a5927122a3cff7da58973ca. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Finch (FinWorkBench); chart 19 (mistral-chart-19.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| GDP.pdf | Supporting evidence | 18.6%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Artificial Analysis evaluation republished by provider. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/multimodal-benchmarks---gdp.pdf---aa%201_1TQNQM.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 3f50d25949be36f83391b09b737b0940b43589899cda5a1940fc767963618742. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; GDP.pdf; chart 16 (mistral-chart-16.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| GDPval-AA | Agentic | 1424 | Artificial Analysis GDPval-AAcontributes to capability | official board | 2026-10-08 |
| Harvey’s Legal Agent Benchmark | Supporting evidence | 15.8%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Vals.ai evaluation republished by provider; version not explicitly stated. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/vals.ai---harvey's-legal-agent-benchmark%201_ZQeTGa.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 b6f8d857fa328c886ab3104ecc0454f0eb8003640e9a3ca68c075efc7494884e. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Harvey’s Legal Agent Benchmark; chart 5 (mistral-chart-5.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Internal STEM preference vs GLM-5.3 | Supporting evidence | 69 weighted win rate percentReported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Internal expert preference assessment onmath/physics; source prints weighted win rate69%. Sample size and weighting formula unspecified. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Do not replace weighted win rate with the65% sum of positive-preference categories; comparator is GLM-5.3. Exact printed chart label; asset https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-stem-win-rate-breakdown%201_Z21ewlj.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 80c3f8474bb21a8f7f4eb2e19cff07627a34bc492413e4698f024880227056d5. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Internal STEM preference vs GLM-5.3; chart 18 (mistral-chart-18.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| SciCode-Verified · pass@1 (n=6) | Supporting evidence | 91.8%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Distinct from the generic Artificial Analysis SciCode series; version retained literally from chart title. Exact printed chart label; asset https://mistral.ai/_astro/scicode-verified-pass@1-(n_6)-alt_ZYLgs8.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 5010c6725092f676802d30b0d1181a8c1c873b83b62fcd440f454bf223841c43. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; SciCode-Verified; chart 17 (mistral-chart-17.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Surge human coding quality | Supporting evidence | 3.74 rating (1–5)Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Blind professional-annotator assessment on1–5 scale; model identities hidden; SurgeAI evaluation. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/human-evaluation---surge-(human-eval-on-code)%201_20jgWu.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 b79d8103a1d51c51c16a6b78ef8000422726c74129dff70e00797d8d04869bfc. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Surge human coding quality; chart 11 (mistral-chart-11.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| SWE-Atlas-QnA | Supporting evidence | 59.4%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Chart footnote attributes private Artificial Analysis evaluation before public harness launch; provider-republished claim, not a new independent result. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Narrative states59.4%; chart rounds to59. Exact printed chart label; asset https://mistral.ai/_astro/code-benchmarks---swe-atlas-qna-v3_2wiWq0.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 9f9d32e5c05b99d91047ead47225644319e95ed3ff0e51415c10681bffbb8d56. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; SWE-Atlas-QnA; chart 10 (mistral-chart-10.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| Terminal-Bench · 4.0 | Coding | 28.3%Reported settings & sourcePublic Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Chart footnote attributes private Artificial Analysis evaluation before public harness launch; provider-republished claim, not a new independent result. Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Narrative states28.3%; chart rounds to28. Exact printed chart label; asset https://mistral.ai/_astro/code-benchmarks---terminal-bench-4%201-v3_Z2iGo8O.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 c3201abcd5c42890988e41c33719804fe09ff7ffe90bd0c71553e964a1f4d340. Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission. Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Terminal-Bench; chart 1 (mistral-chart-1.webp) · reviewed 2026-10-07 | Published configuration | lab self-report | Reviewed 2026-10-07 |
| ARC-AGI-2 | Hard reasoning | — | — | — | — |
| DeepSWE v1.1 | Agentic | — | — | — | — |
| GPQA Diamond | Hard reasoning | — | — | — | — |
| Humanity's Last Exam | Hard reasoning | — | — | — | — |
| LiveCodeBench | Coding | — | — | — | — |