RankingMistral Large 4

Mistral Large 4

Compare models

Data updated 8 Oct 2026

20 published benchmark measures · 1 benchmark families contribute across 1 task areas. 1 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Compare with

Choose a second model and its effort to see both profiles and benchmark differences below.

Compare with

Choose a configuration to compare. The overview below combines evidence across settings and has no model-wide rank.

Model evidence summary

This profile combines published settings. It is not a runnable configuration or a leaderboard rank. Compare measured configurations →

Performance profile

Capabilities

Adjusted comparison score · 0–100
Mistral Large 4 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Mistral Large 4 · Agentic: 24.3 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic1 families · 1 with independent evidence · Preliminary24.3

1 core families; 41 direct opponents across 17 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 13/13 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GDPval: 73.2 observed win share
    artificialanalysis.ai · Source 1
Hard reasoningNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 6/6 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    CodingNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 6/6 scenarios unsupported. Without one publisher: No supported estimate; 19/19 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Human prefNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        KnowledgeNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          MultimodalNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            Long contextNo comparable evidenceUnknown

            0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

            Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

            Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

            Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

            Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

              Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

              Reported effort · Not specified

              Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

              Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

              Inspect each result and its source ↓ · Download effort evidence
              Score contributions and missing evidence

              1 contributing families across 1 capabilities. Fixed reference panels do not change when the catalog expands.

              Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

              Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

              Model information & shareable badge
              Lab
              Mistral
              Catalog status
              preview
              Availability
              Public preview API on Mistral Studio. Open weights promised by end of October 2026; weights and license are not yet established.
              Family
              Mistral Large
              Released
              2026-10-06
              Context
              1,000,000 tokens
              API list price
              $1.36 input / $4.18 output per million tokens
              Standard list prices per million tokens; cached input $0.14. Launch promotion is 50% off for two weeks, currently $0.68 input, $0.07 cached input and $2.09 output. Independent evaluators may use list prices. No named API effort scale established by the model card.
              License
              —
              Model card
              https://docs.mistral.ai/models/mistral-large-4-0
              Default Capability family coverage
              Documented-evidence family coverage/badge/mistral-large-4.svg
              Benchmark scores & sources

              Original results, evaluation harnesses, and evidence behind this model.

              BenchmarkBucketScoreHarnessEvidenceSource-recorded date
              AA-BriefcaseSupporting evidence1393
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Long-horizon knowledge work; evaluator result republished by provider. Benchmark version not explicitly stated.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/aa-briefcase-v1_e63X7.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 21b2107139551df7652f35284d95887df95faac0cfe388f9dbd59281797ea3f1.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; AA-Briefcase; chart 13 (mistral-chart-13.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              Artificial Analysis Cyber IndexSupporting evidence50%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Success score; distinct from safety-block fractions.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Secondary evaluator report; index version not explicitly stated. Exact printed chart label; asset https://mistral.ai/_astro/artificial-analysis-cyber-index-v3_Z1f1y2m.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 111f2f4ec7d98940d474fd1f04ec91d59f54e7c7048b27cabd243fd0d5c250f7.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Artificial Analysis Cyber Index; chart 7 (mistral-chart-7.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              AutomationBenchAgentic59.9%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. 657 business workflows; completed without violations; evaluator result republished by provider.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/artificial-analysis---automationbench-v3_ZUf0Vl.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 34714a4abfcf245a905abf66c7995266fc57e9a2c14ec87409cdbfb351d217b6.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; AutomationBench; chart 12 (mistral-chart-12.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              AutomationBench-AASupporting evidence59.90136246808686%Artificial Analysis AutomationBench-AA

              Effort: Not specified

              official board

              Supporting evidence outside the reviewed capability core

              2026-10-08
              ChartQA ProSupporting evidence63.1%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/multimodal-benchmarks---chartqa-pro%201_1sTH6.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 5e324c8384dcd66150a7107b20da9207bb078f55a748f8ec64929d8b86a0e442.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; ChartQA Pro; chart 15 (mistral-chart-15.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              Coding Agent IndexSupporting evidence49.8%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Chart footnote attributes private Artificial Analysis evaluation before public harness launch; provider-republished claim, not a new independent result.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Provider narrative composite; components must not be counted as independent votes alongside this index.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Coding Agent Index · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              CybenchSupporting evidence93%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. 40 security-competition exercises.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/cybersecurity-benchmarks---cybench%201_1dsNEh.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 c2012a17e49a0c39f528645d27de7740cb01597270091890914d3eee017b9721.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Cybench; chart 9 (mistral-chart-9.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              CyberGym-E2ESupporting evidence82%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Artificial Analysis evaluation republished by provider; reproduce real vulnerability and patch it.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/cybersecurity-benchmarks---cybergym-e2e-(aa)%201_Z2dNL7P.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 337eaf7c67568ce49fb8be89ad5491c3cfd933fd8e4edd9302176c0bb5939e76.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; CyberGym-E2E; chart 8 (mistral-chart-8.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              DeepSWE · 1.1Coding61.7%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Chart footnote attributes private Artificial Analysis evaluation before public harness launch; provider-republished claim, not a new independent result.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Narrative states61.7%; chart rounds to62. Preserve narrative precision. Exact printed chart label; asset https://mistral.ai/_astro/artificial-analysis---deepswe-1.1%201-alt_Z28T39q.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 df82a5aa879ff37b505c514b9b99d73580b194af137fabe772aef827c679d2f0.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; DeepSWE; chart 0 (mistral-chart-0.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              Dense200 · bounding box detectionSupporting evidence42%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/multimodal-benchmarks---dense200-(bbox)-alt_2o08QB.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 0346d5977e9995bd389d1c0e7cdaa159d21cc17c3a1ad41309a0135a04aed14b.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Dense200; chart 14 (mistral-chart-14.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              Finance Agent · 2Supporting evidence54.7%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Vals.ai evaluation republished by provider.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/vals.ai---finance-agent-v2%201_18mXuR.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 a9ee93a03e7b2a2181b776e191982fbc30ff32985564135757ea268da9cf7e2d.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Finance Agent; chart 4 (mistral-chart-4.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              Finch (FinWorkBench)Supporting evidence67.4%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Finance/accounting spreadsheet creation and editing.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/finch-(finworkbench)%201_1fqgOj.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 e609c2f070702c14f61ddec61ae413bd2163371e9a5927122a3cff7da58973ca.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Finch (FinWorkBench); chart 19 (mistral-chart-19.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              GDP.pdfSupporting evidence18.6%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Artificial Analysis evaluation republished by provider.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/multimodal-benchmarks---gdp.pdf---aa%201_1TQNQM.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 3f50d25949be36f83391b09b737b0940b43589899cda5a1940fc767963618742.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; GDP.pdf; chart 16 (mistral-chart-16.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              GDPval-AAAgentic1424Artificial Analysis GDPval-AA

              Effort: Not specified

              contributes to capability
              official board2026-10-08
              Harvey’s Legal Agent BenchmarkSupporting evidence15.8%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Vals.ai evaluation republished by provider; version not explicitly stated.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/vals.ai---harvey's-legal-agent-benchmark%201_ZQeTGa.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 b6f8d857fa328c886ab3104ecc0454f0eb8003640e9a3ca68c075efc7494884e.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Harvey’s Legal Agent Benchmark; chart 5 (mistral-chart-5.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              Internal STEM preference vs GLM-5.3Supporting evidence69 weighted win rate percent
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Internal expert preference assessment onmath/physics; source prints weighted win rate69%. Sample size and weighting formula unspecified.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Do not replace weighted win rate with the65% sum of positive-preference categories; comparator is GLM-5.3. Exact printed chart label; asset https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-stem-win-rate-breakdown%201_Z21ewlj.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 80c3f8474bb21a8f7f4eb2e19cff07627a34bc492413e4698f024880227056d5.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Internal STEM preference vs GLM-5.3; chart 18 (mistral-chart-18.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              SciCode-Verified · pass@1 (n=6)Supporting evidence91.8%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Distinct from the generic Artificial Analysis SciCode series; version retained literally from chart title. Exact printed chart label; asset https://mistral.ai/_astro/scicode-verified-pass@1-(n_6)-alt_ZYLgs8.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 5010c6725092f676802d30b0d1181a8c1c873b83b62fcd440f454bf223841c43.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; SciCode-Verified; chart 17 (mistral-chart-17.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              Surge human coding qualitySupporting evidence3.74 rating (1–5)
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Blind professional-annotator assessment on1–5 scale; model identities hidden; SurgeAI evaluation.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Exact printed chart label; asset https://mistral.ai/_astro/human-evaluation---surge-(human-eval-on-code)%201_20jgWu.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 b79d8103a1d51c51c16a6b78ef8000422726c74129dff70e00797d8d04869bfc.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Surge human coding quality; chart 11 (mistral-chart-11.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              SWE-Atlas-QnASupporting evidence59.4%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Chart footnote attributes private Artificial Analysis evaluation before public harness launch; provider-republished claim, not a new independent result.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Narrative states59.4%; chart rounds to59. Exact printed chart label; asset https://mistral.ai/_astro/code-benchmarks---swe-atlas-qna-v3_2wiWq0.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 9f9d32e5c05b99d91047ead47225644319e95ed3ff0e51415c10681bffbb8d56.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; SWE-Atlas-QnA; chart 10 (mistral-chart-10.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              Terminal-Bench · 4.0Coding28.3%
              Reported settings & source

              Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Chart footnote attributes private Artificial Analysis evaluation before public harness launch; provider-republished claim, not a new independent result.

              Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Narrative states28.3%; chart rounds to28. Exact printed chart label; asset https://mistral.ai/_astro/code-benchmarks---terminal-bench-4%201-v3_Z2iGo8O.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 c3201abcd5c42890988e41c33719804fe09ff7ffe90bd0c71553e964a1f4d340.

              Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Terminal-Bench; chart 1 (mistral-chart-1.webp) · reviewed 2026-10-07

              Published configuration

              Effort: Not specified

              lab self-report

              Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

              Reviewed 2026-10-07
              ARC-AGI-2Hard reasoning————
              DeepSWE v1.1Agentic————
              GPQA DiamondHard reasoning————
              Humanity's Last ExamHard reasoning————
              LiveCodeBenchCoding————