RankingMagistral Medium 1.2

Magistral Medium 1.2

Data updated 12 Sept 2026

6 published benchmark measures · 2 benchmark families contribute across 1 task areas. 1 capability estimates available in the full profile. See all results ↓ · Compare published benchmarks →

Model evidence summary

Choose a configuration

A model-wide score would mix effort settings. This page summarizes published evidence; only identified configurations receive leaderboard ranks.

Compare measured configurations →

The capability chart summarizes evidence across reported settings, not a runnable configuration or a leaderboard rank.

Performance profile

Capabilities

Adjusted comparison score · 0–100
Magistral Medium 1.2 capability radarComparison scores from 0 to 100. Unknown capabilities have no point; lines stop at gaps. Hollow points are preliminary. Exact values and support labels follow the chart.AgenticHard reasoningCodingHuman prefKnowledgeMultimodalLong context50100Magistral Medium 1.2 · Agentic: 1.4 · Preliminary

Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.

Agentic2 families · 1 with independent evidence · Preliminary1.4

2 core families; 41 direct opponents across 19 labs. Needs broader benchmark and opponent support

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

  • GDPval: 15.0 observed win share
    artificialanalysis.ai · Source 1
  • Browsecomp: 50.0 observed win share
    Mistral · Source 1
Hard reasoningNo comparable evidenceUnknown

0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

    CodingNo comparable evidenceUnknown

    0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

    Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

    Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 20/20 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

    Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.

    Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

      Human prefNo comparable evidenceUnknown

      0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

      Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

      Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

      Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.

      Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

        KnowledgeNo comparable evidenceUnknown

        0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

        Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

        Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

        Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.

        Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

          MultimodalNo comparable evidenceUnknown

          0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

          Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

          Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

          Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.

          Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

            Long contextNo comparable evidenceUnknown

            0 core families; 0 direct opponents across 0 labs. No admitted core comparisons

            Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.

            Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.

            Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.

            Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.

              Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.

              Compare 1 effort levels across 17 benchmark/harness combinations →

              Reported effort · Maximum reasoning + unspecified

              Settings reported in this model's published benchmark results, including results outside the aggregate. Effort names are provider-specific. These are not API defaults or equal compute budgets.

              Mixed settings means multiple settings occur in the evidence. Best across efforts means the source selected its best reported result across settings; it does not mean Max. Unspecified settings stay unknown. This model-summary chart combines reported settings. The leaderboard keeps identified configurations separate and excludes unknown effort.

              Inspect each result and its source ↓ · Download effort evidence

              Compare capability profiles →

              Score contributions and missing evidence

              2 contributing families across 1 capabilities. Fixed reference panels do not change when the catalog expands.

              Capability is fitted jointly across families. Capability estimates below describe different task areas; their weighted sum is not the Capability score.

              Results without reviewed compatibility or a reference match remain in the raw evidence below. Coverage counts only contributing results.

              Model information & shareable badge
              Lab
              Mistral
              Catalog status
              deprecated
              Availability
              Documented provider API; deprecated and still accessible until retirement
              Family
              Magistral
              Released
              2025-09-18
              Context
              128,000 tokens
              License
              proprietary
              Model card
              https://docs.mistral.ai/models/magistral-medium-1-2-25-09
              Default Capability family coverage
              Documented-evidence family coverage/badge/magistral-medium-2509.svg
              Benchmark scores & sources

              Original results, evaluation harnesses, and evidence behind this model.

              BenchmarkBucketScoreHarnessEvidenceSource-recorded date
              BrowseComp · source release snapshot; version as labeledAgentic10%
              Reported settings & source

              Maximum reasoning; Mistral same-lab comparison; tau3 4trials GPT5.2low user simulator; BrowseComp context-discard setup in card.

              First-party chart numeric label, visually verified.

              Mistral Medium 3.5 model card performance charts · images/image2.png · reviewed 2026-09-06

              Published configuration

              Effort: Maximum reasoning

              contributes to capability
              lab self-reportReviewed 2026-09-06
              GDPval-AAAgentic362Artificial Analysis GDPval-AA

              Effort: Not specified

              contributes to capability
              official board2026-09-12
              tau3 Airline · source release snapshot; version as labeledSupporting evidence53.5%
              Reported settings & source

              Maximum reasoning; Mistral same-lab comparison; tau3 4trials GPT5.2low user simulator; BrowseComp context-discard setup in card.

              First-party chart numeric label, visually verified.

              Mistral Medium 3.5 model card performance charts · images/image2.png · reviewed 2026-09-06

              Published configuration

              Effort: Maximum reasoning

              lab self-report

              Supporting evidence outside the reviewed capability core

              Reviewed 2026-09-06
              tau3 Banking · source release snapshot; version as labeledSupporting evidence7.7%
              Reported settings & source

              Maximum reasoning; Mistral same-lab comparison; tau3 4trials GPT5.2low user simulator; BrowseComp context-discard setup in card.

              First-party chart numeric label, visually verified.

              Mistral Medium 3.5 model card performance charts · images/image2.png · reviewed 2026-09-06

              Published configuration

              Effort: Maximum reasoning

              lab self-report

              Supporting evidence outside the reviewed capability core

              Reviewed 2026-09-06
              tau3 Retail · source release snapshot; version as labeledSupporting evidence70.2%
              Reported settings & source

              Maximum reasoning; Mistral same-lab comparison; tau3 4trials GPT5.2low user simulator; BrowseComp context-discard setup in card.

              First-party chart numeric label, visually verified.

              Mistral Medium 3.5 model card performance charts · images/image2.png · reviewed 2026-09-06

              Published configuration

              Effort: Maximum reasoning

              lab self-report

              Supporting evidence outside the reviewed capability core

              Reviewed 2026-09-06
              tau3 Telecom · source release snapshot; version as labeledSupporting evidence60.5%
              Reported settings & source

              Maximum reasoning; Mistral same-lab comparison; tau3 4trials GPT5.2low user simulator; BrowseComp context-discard setup in card.

              First-party chart numeric label, visually verified.

              Mistral Medium 3.5 model card performance charts · images/image2.png · reviewed 2026-09-06

              Published configuration

              Effort: Maximum reasoning

              lab self-report

              Supporting evidence outside the reviewed capability core

              Reviewed 2026-09-06
              ARC-AGI-2Hard reasoning
              DeepSWE v1.1Agentic
              GPQA DiamondHard reasoning
              Humanity's Last ExamHard reasoning
              LiveCodeBenchCoding
              LMArena Text ArenaHuman pref
              MMLU-ProKnowledge
              OSWorld-VerifiedAgentic
              SWE-bench ProAgentic
              SWE-bench VerifiedAgentic
              Terminal-Bench 2.1Agentic