RankingInternal STEM preference vs GLM-5.3

Internal STEM preference vs GLM-5.3

Data updated 8 Oct 2026

Bucket
Supporting evidence
Unit
index
Direction
Higher is better
Version
—
Display harness
Introducing Mistral Large 4 — public preview
Board
https://mistral.ai/news/mistral-large-4/

The available records have no admitted matched comparison in the capability core. Raw results remain available below.

Compare published benchmark results with category weights →

Models

Published configurations retain their source and harness labels. Missing results remain unknown.

1–1 of 1 entries

ModelScoreHarnessEvidenceSource-recorded date
Mistral Large 4Mistral69 weighted win rate percent
Reported settings & source

Public Mistral Large4 Preview. Effort, sampling and evaluation budgets are not stated. Internal expert preference assessment onmath/physics; source prints weighted win rate69%. Sample size and weighting formula unspecified.

Provider-published lab_self_report; retained as launch evidence, not an independent board admission. Do not replace weighted win rate with the65% sum of positive-preference categories; comparator is GLM-5.3. Exact printed chart label; asset https://mistral.ai/_astro/ml4-vs-glm-5.3-%E2%80%94-stem-win-rate-breakdown%201_Z21ewlj.webp?dpl=6ac6a8353c2a68969e300fd2; asset SHA256 80c3f8474bb21a8f7f4eb2e19cff07627a34bc492413e4698f024880227056d5.

Comparison limit: Provider publication; exact benchmark-specific configuration and independent evaluation provenance require separate review before matched-board admission.

Introducing Mistral Large 4 — public preview · Capabilities deep-dive; Internal STEM preference vs GLM-5.3; chart 18 (mistral-chart-18.webp) · reviewed 2026-10-07

Introducing Mistral Large 4 — public previewlab self-report2026-10-07