Rankingtau3 Airline · source release snapshot; version as labeled
tau3 Airline · source release snapshot; version as labeled
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version as labeled
- Display harness
- Mistral Medium 3.5 model card performance charts
- Board
- https://huggingface.co/mistralai/Mistral-Medium-3.5-128B
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Magistral Medium 1.2Mistral | 53.5%Reported settings & sourceMaximum reasoning; Mistral same-lab comparison; tau3 4trials GPT5.2low user simulator; BrowseComp context-discard setup in card. First-party chart numeric label, visually verified. Mistral Medium 3.5 model card performance charts · images/image2.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Mistral Medium 3.1Mistral | 41.5%Reported settings & sourceMaximum reasoning; Mistral same-lab comparison; tau3 4trials GPT5.2low user simulator; BrowseComp context-discard setup in card. First-party chart numeric label, visually verified. Mistral Medium 3.5 model card performance charts · images/image2.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Mistral Small 4Mistral | 38.5%Reported settings & sourceMaximum reasoning; Mistral same-lab comparison; tau3 4trials GPT5.2low user simulator; BrowseComp context-discard setup in card. First-party chart numeric label, visually verified. Mistral Medium 3.5 model card performance charts · images/image2.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |