Rankingtau3 Airline · source release snapshot; version not specified
tau3 Airline · source release snapshot; version not specified
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- Mistral Medium 3.5 model card performance charts
- Board
- https://huggingface.co/mistralai/Mistral-Medium-3.5-128B
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Sonnet 4.6Anthropic | 83%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Qwen3.5-397B-A17BQwen | 81.5%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| GLM-5.1Z.ai | 79.5%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Claude Sonnet 4.5Anthropic | 72%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Mistral Medium 3.5Mistral | 72%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |