Rankingtau3 Banking · source release snapshot; version not specified
tau3 Banking · source release snapshot; version not specified
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- Mistral Medium 3.5 model card performance charts
- Board
- https://huggingface.co/mistralai/Mistral-Medium-3.5-128B
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Claude Sonnet 4.6Anthropic | 28.4%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Claude Sonnet 4.5Anthropic | 22.4%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| GLM-5.1Z.ai | 16.2%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Mistral Medium 3.5Mistral | 13.4%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |
| Qwen3.5-397B-A17BQwen | 9.8%Reported settings & sourceMaximum reasoning; tau3 4 trials, GPT5.2 low simulator except Sierra reported comparisons; BrowseComp discard-all context at100k. First-party reported result; comparator results retain the source evaluation setup. Mistral Medium 3.5 model card performance charts · images/image4.png · reviewed 2026-09-06 | Mistral Medium 3.5 model card performance charts | lab self-report | 2026-09-06 |