RankingToolathlon-Verified (official board)

Toolathlon-Verified (official board)

Data updated 6 Oct 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
2026-10-05
Display harness
Toolathlon-Verified Default agent
Board
https://toolathlon.xyz/docs/leaderboard

Toolathlon-Verified is the June 30, 2026 score series. These rows are the official board's headline Pass@1 on the shared Default agent. They do not enter the weighted Overall or capability scores.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

51–75 of 460 entries

ModelScoreEvidenceSource-recorded date
Claude Opus 5.5Anthropic———
Claude Sonnet 4.5Anthropic———
Claude Sonnet 4.6Anthropic———
Claude Sonnet 5.5Anthropic———
codegeex2-6bZ.ai———
codegeex4-all-9bZ.ai———
CodeGemma 7B ITGoogle———
Codestral 2508Mistral———
cogagent-9b-20241220Z.ai———
cogagent-chat-hfZ.ai———
cogagent-vqa-hfZ.ai———
cogvlm-chat-hfZ.ai———
cogvlm2-llama3-chat-19BZ.ai———
cogvlm2-llama3-chinese-chat-19BZ.ai———
cogvlm2-video-llama3-chatZ.ai———
CommandCohere———
Command A 03 2025Cohere———
Command A Plus 05 2026Cohere———
Command A Reasoning 08 2025Cohere———
Command A Translate 08 2025Cohere———
Command A Vision 07 2025Cohere———
Command LightCohere———
Command R 03 2024Cohere———
Command R 08 2024Cohere———
Command R Plus 04 2024Cohere———