RankingToolathlon Verified · source release snapshot; version not specified
Toolathlon Verified · source release snapshot; version not specified
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- source release snapshot; version not specified
- Display harness
- zai-org/GLM-5.3
- Board
- https://huggingface.co/zai-org/GLM-5.3
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| Kimi K3Moonshot | 76.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 76.2%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| GPT-5.6 SolOpenAI | 74.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Claude Fable 5Anthropic | 74.7%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. Comparison limit: The source column includes fallback execution without a documented exact effort. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| DeepSeek V4-Pro 0813DeepSeek | 74.1%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| GLM-5.3Z.ai | 73%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Qwen 3.8-MaxQwen | 72.5%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| GLM-5.2Z.ai | 59.9%Reported settings & sourceGLM-5.3 release table; benchmark-specific harness and reasoning configurations in Evaluation Details. First-party reported result. zai-org/GLM-5.3 · Performance table: Toolathlon Verified · reviewed 2026-09-12 | zai-org/GLM-5.3 | lab self-report | 2026-09-12 |
| Claude Opus 4.8Anthropic | 76.2%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| Gemini 3.1 ProGoogle | 61.1%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| GPT-5.5OpenAI | 73.5%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |
| Muse Spark 1.1Meta | 75.6%Reported settings & sourceMeta Model API xhigh; Gemini high; Opus max; GPT xhigh. Best of self-reported or Meta reproduction unless specified. OSWorld GUI only; Terminal bash only5attempts89tasks6CPU8GB; DeepSWE no internet5attempts. First-party reported result; comparator results retain the source evaluation setup. Muse Spark 1.1 evaluation report Figure44 · Figure44 physical page101; methodology pages101–105 · reviewed 2026-09-06 | Muse Spark 1.1 evaluation report Figure44 | lab self-report | 2026-09-06 |