RankingAutomationBench-AA
AutomationBench-AA
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- aa-score
- Display harness
- Artificial Analysis AutomationBench-AA
- Board
- https://artificialanalysis.ai/evaluations/automationbench-aa
AutomationBench-AA score: share of task objectives completed with no guardrail violation, on Artificial Analysis's private 657-task holdout. This is not Zapier's tasks-completed rate. The effort dataset already includes this evaluation in the agentic family under the existing equal-family budget. These rows are the board, not a second weight.
Compare published benchmark results with category weights →
Models
176–200 of 455 entries
| Model | Score | Evidence | Source-recorded date |
|---|---|---|---|
| GLM-4.1V-9B-ThinkingZ.ai | — | — | — |
| GLM-4.5Z.ai | — | — | — |
| GLM-4.5-AirZ.ai | — | — | — |
| GLM-4.5-AirXZ.ai | — | — | — |
| GLM-4.5-FlashZ.ai | — | — | — |
| GLM-4.5-XZ.ai | — | — | — |
| GLM-4.5VZ.ai | — | — | — |
| GLM-4.6Z.ai | — | — | — |
| GLM-4.6VZ.ai | — | — | — |
| GLM-4.6V-FlashZ.ai | — | — | — |
| GLM-4.6V-FlashXZ.ai | — | — | — |
| GLM-4.7Z.ai | — | — | — |
| GLM-4.7-FlashZ.ai | — | — | — |
| GLM-4.7-FlashXZ.ai | — | — | — |
| glm-4v-9bZ.ai | — | — | — |
| GLM-5Z.ai | — | — | — |
| GLM-5-TurboZ.ai | — | — | — |
| GLM-5.1Z.ai | — | — | — |
| GLM-5.2Z.ai | — | — | — |
| GLM-5.3Z.ai | — | — | — |
| GLM-5.3-FlashZ.ai | — | — | — |
| GLM-5V-TurboZ.ai | — | — | — |
| glm-edge-1.5b-chatZ.ai | — | — | — |
| glm-edge-4b-chatZ.ai | — | — | — |
| glm-edge-v-2bZ.ai | — | — | — |