RankingAutomationBench-AA
AutomationBench-AA
- Bucket
- Supporting evidence
- Unit
- percent
- Direction
- Higher is better
- Version
- aa-score
- Display harness
- Artificial Analysis AutomationBench-AA
- Board
- https://artificialanalysis.ai/evaluations/automationbench-aa
AutomationBench-AA score: share of task objectives completed with no guardrail violation, on Artificial Analysis's private 657-task holdout. This is not Zapier's tasks-completed rate. The effort dataset already includes this evaluation in the agentic family under the existing equal-family budget. These rows are the board, not a second weight.
Compare published benchmark results with category weights →
Models
451–455 of 455 entries
| Model | Score | Evidence | Source-recorded date |
|---|---|---|---|
| Voxtral MiniMistral | — | — | — |
| Voxtral SmallMistral | — | — | — |
| WeDLM-7B-InstructTencent | — | — | — |
| WeDLM-8B-InstructTencent | — | — | — |
| Youtu-LLM-2BTencent | — | — | — |