RankingAutomationBench-AA

AutomationBench-AA

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
aa-score
Display harness
Artificial Analysis AutomationBench-AA
Board
https://artificialanalysis.ai/evaluations/automationbench-aa

AutomationBench-AA score: share of task objectives completed with no guardrail violation, on Artificial Analysis's private 657-task holdout. This is not Zapier's tasks-completed rate. The effort dataset already includes this evaluation in the agentic family under the existing equal-family budget. These rows are the board, not a second weight.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

226–250 of 455 entries

ModelScoreEvidenceSource-recorded date
GPT-5.4 ProOpenAI
GPT-5.5OpenAI
GPT-5.5 ProOpenAI
GPT-5.6 LunaOpenAI
GPT-5.6 TerraOpenAI
Grok-4.20-0309-non-reasoningxAI
Grok-4.20-0309-reasoningxAI
Grok-4.20-multi-agent-0309xAI
Grok-4.3xAI
Grok-4.5xAI
Grok-build-0.1xAI
Hunyuan A52B InstructTencent
Hunyuan-0.5B-InstructTencent
Hunyuan-1.8B-InstructTencent
Hunyuan-4B-InstructTencent
Hunyuan-7B-InstructTencent
Hunyuan-7B-Instruct-0124Tencent
Hunyuan-A13B-InstructTencent
Hy3-previewTencent
Hy4-previewTencent
K EXAONE 2.0 750B A37BLG AI Research
K EXAONE 236B A23BLG AI Research
Kimi-Dev-72BMoonshot
Kimi-K2.6Moonshot
Kimi-K2.7-CodeMoonshot