RankingAutomationBench-AA

AutomationBench-AA

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
aa-score
Display harness
Artificial Analysis AutomationBench-AA
Board
https://artificialanalysis.ai/evaluations/automationbench-aa

AutomationBench-AA score: share of task objectives completed with no guardrail violation, on Artificial Analysis's private 657-task holdout. This is not Zapier's tasks-completed rate. The effort dataset already includes this evaluation in the agentic family under the existing equal-family budget. These rows are the board, not a second weight.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

76–100 of 455 entries

ModelScoreEvidenceSource-recorded date
Claude Sonnet 4.5Anthropic
Claude Sonnet 4.6Anthropic
codegeex2-6bZ.ai
codegeex4-all-9bZ.ai
CodeGemma 7B ITGoogle
Codestral 2508Mistral
cogagent-9b-20241220Z.ai
cogagent-chat-hfZ.ai
cogagent-vqa-hfZ.ai
cogvlm-chat-hfZ.ai
cogvlm2-llama3-chat-19BZ.ai
cogvlm2-llama3-chinese-chat-19BZ.ai
cogvlm2-video-llama3-chatZ.ai
CommandCohere
Command A 03 2025Cohere
Command A Plus 05 2026Cohere
Command A Reasoning 08 2025Cohere
Command A Translate 08 2025Cohere
Command A Vision 07 2025Cohere
Command LightCohere
Command R 03 2024Cohere
Command R 08 2024Cohere
Command R Plus 04 2024Cohere
Command R Plus 08 2024Cohere
Command R7B 12 2024Cohere