RankingAutomationBench-AA

AutomationBench-AA

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
aa-score
Display harness
Artificial Analysis AutomationBench-AA
Board
https://artificialanalysis.ai/evaluations/automationbench-aa

AutomationBench-AA score: share of task objectives completed with no guardrail violation, on Artificial Analysis's private 657-task holdout. This is not Zapier's tasks-completed rate. The effort dataset already includes this evaluation in the agentic family under the existing equal-family budget. These rows are the board, not a second weight.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

351–375 of 455 entries

ModelScoreEvidenceSource-recorded date
Qwen-plus-character-jaQwen
Qwen-turboQwen
Qwen-vl-maxQwen
Qwen-vl-plusQwen
Qwen3-0.6BQwen
Qwen3-1.7BQwen
Qwen3-14BQwen
Qwen3-235B-A22BQwen
Qwen3-235B-A22B-Instruct-2507Qwen
Qwen3-235B-A22B-Thinking-2507Qwen
Qwen3-30B-A3BQwen
Qwen3-30B-A3B-Instruct-2507Qwen
Qwen3-30B-A3B-Thinking-2507Qwen
Qwen3-32BQwen
Qwen3-4BQwen
Qwen3-4B-Instruct-2507Qwen
Qwen3-4B-Thinking-2507Qwen
Qwen3-8BQwen
Qwen3-Coder-30B-A3B-InstructQwen
Qwen3-Coder-480B-A35B-InstructQwen
Qwen3-coder-flashQwen
Qwen3-coder-plusQwen
Qwen3-maxQwen
Qwen3-Next-80B-A3B-InstructQwen
Qwen3-Next-80B-A3B-ThinkingQwen