RankingAutomationBench-AA

AutomationBench-AA

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
aa-score
Display harness
Artificial Analysis AutomationBench-AA
Board
https://artificialanalysis.ai/evaluations/automationbench-aa

AutomationBench-AA score: share of task objectives completed with no guardrail violation, on Artificial Analysis's private 657-task holdout. This is not Zapier's tasks-completed rate. The effort dataset already includes this evaluation in the agentic family under the existing equal-family budget. These rows are the board, not a second weight.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

101–125 of 455 entries

ModelScoreEvidenceSource-recorded date
DeepSeek V4 Flash 0424DeepSeek
DeepSeek-R1-0528DeepSeek
DeepSeek-R1-0528-Qwen3-8BDeepSeek
DeepSeek-R1-Distill-Llama-70BDeepSeek
DeepSeek-R1-Distill-Llama-8BDeepSeek
DeepSeek-R1-Distill-Qwen-1.5BDeepSeek
DeepSeek-R1-Distill-Qwen-14BDeepSeek
DeepSeek-R1-Distill-Qwen-32BDeepSeek
DeepSeek-R1-Distill-Qwen-7BDeepSeek
DeepSeek-V3.2DeepSeek
DeepSeek-V3.2-SpecialeDeepSeek
DeepSeek-V4-Flash-Vision-ExpDeepSeek
Devstral Medium 1.0Mistral
Devstral Small 1.1Mistral
DiffusionGemma 26B A4BGoogle
Doubao Seed CodeByteDance
ERNIE-4.5-0.3B-PTBaidu
ERNIE-4.5-21B-A3B-PTBaidu
ERNIE-4.5-21B-A3B-ThinkingBaidu
ERNIE-4.5-300B-A47B-PTBaidu
ERNIE-4.5-turbo-128kBaidu
ERNIE-4.5-turbo-20260402Baidu
ERNIE-4.5-turbo-32kBaidu
ERNIE-4.5-turbo-vlBaidu
ERNIE-4.5-turbo-vl-32kBaidu