RankingAutomationBench-AA

AutomationBench-AA

Data updated 23 Sept 2026

Bucket
Supporting evidence
Unit
percent
Direction
Higher is better
Version
aa-score
Display harness
Artificial Analysis AutomationBench-AA
Board
https://artificialanalysis.ai/evaluations/automationbench-aa

AutomationBench-AA score: share of task objectives completed with no guardrail violation, on Artificial Analysis's private 657-task holdout. This is not Zapier's tasks-completed rate. The effort dataset already includes this evaluation in the agentic family under the existing equal-family budget. These rows are the board, not a second weight.

Compare published benchmark results with category weights →

Models

Results are ordered by score in the display harness, followed by models with no result. Other harnesses are listed separately below. Missing results remain unknown.

276–300 of 455 entries

ModelScoreEvidenceSource-recorded date
MiMo-VL-7B-SFT-2508Xiaomi
MiniMax-M1-40kMiniMax
MiniMax-M1-80kMiniMax
MiniMax-M2MiniMax
MiniMax-M2-herMiniMax
MiniMax-M2.1MiniMax
MiniMax-M2.5MiniMax
MiniMax-Text-01MiniMax
MiniMax-VL-01MiniMax
Ministral 3BMistral
Ministral 8BMistral
Mistral Large 2 (July 2024)Mistral
Mistral Large 2.1Mistral
Mistral Medium 3Mistral
Mistral Nemo 12BMistral
Mistral Small CreativeMistral
Muse Glimmer 30BMeta
Muse Spark 1.1Meta
Muse Spark 1.2Meta
Muse Spark 1.3Meta
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16NVIDIA
North-Micro-Vision-InstructCohere
North-Mini-Code-1.0Cohere
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16NVIDIA
NVIDIA-Nemotron-3-Nano-4B-BF16NVIDIA