RankingMethodology

How the score works

An experimental estimate from comparable public benchmarks. Every ranked entry identifies a model and effort setting.

  1. Match the setup. Checkpoints, efforts, thinking budgets, benchmark versions, task sets and agent harnesses stay separate. Unknown efforts, conflicting values and incomplete runs remain visible but cannot enter the fit. “Best across efforts” is not Max.
  2. Balance the evidence. Each configuration gets an evidence budget, split across independent evaluators, then their families, evaluations and opponents. A comparison uses the smaller budget offered by its two settings. Identical evidence profiles share weight. Provider reports supplement independent evidence, capped at 25% of each configuration's direct weight.
  3. Estimate strength. Bradley–Terry fits matched wins, ties and losses jointly, with 20% neutral smoothing. It uses order, not raw score margins. Missing results are unknown; scores are never copied between efforts.
  4. Require support. A rank needs three families, one reviewed independent evaluator, three opponent models and two independent families with three opponents each. Lab-only evidence cannot qualify. Broad single-source evidence can rank, with its limitation labeled.

Reading the leaderboard

Score: estimated win share against GPT-5.6 Sol High, Claude Opus 4.8 High and Gemini 3.1 Pro High. These fixed references set the display scale. No connected path to the complete panel means no score. The number is not an accuracy percentage.

Best supported: the highest-scoring supported setting per model. All efforts: every measured setting. Unsupported entries remain available in the full evidence view.

Independent corroboration: more than one reviewed independent evaluator contributes. It does not guarantee agreement. Source percentages describe direct comparison weight, not confidence or fractions of the final score. Reviewed leaderboard reprints defer to the original evaluator’s matching configuration.

Open Why this score? for source weights, source/family removal checks and observation exclusions. Sensitivity values appear only when the dated audit matches the current inputs. Family cells are descriptive matched win shares, not separate aggregate ratings.

Factual accuracy uses the share of correct answers. Related benchmark versions share a family budget; comparisons stay within each version and harness. Imported evaluations use the original evaluator’s identity; missing settings and mixed-scorer summaries do not enter the rank.

Freshness: reviewed feeds are scheduled daily. A collection date records retrieval, not a new evaluation. Changed setups and unreviewed source tables wait for review; failed checks retain the last accepted evidence.

Capability profiles

All capability axes use their own fixed panel: Sol High, Opus 4.7 Max and Gemini 3.5 Flash High, measured across the populated areas. Gaps remain unknown; preliminary points stay labeled. Custom profiles combine supported capability estimates. Their weights do not affect Capability.

Limits and validation

Coverage and opponent selection can change the order. Balancing evaluators does not imply equal task-area weight or prove equal reliability. Effort labels do not establish equal compute. Settings can cover different tests: the highest aggregate is not proof that one effort beats another on shared tasks. Small score gaps and sensitivity ranges do not establish statistical significance.

The dated audit predicts 95.9% of held-out family comparisons and 77.8% of held-out source comparisons on average, including empty folds. It retains 12 consistent shared-benchmark reversals. This is reused-data diagnosis, not an untouched final test.

The source-balanced method was frozen before acquiring LiveBench and SWE-rebench results. On comparable predictions, accuracy rose from 68.4% to 69.3%, with lower probability error. This small gain does not establish statistical superiority. The results were newly acquired, not necessarily newly published. Earlier listwise and family-varying candidates remain unpromoted under their separate temporal gates.

A joint source-and-family budget remains unpromoted: it reduced effort-order reversals but worsened held-out probability error and source sensitivity. New independent Vals AI evidence also favored the current formula in a frozen prediction check; acquisition dates alone do not make this future-data validation.

Frozen external validation · Dated configuration audit, alternatives and coverage gaps · Candidate diagnostics and release status

Observation admission decisions · Scores and contributing evidence · Source measurements

Comparing value

Value ranks the strongest supported configurations within your budget. It keeps Capability scores and evidence requirements unchanged. The budget is applied before selecting the best effort per model; lower cost breaks exact, unrounded score ties. With no budget, higher Capability scores lead among priced configurations.

Measured usage: the complete AA Intelligence Index v4.3 workload, priced from its reported tokens, including reasoning, cache treatment and reviewed per-evaluation fallback token fractions at the fallback model’s prices. Missing families and unresolved fallback routing receive no task cost. This is one evaluator’s workload, not a universal cost of quality.

We do not divide the score by price: estimated win share is not task accuracy or a count of successful tasks. Similar scores do not prove equivalent task performance. A cost frontier shows configurations with no cheaper, equal-or-higher point estimate.

Price scenario: fixed uncached input and total billable output tokens at AA’s published rates. It cannot establish effort efficiency. Endpoint, context tier, tool charges and fallback costs are not verified in this estimate. Open weights are not treated as free inference.

Coverage and freshness: exact reviewed configurations only. Missing and over-budget costs do not receive a Value rank. Missing costs stay unknown; pricing older than 30 days is suspended from comparison. Source dates and skipped records are downloadable. The frontier identifies cheaper, equal-or-better point estimates, not statistically proven winners.

Cite the dated data downloads, source policy, ranking mode, configuration view and any custom weights. For Value, also cite the pricing dataset, cost basis, token quantities and budget.