Filled points are supported; hollow points are preliminary. Lines stop at unknown capabilities. Exact values and sources follow below.
AgenticNo comparable evidenceUnknown
0 core families; 0 direct opponents across 0 labs. No admitted core comparisons
Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.
Refit without one family: No supported estimate; 11/11 scenarios unsupported. Without one publisher: No supported estimate; 15/15 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.
Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage; source sensitivity still requires inspection. Validation audit.
Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.
Hard reasoningNo comparable evidenceUnknown
0 core families; 0 direct opponents across 0 labs. No admitted core comparisons
Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.
Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 16/16 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.
Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.
Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.
CodingNo comparable evidenceUnknown
0 core families; 0 direct opponents across 0 labs. No admitted core comparisons
Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.
Refit without one family: No supported estimate; 7/7 scenarios unsupported. Without one publisher: No supported estimate; 20/20 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.
Audit: family and publisher holdouts improved on neutral predictions within evaluated coverage. Validation audit.
Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.
Human prefNo comparable evidenceUnknown
0 core families; 0 direct opponents across 0 labs. No admitted core comparisons
Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.
Refit without one family: No supported estimate; 1/1 scenarios unsupported. Without one publisher: No supported estimate; 1/1 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.
Audit: cross-family and cross-publisher predictions could not be evaluated from the available comparison network. Validation audit.
Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.
Knowledge1 families · 1 with independent evidence · PreliminaryUnknown
1 core families; 51 direct opponents across 12 labs. No connected comparison to the complete reference panel
Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.
Refit without one family: No supported estimate; 4/4 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.
Audit: almost no publisher-held-out comparisons remained estimable. Cross-source validity is unresolved. Validation audit.
Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.
- MMLU-Pro: 0.0 observed win share
huggingface.co · Source 1
MultimodalNo comparable evidenceUnknown
0 core families; 0 direct opponents across 0 labs. No admitted core comparisons
Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.
Refit without one family: No supported estimate; 8/8 scenarios unsupported. Without one publisher: No supported estimate; 8/8 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.
Audit: limited publisher-held-out coverage. Cross-source validity is unresolved. Validation audit.
Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.
Long contextNo comparable evidenceUnknown
0 core families; 0 direct opponents across 0 labs. No admitted core comparisons
Estimated against the same complete panel: gpt-5.6-sol claude-opus-4-8 gemini-3.1-pro . Indirect comparisons assume performance can be summarized by one strength within this capability.
Refit without one family: No supported estimate; 5/5 scenarios unsupported. Without one publisher: No supported estimate; 5/5 unsupported. Smoothing check: No supported estimate; 3/3 unsupported. These are sensitivity checks, not confidence intervals.
Audit: family and publisher holdouts underperformed neutral predictions. Treat this as an exploratory estimate. Validation audit.
Matched family win shares below describe source evidence; they are not averaged to produce the adjusted score.
Scores estimate outcomes against a shared reference panel; they are not accuracy percentages. Sparse or disconnected evidence cannot qualify an overall profile. Open a capability to inspect its evidence.