Compare published benchmarks
Use the full performance tables published with model releases. Compare the tests both models have, with category weights you choose.
These are publisher-reported configurations. A shared table does not establish equal inference budgets or an independent reproduction. Win share measures the fraction of weighted test outcomes won, not the size or statistical significance of an advantage.
Select two different models and a published evaluation.
Explore all benchmark coverage · Compare the calibrated benchmark set · Download all observations