RankingOSWorld partial · 2.0 task revision unverified gpt-5.6-sol
OSWorld partial · 2.0 task revision unverified gpt-5.6-sol
- Bucket
- Agentic
- Unit
- percent
- Direction
- Higher is better
- Version
- 2.0 task revision unverified gpt-5.6-sol
- Display harness
- Gemini3.8Flash Model Card
- Board
- https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-8-Flash-Model-Card.pdf
The available records have no admitted matched comparison in the capability core. Raw results remain available below.
Compare published benchmark results with category weights →
Models
| Model | Score | Harness | Evidence | Source-recorded date |
|---|---|---|---|---|
| GPT-5.6 SolOpenAI | 62.6%Reported settings & sourcePartialscore; batchtools;1080p/500steps; Gemini/Sonnet bestof3runs; screenshotonly; officialCUAharness Methodology says runs pre08.08patch but Opusvalue fromFable5.1card usesAugustfixedtasks; no controlledsameversionclaim. GPT values providerreports. Comparison limit: Provider-sourced OSWorld task revision is unverified; cannot join a known-version comparison. Gemini3.8Flash Model Card · Model card page5 Results table · reviewed 2026-09-06 | Gemini3.8Flash Model Card | lab self-report | 2026-09-06 |