31st January 2025 - Link Blog
o3-mini-system-card.pdf. The o3-mini system card is out - the model itself is likely to be available shortly.
While o3-mini scores higher than o1 and gpt-4o on many of the included benchmarks, especially around coding, it wasn't universally better than them across ever benchmark.
The biggest win was on Codeforces ELO, a competitive programming benchmark where o3-mini scored 2036 against 1841 for o1, 1250 for o1-preview and 900 for GPT-4o. This fits my intuition that inference-scaling models (like R1) are really good at complex code challenges.
Recent articles
- We're going to need default hard budget caps on pretty much everything - 3rd October 2026
- OpenAI DevDay 2026 live blog - 29th September 2026
- 2026 in LLMs (so far) - 27th September 2026