<?xml version="1.0" encoding="utf-8"?>
<feed xml:lang="en-us" xmlns="http://www.w3.org/2005/Atom"><title>Simon Willison's Weblog: gpt-6-astra</title><link href="http://simonwillison.net/" rel="alternate"/><link href="http://simonwillison.net/tags/gpt-6-astra.atom" rel="self"/><id>http://simonwillison.net/</id><updated>2026-09-04T23:59:05+00:00</updated><author><name>Simon Willison</name></author><entry><title>The Pelican comparison grid for Astra is pretty interesting</title><link href="https://simonwillison.net/2026/Sep/4/astra-pelicans/" rel="alternate"/><published>2026-09-04T23:59:05+00:00</published><updated>2026-09-04T23:59:05+00:00</updated><id>https://simonwillison.net/2026/Sep/4/astra-pelicans/</id><summary type="html">
    &lt;p&gt;I got access to GPT-6 Astra this afternoon, so naturally I used it to generate &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/"&gt;SVGs of pelicans riding bicycles&lt;/a&gt; - at low, medium, high, xhigh and max reasoning levels (Astra doesn't support reasoning=none). Then I rendered those pelicans in &lt;a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6-pelicans.html"&gt;a comparison grid&lt;/a&gt; with GPT-5.6 Sol, Terra, and Luna, and beyond being fun the result was surprisingly useful.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/gpt-6-and-5-6-pelicans-html.webp" alt="Comparison grid showing gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna at 6 different reasoning levels with pelicans and token counts and prices for each one." style="max-width: 100%;" /&gt;
See &lt;a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6-pelicans.html"&gt;the grid&lt;/a&gt; for full quality images. Here's &lt;a href=""&gt;the transcript&lt;/a&gt; that created the GPT-6 Nova pelicans.&lt;/p&gt;
&lt;p&gt;There are a few interesting things that stand out from this grid.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The Astra pelicans are &lt;em&gt;much better&lt;/em&gt;. The very best GPT-5.6-Sol pelican (I liked xhigh better than max) is still pretty clearly a bunch of abstract shapes. Every single one of the Astra pelicans, from low to xhigh, looks better than that. The Astra max one is really good.&lt;/li&gt;
&lt;li&gt;Astra below max still doesn't reliably get the pelican legs on both sides of the frame.&lt;/li&gt;
&lt;li&gt;In terms of cost, Astra may be around twice the price of Sol ($10/million input, $50/million output, compared to $5/$30 for Sol), but it uses significantly less tokens at each of the levels, making the prices at the different levels closer than they might otherwise be.&lt;/li&gt;
&lt;li&gt;Astra low produces a better pelican than ANY of the GPT-5.6 Sol models at any level, for 9.55 cents. Spending 10 cents on any other model gets a much worse result.&lt;/li&gt;
&lt;li&gt;Look at the input token counts: Astra and Luna both used 16 input tokens, Sol and Terra used 26. That's interesting.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I wonder if Astra and Luna are more related to each other than OpenAI let on?&lt;/p&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle"&gt;pelican-riding-a-bicycle&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-6-astra"&gt;gpt-6-astra&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="pelican-riding-a-bicycle"/><category term="gpt-6-astra"/></entry><entry><title>GPT‑6 Astra</title><link href="https://simonwillison.net/2026/Sep/3/gpt6-astra/" rel="alternate"/><published>2026-09-03T20:18:41+00:00</published><updated>2026-09-03T20:18:41+00:00</updated><id>https://simonwillison.net/2026/Sep/3/gpt6-astra/</id><summary type="html">
    
&lt;p&gt;&lt;strong&gt;&lt;a href="https://openai.com/index/gpt-6-astra/"&gt;GPT‑6 Astra&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
GPT-6 Astra is "rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS" - I've not tried it yet myself, so I don't have a great deal to say about it yet.&lt;/p&gt;
&lt;p&gt;It's going to be API priced at the same rate as Claude Fable 5 and 5.1: $10/million input and $50/million output. This is clearly OpenAI's Fable competitor, and appears to score higher than Fable on most of OpenAI's self-reported benchmarks.&lt;/p&gt;
&lt;p&gt;Most impressively, Astra scores 99.9% on the recent (released in March) &lt;a href="https://arcprize.org/arc-agi/3"&gt;ARC-AGI 3 benchmark&lt;/a&gt; - though notably Fable 5 does not yet have a published result, and the &lt;a href="https://arcprize.org/blog/astra"&gt;ARC-AGI blog notes&lt;/a&gt; that the 99.9% score was achieved for $19K using OpenAI's custom "Provider Adapter harness", while the default ARC-AGI harness scored 62.7% for $26K.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Unsurprisingly, given &lt;a href="https://simonwillison.net/tags/openai-hugging-face-incident/"&gt;the recent Hugging Face incident&lt;/a&gt;, Astra is a beast at security tasks. It scores 100% on ExploitBench (GPT-5.6 Sol got 78.5%), 42.4% on ExploitGym (Sol got 30.3%), and 99.2% within four attempts on SRE-Bench binary reverse engineering compared to Sol's 68.7%.&lt;/p&gt;
&lt;p&gt;It's also better at long context: on OpenAI's eight-needle benchmark it got 100% at 256K–512K tokens and 96.3% at 512K–1M tokens. OpenAI may have vanquished one of the ongoing challenges with long context processing.&lt;/p&gt;
&lt;p&gt;It doesn't win at everything though. &lt;a href="https://twitter.com/ArtificialAnlys/status/2095595489031000350"&gt;Artificial Analysis&lt;/a&gt; note that Astra is still beaten by Fable on their Intelligence Index:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Sits beside GPT-5.6 Sol in Intelligence&lt;/strong&gt;: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It did better on their Coding Agent Index:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Leads Coding Agent Index cost efficiency frontier&lt;/strong&gt;: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I'll write more about Astra once I get access to it. The API model label once it rolls out will be &lt;code&gt;gpt-6-astra&lt;/code&gt;.&lt;/p&gt;
&lt;!-- &lt;small&gt;OpenAI's blog keeps throwing 500 errors, but [here's a mirror](https://astratest.codergautam.workers.dev/GPT-6%20Astra_%20A%20new%20generation%20of%20intelligence%20_%20OpenAI) of the post I found [via Hacker News](https://news.ycombinator.com/item?id=49554273#49555070).&lt;/small&gt; --&gt;

    &lt;p&gt;&lt;small&gt;&lt;/small&gt;Via &lt;a href="https://news.ycombinator.com/item?id=49554643"&gt;Hacker News&lt;/a&gt;&lt;/small&gt;&lt;/p&gt;


    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-release"&gt;llm-release&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-6-astra"&gt;gpt-6-astra&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="llm-release"/><category term="gpt-6-astra"/></entry><entry><title>Ten advances in mathematics and theoretical computer science</title><link href="https://simonwillison.net/2026/Aug/1/ten-advances-in-mathematics/" rel="alternate"/><published>2026-08-01T20:34:49+00:00</published><updated>2026-08-01T20:34:49+00:00</updated><id>https://simonwillison.net/2026/Aug/1/ten-advances-in-mathematics/</id><summary type="html">
    
&lt;p&gt;&lt;strong&gt;&lt;a href="https://openai.com/index/ten-advances-in-mathematics/"&gt;Ten advances in mathematics and theoretical computer science&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
A few days ago it was Anthropic &lt;a href="https://simonwillison.net/2026/Jul/28/discovering-cryptographic-weaknesses-with-claude/"&gt;discovering cryptographic weaknesses with Claude&lt;/a&gt; using Mythos Preview, spending $100,000 on tokens and with prompts that included "again we are not looking for low hanging fruit, we want proper research to find genuinly hard findings."&lt;/p&gt;
&lt;p&gt;Now it's OpenAI's turn to flex. They set "an internal version of Astra, our next major model" on finding solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade". They claim to have spent less than $2,000 at GPT-5.6 Sol token prices on each one.&lt;/p&gt;
&lt;p&gt;(No news on how many problems they spent $2,000 on &lt;em&gt;without&lt;/em&gt; reaching a solution though.)&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://github.com/openai/ten-proofs"&gt;openai/ten-proofs&lt;/a&gt; repository has Lean 4 formalizations of their results, and there's also &lt;a href="https://cdn.openai.com/pdf/ten-proofs-oai.pdf"&gt;a paper&lt;/a&gt; describing the solutions and an additional &lt;a href="https://cdn.openai.com/pdf/reasoning-walkthroughs.pdf"&gt;LLM-generated PDF&lt;/a&gt; where the model "reconstructs how the proof came together" based on the unpublished reasoning traces.&lt;/p&gt;
&lt;p&gt;That's a decent level of transparency, but I want to see the prompts they used!&lt;/p&gt;
&lt;p&gt;A lot of mathematicians online are experiencing a collective burst of &lt;a href="https://simonwillison.net/2026/Feb/15/deep-blue/"&gt;Deep Blue&lt;/a&gt;. Mathematician Kirwin Hampshire published an impassioned essay last week, &lt;a href="https://kirwinhampshire.substack.com/p/the-dark-night-of-mathematics"&gt;The Dark Night of Mathematics&lt;/a&gt;, describing "a profound spiritual crisis" brought on by previous (and less significant) results.&lt;/p&gt;
&lt;p&gt;OpenAI's results reminds me of what Terence Tao described as "big mathematics" in &lt;a href="https://spectrum.ieee.org/ai-in-mathematics"&gt;IEEE Spectrum in June&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Unlike some of his peers, Tao is neither dismissive of AI nor fearful. Instead, he sees it as the catalyst for a fundamental shift in the discipline—a transition toward what he calls “big mathematics.” He envisions a future of large-scale, decentralized collaborations between humans and machines, where complex mathematical tasks can be diced and sliced, with humans claiming the creative parts and AI doing the lion’s share of the technical grunt work.&lt;/p&gt;
&lt;/blockquote&gt;

    &lt;p&gt;&lt;small&gt;&lt;/small&gt;Via &lt;a href="https://news.ycombinator.com/item?id=49132058"&gt;Hacker News&lt;/a&gt;&lt;/small&gt;&lt;/p&gt;


    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/mathematics"&gt;mathematics&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/deep-blue"&gt;deep-blue&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-6-astra"&gt;gpt-6-astra&lt;/a&gt;&lt;/p&gt;



</summary><category term="mathematics"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="deep-blue"/><category term="gpt-6-astra"/></entry></feed>