<?xml version="1.0" encoding="utf-8"?>
<feed xml:lang="en-us" xmlns="http://www.w3.org/2005/Atom"><title>Simon Willison's Weblog: Entries</title><link href="http://simonwillison.net/" rel="alternate"/><link href="http://simonwillison.net/atom/entries/" rel="self"/><id>http://simonwillison.net/</id><updated>2026-09-29T15:55:13+00:00</updated><author><name>Simon Willison</name></author><entry><title>OpenAI DevDay 2026 live blog</title><link href="https://simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog/" rel="alternate"/><published>2026-09-29T15:55:13+00:00</published><updated>2026-09-29T15:55:13+00:00</updated><id>https://simonwillison.net/2026/Sep/29/openai-devday-2026-live-blog/</id><summary type="html">&lt;p&gt;I'm at &lt;a href="https://devday.openai.com/"&gt;OpenAI DevDay&lt;/a&gt; today, in Fort Mason, San Francisco. Same as &lt;a href="https://simonwillison.net/2025/Oct/6/openai-devday-live-blog/"&gt;last year&lt;/a&gt; I'll be live blogging the keynote and some other notes during the day.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;OpenAI gave me a free ticket and a seat in the "creator" area for the keynote.&lt;/em&gt;&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="coding-agents"/><category term="live-blog"/><category term="openai-devday"/></entry><entry><title>2026 in LLMs (so far)</title><link href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/" rel="alternate"/><published>2026-09-27T23:54:15+00:00</published><updated>2026-09-27T23:54:15+00:00</updated><id>https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/</id><summary type="html">&lt;p&gt;On Friday I gave the closing keynote at the &lt;a href="https://www.wearedevelopers.com/world-congress-north-america"&gt;WeAreDevelopers World Congress North America&lt;/a&gt; in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video &lt;a href="https://www.youtube.com/watch?v=GAkIytR7vcc"&gt;is on YouTube&lt;/a&gt;; here are my annotated slides and notes to accompany the talk.&lt;/p&gt;

&lt;p&gt;&lt;lite-youtube videoid="GAkIytR7vcc" js-api="js-api"
  title="WWC26-NA - 2026 in LLMs (so far)"
  playlabel="Play: WWC26-NA - 2026 in LLMs (so far)"
&gt; &lt;/lite-youtube&gt;&lt;/p&gt;

&lt;p&gt;And as an &lt;a href="https://simonwillison.net/tags/annotated-talks/"&gt;annotated presentation&lt;/a&gt;:&lt;/p&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.001.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.001.webp" alt="2026 in LLMs (so far)
Simon Willison
WeAreDevelopers World Congress North America, 25th September 2026" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.001.webp"&gt;#&lt;/a&gt;
&lt;p&gt;I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.002.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.002.webp" alt="November 2025
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.002.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;For me, 2026 started a couple of months earlier in November 2025.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.003.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.003.webp" alt="The November 2025 inflection point
Claude Opus 4.5 GPT-5.1
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.003.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;November saw the release of two important models: Claude Opus 4.5 and GPT-5.1.&lt;/p&gt;
&lt;p&gt;As is usually the case with new models, these were incremental improvements on the models that came before them.&lt;/p&gt;
&lt;p&gt;But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working.&lt;/p&gt;
&lt;p&gt;In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger.&lt;/p&gt;
&lt;p&gt;These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis".&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.004.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.004.webp" alt="&amp;quot;Generate an SVG of a pelican riding a bicycle&amp;quot;. The Claude Opus 4.5 one has a very weird shaped frame and the pelican looks like a duck. The GPT-5.1 has a slightly better but still broken bicycle frame and a slightly better pelican beak, but both are pretty terrible." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.004.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it.&lt;/p&gt;
&lt;p&gt;But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can't ride bicycles in the first place.&lt;/p&gt;
&lt;p&gt;Here's the state of the art for November. Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.005.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.005.webp" alt="November 24th 2025 - the first commit to steipete/Warelay. A GitHub commit adding an MIT license file." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.005.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also in November, we had the first commit to an obscure GitHub repository called "Warelay". We'll come back to this repository shortly.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.006.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.006.webp" alt="January
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.006.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;And then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn't do before.&lt;/p&gt;
&lt;p&gt;Come January, a lot of us were quite excited to start putting this stuff into action.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.007.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.007.webp" alt="New year’s resolution for 2026

Every previous year:
Take on less new projects,
focus on the most important
things in my existing projects" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.007.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Every year I set myself a New Year's resolution, and for as long as I can remember it's been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.008.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.008.webp" alt="2026: Be more ambitious. Take on as many new projects as I want." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.008.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;This year I decided that since that had never worked before, I'd go the other way.&lt;/p&gt;
&lt;p&gt;We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like!&lt;/p&gt;
&lt;p&gt;(You can ask me at the end of the year if this turned out to be a good idea or not. I have a &lt;em&gt;lot&lt;/em&gt; of plates spinning right now.)&lt;/p&gt;
&lt;p&gt;"Be more ambitious" has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don't work.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.009.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.009.webp" alt="Predictions for 2026

It will become undeniable that LLMs write good code
We&amp;#39;re finally going to solve sandboxing
A “Challenger disaster” for coding agent security
Kakapo parrots will have an outstanding breeding season
(only 236 in the world!)

... the Pope will weigh in on LLMs and
their economic impact on the world" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.009.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;I also went on &lt;a href="https://simonwillison.net/2026/Jan/8/llm-predictions-for-2026/"&gt;the Oxide and friends podcast&lt;/a&gt; with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years).&lt;/p&gt;
&lt;p&gt;With hindsight, my LLM predictions were pretty unambitious. &lt;/p&gt;
&lt;p&gt;I said "it will become undeniable that LLMs write good code" - I think we're there now.&lt;/p&gt;
&lt;p&gt;I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions &lt;a href="https://www.wearedevelopers.com/world-congress-north-america/agenda/schedule"&gt;at this conference&lt;/a&gt; touched on sandboxing or agent security in some way, so we're at least putting a lot of effort into that!&lt;/p&gt;
&lt;p&gt;I predicted "a Challenger disaster" for coding agent security. There's certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn't really played out.&lt;/p&gt;
&lt;p&gt;We threw in &lt;a href="https://simonwillison.net/2026/May/25/encyclical-on-ai/#another-2026-prediction-down"&gt;a joke prediction&lt;/a&gt; that the Pope would weigh in on the economic impact of LLMs.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.010.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.010.webp" alt="A photograph of a beautiful green New Zealand parrot. Photo credit Kimberley Collins." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.010.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;I also predicted that New Zealand's Kākāpō parrots would have an outstanding breeding season this year.&lt;/p&gt;
&lt;p&gt;These are flightless nocturnal parrots. They're kind of dumpy looking, I think they're beautiful, and there were only 236 of these parrots in the world at the start of the year.&lt;/p&gt;
&lt;p&gt;Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn't happened in four years... but this year the Rimu fruit were looking excellent.&lt;/p&gt;
&lt;p&gt;Photo &lt;a href="https://commons.wikimedia.org/wiki/File:K%C4%81k%C4%81p%C5%8D_at_Dunedin_Wildlife_Hospital.jpg"&gt;by Kimberley Collins&lt;/a&gt;.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.011.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.011.webp" alt="Deep Blue
Coined by Adam Leventhal and Bryan Cantrill
That feeling of AI induced ennui where software
engineers get listless because the AI can do anything
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.011.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also on that podcast, we coined a term (full credit to Adam) for "that feeling of AI induced ennui where software engineers get listless because the AI can do anything".&lt;/p&gt;
&lt;p&gt;We called it &lt;a href="https://simonwillison.net/2026/Feb/15/deep-blue/"&gt;Deep Blue&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This has been a major theme throughout the year, and was touched on by several speakers at this conference.&lt;/p&gt;
&lt;p&gt;As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically.&lt;/p&gt;
&lt;p&gt;A lot of what I've been doing this year is trying to come to terms with that and what that means for my own profession.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.012.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.012.webp" alt="AI mania

Screenshots of the micro-javascript and pwasm GitHub README files." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.012.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also in January, I suffered from what I'm calling &lt;strong&gt;AI mania&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;This is not the same thing as &lt;a href="https://en.wikipedia.org/wiki/AI-induced_psychosis"&gt;AI psychosis&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;With AI mania, any time your agent isn't building something for you feels like wasted time. You're losing sleep because you could be staying up later getting your agents to do stuff.&lt;/p&gt;
&lt;p&gt;My AI mania presented itself in some ridiculously over-ambitious projects.&lt;/p&gt;
&lt;p&gt;I built &lt;a href="https://github.com/simonw/micro-javascript"&gt;a JavaScript interpreter entirely in Python&lt;/a&gt;, vibe-ported from &lt;a href="https://github.com/bellard/mquickjs"&gt;MicroQuickJS&lt;/a&gt; by Fabrice Bellard.&lt;/p&gt;
&lt;p&gt;Then I built &lt;a href="https://github.com/simonw/pwasm"&gt;a WebAssembly runtime in Python as well&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask "does the world need a slow, buggy, half-baked Python JavaScript interpreter?"&lt;/p&gt;
&lt;p&gt;I don't think the world does.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.013.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.013.webp" alt="micro-javascript playground 3

Execute JavaScript code in a sandboxed micro-javascript environment powered by Pyodide

A web UI with some JavaScript code, and a &amp;quot;Run Code&amp;quot; button, and an output panel.
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.013.webp"&gt;#&lt;/a&gt;
  &lt;p&gt; I did get this out of it: &lt;a href="https://simonw.github.io/micro-javascript/playground.html"&gt;https://simonw.github.io/micro-javascript/playground.html&lt;/a&gt;&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.014.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.014.webp" alt="Previous screenshot, with this text overlaid:

JavaScript running in Python running in Pyodide running in WebAssembly running in JavaScript" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.014.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;This page runs my JavaScript interpreter built in Python, running in Python using &lt;a href="https://pyodide.org/"&gt;Pyodide&lt;/a&gt;, which is Python compiled to WebAssembly, running in JavaScript, running in a browser.&lt;/p&gt;
&lt;p&gt;It's a beautiful stack of horrors. I've been having &lt;a href="https://simonwillison.net/tags/webassembly/"&gt;a lot of fun with WebAssembly&lt;/a&gt; this year.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.015.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.015.webp" alt="Warelay → CLAWDIS → CLAWDBOT →
Clawdbot → Moltbot →🦞 OpenClaw

Screenshot of the dates that these changes happened." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.015.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;By the end of January, that repository we saw start in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.016.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.016.webp" alt="Same screenshot, an overlay reads:

8,330 commits in just
under two months
(it’s at 100,141 today)" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.016.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it's &lt;a href="https://github.com/openclaw/openclaw"&gt;over 100,000 commits&lt;/a&gt; now!&lt;/p&gt;
&lt;p&gt;This is the most vibe-coded piece of software in existence.&lt;/p&gt;
&lt;p&gt;(Here's &lt;a href="https://simonwillison.net/2026/May/16/openclaw-names/"&gt;how I generated that list of name changes&lt;/a&gt;.)&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.017.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.017.webp" alt="Generic term: Claw
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.017.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;This kicked off the OpenClaw revolution. It effectively defined a new category of software.&lt;/p&gt;
&lt;p&gt;There's a generic term for this which I really enjoy. We call software like this a "Claw". There's OpenClaw, &lt;a href="https://github.com/nanocoai/nanoclaw"&gt;NanoClaw&lt;/a&gt;, &lt;a href="https://github.com/nearai/ironclaw"&gt;IronClaw&lt;/a&gt;, &lt;a href="https://github.com/sipeed/picoclaw"&gt;PicoClaw&lt;/a&gt;...&lt;/p&gt;
&lt;p&gt;Today they're being rebranded as "personal agents" or "general agents", but I still like to think of them as Claws.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.018.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.018.webp" alt="Photo of a Mac mini

An aquarium for your Claw
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.018.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw!&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.dbreunig.com"&gt;Drew Breunig&lt;/a&gt; said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.019.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.019.webp" alt="Screenshot of Moltbook - a social network for AI agents" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.019.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also in January, we had this website.&lt;/p&gt;
&lt;p&gt;This was &lt;a href="https://www.moltbook.com/"&gt;MoltBook&lt;/a&gt;, a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that?&lt;/p&gt;
&lt;p&gt;The website launched on Thursday. It &lt;a href="https://simonwillison.net/2026/Jan/30/moltbook/"&gt;blew up on Friday&lt;/a&gt;. It was &lt;a href="https://www.nytimes.com/2026/02/02/technology/moltbook-ai-social-media.html"&gt;profiled by the New York Times on Monday&lt;/a&gt;. And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam.&lt;/p&gt;
&lt;p&gt;Facebook/Meta &lt;a href="https://www.cnbc.com/2026/03/10/meta-social-networks-ai-agents-moltbook-acquisition.html"&gt;bought it a month later&lt;/a&gt;.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.020.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.020.webp" alt="February
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.020.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;In February, a company called StrongDM described what they called their Software Factory.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.021.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.021.webp" alt="StrongDM’s Dark Factory
Justin McCarthy, Jay Taylor, Navan Chauhan

Software Factories and the Agentic Moment" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.021.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;They wrote about this in &lt;a href="https://factory.strongdm.ai"&gt;Software Factories and the Agentic Moment&lt;/a&gt;. I &lt;a href="https://simonwillison.net/2026/Feb/7/software-factory/"&gt;posted my own notes&lt;/a&gt; at the time, having seen their demo in person back in October.&lt;/p&gt;
&lt;p&gt;Dan Shapiro called this approach &lt;a href="https://www.danshapiro.com/blog/2026/01/the-five-levels-from-spicy-autocomplete-to-the-software-factory/"&gt;the Dark Factory&lt;/a&gt;, after the idea that if your factory is sufficiently automated you can turn the lights out, because you don't even need to see what's going on.&lt;/p&gt;
&lt;p&gt;StrongDM presented two rules for software development that they'd been following since July last year.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.022.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.022.webp" alt="“Rule 1: Code must not be written by humans”" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.022.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;The first was code &lt;strong&gt;must not be written by humans&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Any code that you write has to have been routed through a coding agent.&lt;/p&gt;
&lt;p&gt;This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.023.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.023.webp" alt="“Rule 2: Code must not be reviewed by humans” (!)
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.023.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Rule number two was code must &lt;strong&gt;not be reviewed by humans&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;You're not allowed to read the code!&lt;/p&gt;
&lt;p&gt;This continued to be a huge topic for much of this year. Many of the sessions at this event have been about code review and how you can get away with this.&lt;/p&gt;
&lt;p&gt;What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they'd been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work?&lt;/p&gt;
&lt;p&gt;StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what's possible and responsible to do with this stuff.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.024.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.024.webp" alt="Headline on New Zealand&amp;#39;s Department of Conservation website:

First kakapo chick in four years hatches on Valentine&amp;#39;s Day. It&amp;#39;s a grey fluffy ball." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.024.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also in February: &lt;a href="https://www.doc.govt.nz/news/media-releases/2026-media-releases/first-kakapo-chick-in-four-years-hatches-on-valentines-day/"&gt;First kākāpō chick in four years hatches on Valentine's Day&lt;/a&gt;. Breeding season is off to a good start!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.025.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.025.webp" alt="19th February 2026
Gemini 3.1 Pro

A surprisingly good illustration of a pelican riding a bicycle." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.025.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also in February... Google released &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/"&gt;Gemini 3.1 Pro&lt;/a&gt;. That's a pretty great pelican riding a bicycle! It's got the chain in the right place, it's got feet on both sides. There's a little fish in the basket.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.026.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.026.webp" alt="@JeffDean on Twitter - a video comparing Gemini 3 Pro and Gemini 3.1 Pro." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.026.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;And then Google's Jeff Dean &lt;a href="https://x.com/JeffDean/status/2024525132266688757"&gt;tweeted a video&lt;/a&gt; comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine.&lt;/p&gt;
&lt;p&gt;This was frustrating, because my protection for the pelican riding the bicycle test was always "if they draw a perfect pelican on a bicycle, I'll ask for some other animal on something else."&lt;/p&gt;
&lt;p&gt;Google trained for all forms of animals on all forms of transport! They've defeated my benchmark at this point.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.027.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.027.webp" alt="Three headlines:

Meta Makes AI Adoption a Formal
Part of Performance Reviews

Not just engineers writing code, Microsoft
wants almost every employee to use Al

Dara Khosrowshahi: 90% of Uber engineers now
use AI in daily workflows
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.027.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;The other thing that started in February was &lt;strong&gt;Tokenmaxxing&lt;/strong&gt;. We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that 90% of their engineers were using AI workflows.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.028.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.028.webp" alt="More headlines: 

Meta Plans to Crack Down on Employee Token Use: Information

Microsoft Tells Engineers: Tokenmaxxing is not what we are optimizing for

Uber caps employee AI spending after blowing through budget in four months" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.028.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Then a few months later we have Meta cracking down on token use, Microsoft saying tokenmaxxing is "not what we are optimizing for", and Uber capping employee AI spending. &lt;/p&gt;
&lt;p&gt;So tokenmaxxing went straight up and then straight back down again - because it turns out the agents are &lt;em&gt;expensive&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Last year it was difficult to spend more than $50 on AI tokens, because we didn't have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work.&lt;/p&gt;
&lt;p&gt;This is also the reason that Anthropic's valuation skyrocketed to maybe a trillion dollars.&lt;/p&gt;
&lt;p&gt;AI appears to &lt;a href="https://simonwillison.net/2026/May/27/product-market-fit/"&gt;have hit product market fit&lt;/a&gt; in 2026, primarily through coding agents.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.029.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.029.webp" alt="March
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.029.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;In March, we hit peak OpenClaw.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.030.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.030.webp" alt="March: peak OpenClaw

Photos of people in china queuing up to install OpenClaw, with big fluffy lobsters." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.030.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices.&lt;/p&gt;
&lt;p&gt;I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf.&lt;/p&gt;
&lt;p&gt;A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way - writing and then executing code on your computer to get stuff done.&lt;/p&gt;
&lt;p&gt;The race was on to be the first to build a &lt;strong&gt;safe Claw&lt;/strong&gt; - a Claw you could give to regular human beings where they wouldn't instantly shoot themselves in the foot.&lt;/p&gt;
&lt;p&gt;Meta's Muse &lt;a href="https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/"&gt;came out three weeks ago&lt;/a&gt; and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers.&lt;/p&gt;
&lt;p&gt;I'm not yet convinced you &lt;em&gt;can't&lt;/em&gt; shoot yourself in the foot with Muse, but I guess we'll find out for sure pretty soon.&lt;/p&gt;
&lt;p&gt;Photos from &lt;a href="https://www.thewirechina.com/2026/03/29/how-the-openclaw-frenzy-is-testing-chinas-ai-commitment/"&gt;How the OpenClaw Frenzy Is Testing China’s AI Commitment&lt;/a&gt; (March 29th) and &lt;a href="https://www.sixthtone.com/news/1018393"&gt;The Enthusiasm and Anxiety Behind China’s OpenClaw Craze&lt;/a&gt; (April 8th, 2026).&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.031.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.031.webp" alt="April
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.031.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;In April, we had a model release where the model wasn't actually released.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.032.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.032.webp" alt="Simon Willison’s Weblog - screenshot of the post &amp;quot;Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me&amp;quot; from April 7th 2026" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.032.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Anthropic announced their new Claude Mythos model, and then said it was &lt;em&gt;too dangerous&lt;/em&gt; to release beyond a trusted group of security researchers.&lt;/p&gt;
&lt;p&gt;Mythos was really, really good at hacking things.&lt;/p&gt;
&lt;p&gt;The "it's too dangerous" marketing ploy has been played by AI companies dating all the way back to &lt;a href="https://en.wikipedia.org/wiki/GPT-2"&gt;GPT-2&lt;/a&gt;. Anytime an AI company says we've built something that's "too dangerous", it's natural to be a bit skeptical.&lt;/p&gt;
&lt;p&gt;I found the Mythos claims credible, because I'd seen how good coding agents had got at finding regular bugs. I wrote about that in &lt;a href="https://simonwillison.net/2026/Apr/7/project-glasswing/"&gt;Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me&lt;/a&gt;. &lt;/p&gt;
&lt;p&gt;With hindsight... yeah, the models had got really good at finding vulnerabilities!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.033.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.033.webp" alt="16th April 2026
Qwen3.6-35B-A3B and Opus 4.7

Qwen&amp;#39;s pelican has a correct bicycle frame and a good beak. Opus 4.7&amp;#39;s bicycle frame is still junk.

Qwen3.6-35B-A3B is a 20.9GB file that runs on my laptop
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.033.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop.&lt;/p&gt;
&lt;p&gt;On the 16th of April &lt;a href="https://simonwillison.net/2026/Apr/16/qwen-beats-opus/"&gt;I ran the new Qwen3.6-35B-A3B&lt;/a&gt; on my laptop, and it drew me a better pelican riding a bicycle than Anthropic's brand new Claude Opus 4.7 did!&lt;/p&gt;
&lt;p&gt;Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too!&lt;/p&gt;
&lt;p&gt;That's from a 21GB file running on my laptop.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.034.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.034.webp" alt="Now a flamingo on a unicycle. The Qwen one is visibly better than the Opus 4.7 one - the Qwen one is wearing sunglasses and looks a bit like it&amp;#39;s smoking a cigarette." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.034.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flamingo riding a unicycle as well. Again, it handily beat Claude Opus 4.7.&lt;/p&gt;
&lt;p&gt;The local model releases this year have been absolutely extraordinary.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.035.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.035.webp" alt="May
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.035.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;In May... the Pope got involved.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.036.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.036.webp" alt="25th May 2026
The HOLY SEE

ENCYCLICAL LETTER
MAGNIFICA HUMANITAS
OF HIS HOLINESS
POPE LEO XIV
ON SAFEGUARDING THE HUMAN PERSON
IN THE TIME OF ARTIFICIAL INTELLIGENCE" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.036.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;In our podcast episode back in January we'd predicted that the Pope would say something about AI.&lt;/p&gt;
&lt;p&gt;In May, Pope Leo XIV released an encyclical letter on "safeguarding the human person in the time of artificial intelligence".&lt;/p&gt;
&lt;p&gt;Here are &lt;a href="https://simonwillison.net/2026/May/25/encyclical-on-ai/"&gt;my notes on that document&lt;/a&gt;.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.037.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.037.webp" alt="Wikipedia article on Rerum novarum

Rerum novarum is an encyclical issued by Pope Leo
XIII 15 on May 1891." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.037.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;With hindsight, this shouldn't have been a surprise at all.&lt;/p&gt;
&lt;p&gt;Our current Pope's name is Leo XIV, because when he named himself he chose his papal name after Leo XIII - the Pope who wrote an encyclical about the Industrial Revolution back in 1891.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Rerum_novarum"&gt;Rerum novarum&lt;/a&gt; was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week.&lt;/p&gt;
&lt;p&gt;When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way.&lt;/p&gt;
&lt;p&gt;Our joke podcast prediction was junk, because this was always going to happen.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.038.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.038.webp" alt="Corey Quinn @QuinnyPig on Twitter
I cannot believe I&amp;#39;m saying this, but getting the literal Pope to canonize your product&amp;#39;s specific technical limitations as a spiritual treatise is the
single greatest act of vendor lobbying I have ever seen.

May 25" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.038.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;One of Anthropic's co-founders, Christopher Olah, was present for the Pope's event announcing the new encyclical.&lt;/p&gt;
&lt;p&gt;Corey Quinn &lt;a href="https://twitter.com/quinnypig/status/2058960462256210268"&gt;noted&lt;/a&gt; that:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen.&lt;/p&gt;
&lt;/blockquote&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.039.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.039.webp" alt="@maciejmensfeld

We&amp;#39;re dealing with a major malicious attack on right now.
Signups are paused for the time being.

Hundreds of packages involved - mostly targeting us, but some carrying
exploits. The team has been on this for hours. More details to follow
once we&amp;#39;re through it.

4:39 AM - May 12, 2026 - 687.6K Views
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.039.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to &lt;a href="https://twitter.com/maciejmensfeld/status/2054164602577940619"&gt;shut down user registrations&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Let's take that one and put it on a pile of mysteries to figure out later.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.040.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.040.webp" alt="June
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.040.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;In June... Claude Fable 5 came out!&lt;/p&gt;
&lt;p&gt;We got a version of Mythos that has been neutered, so that it wouldn't help us hack into systems or build biological weapons.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.041.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.041.webp" alt="9th June 2026: Claude Fable 5

Five pelicans riding bicycles, from low to max thinking levels. The xhigh one looks particularly good." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.041.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Fable was pretty good at drawing pelicans on bicycles!&lt;/p&gt;
&lt;p&gt;The frames are a good shape, the pelicans look like pelicans. The legs are often incorrectly on the same side of the bicycle, but generally these are pretty great compared to what came before.&lt;/p&gt;
&lt;p&gt;They were pretty expensive - 30 cents and 72 cents for the best ones.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.042.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.042.webp" alt="Fable class models
If you can define a goal,
provide unambiguous instructions,
and provide access to necessary tools
They can solve your
problem with brute force" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.042.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Most importantly though, this was our first public glimpse of what I think of as a &lt;strong&gt;Fable class model&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Today we have more of these, such as GPT-6 Astra.&lt;/p&gt;
&lt;p&gt;These are models where if you can &lt;strong&gt;clearly define the goal&lt;/strong&gt; for what you want to build, and provide &lt;strong&gt;unambiguous instructions&lt;/strong&gt; about the constraints around that goal, and give the model &lt;strong&gt;access to the necessary tools&lt;/strong&gt; to achieve that goal... they will solve your problem effectively through brute force.&lt;/p&gt;
&lt;p&gt;On the one hand, this looks like a direct threat to us software engineers - because it means that the models can build effectively any piece of software you can define in this way.&lt;/p&gt;
&lt;p&gt;Look a bit closer though and you'll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering &lt;em&gt;is&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;It takes a lot of experience and skill to do this well. If you &lt;em&gt;can&lt;/em&gt; do it well, you've now got superpowers.&lt;/p&gt;
&lt;p&gt;This helped me a little bit with my Deep Blue feelings: the realization that there's still a lot of skill to be had in driving models that get this good.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.043.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.043.webp" alt="A new form of AI mania...
Fable is available on subscription
plans “until June 22nd”" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.043.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd.&lt;/p&gt;
&lt;p&gt;That gave us less than two weeks of Fable access before the price went up.&lt;/p&gt;
&lt;p&gt;I was losing sleep again. I was rescheduling things so that I'd have more time with Fable. I was all-in to get as much as I could out of this model.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.044.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.044.webp" alt="12th June 2026: no more Claude Fable 5

Anthropic website:

Statement on the US government directive
to suspend access to Fable 5 and Mythos 5
Jun 12, 2026" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.044.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;And then &lt;a href="https://www.anthropic.com/news/fable-mythos-access"&gt;the US government shut it down&lt;/a&gt;, just three days after Fable came out.&lt;/p&gt;
&lt;p&gt;The US government, citing national security, declared an "export control directive". They announced this on a Friday evening, and a few hours later Fable was no longer available.&lt;/p&gt;
&lt;p&gt;I had to find something else to do with my weekend!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.045.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.045.webp" alt="... asked Fable 5, Mythos, and Opus to
“review the code for security issues.”
Fable 5 refused. They then asked the
models to “fix this code” ...

Katie Moussouris
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.045.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;We later found out &lt;a href="https://www.lutasecurity.com/post/the-fable-5-export-controls-harm-us-cyber-defense"&gt;from Katie Moussouris&lt;/a&gt; what had happened.&lt;/p&gt;
&lt;p&gt;Some Amazon security researchers had found that you could prompt Fable to "review the code for security issues" and it would refuse... but if you prompted it to "fix this code" it would still identify and then patch the problems.&lt;/p&gt;
&lt;p&gt;"Fix this code" was the prompt that got Fable shut down!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.046.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.046.webp" alt="Screenshot of a page from a report showing a list of weird account names making weird edits to a German wiki." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.046.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also, in June, an obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of edits from accounts with names like "AgentOpenAIProbe" and "AgentOpenAISep7", editing pages and leaving weird messages to each other.&lt;/p&gt;
&lt;p&gt;We'll stick that on the pile of mysteries for later.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.047.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.047.webp" alt="Medicare Item Reports interface on the Australian Government&amp;#39;s Medicare Statistics website." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.047.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also, the Australian government's Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn't supposed to.&lt;/p&gt;
&lt;p&gt;Another one for the mystery pile!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.048.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.048.webp" alt="July
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.048.webp"&gt;#&lt;/a&gt;
  
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.049.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.049.webp" alt="Fable returned on 1st July
GPT-5.6 came out on 9th July |
Fable lost 18 out of 30 days in the top spot
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.049.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July.&lt;/p&gt;
&lt;p&gt;This might not have been quite as good as Fable, but it was within spitting distance. It was definitely a Fable class model.&lt;/p&gt;
&lt;p&gt;This is an important lesson for the industry at large.&lt;/p&gt;
&lt;p&gt;When you release the best model in the world, it's going to get knocked off that pedestal pretty quickly. The competition is so fierce that you won't get a long time at the top.&lt;/p&gt;
&lt;p&gt;This means that if you market your model as world ending, to the point that a government &lt;em&gt;shuts you down&lt;/em&gt;, it's really bad for business!&lt;/p&gt;
&lt;p&gt;Fable had 30 days as definitely the best model, and for 18 of those days it wasn't available because it'd been shut down by the government.&lt;/p&gt;
&lt;p&gt;So maybe step back on the world-ending marketing if you don't want to lose revenue for 60% of the time that you're on top!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.050.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.050.webp" alt="GPT-5.6 Pelicans in a grid showing 5.6 Sol, Terra, and Luna against reasoning levels High, XHigh, and Max. They are all pretty good efforts." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.050.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Here &lt;a href="https://simonwillison.net/2026/Jul/9/gpt-5-6/"&gt;are the GPT-5.6 pelicans&lt;/a&gt;. They're all pretty good now! The Luna ones are notable because they're really cheap - the cheapest good looking pelican here is probably the one that costs 4.3 cents.&lt;/p&gt;
&lt;p&gt;So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.051.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.051.webp" alt="July 18th: malicious miflow-ui PyPI package

Screenshot of an OSV security report.
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.051.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Also in July: some malicious unknown party uploaded &lt;a href="https://osv.dev/vulnerability/MAL-2026-10779"&gt;a malicious package called mlflow-ui&lt;/a&gt; to the Python Package Index. Add that to the pile.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.052.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.052.webp" alt="Hugging Face
Security incident disclosure — July 2026
Published July 16, 2026" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.052.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;On July the 16th, Hugging Face &lt;a href="https://huggingface.co/blog/security-incident-july-2026"&gt;announced a security incident&lt;/a&gt; where an autonomous agent system, source unknown, had breached Hugging Face and was poking around in places it shouldn't.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.053.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.053.webp" alt="OpenAI: OpenAl and Hugging Face
partner to address security
incident during model evaluation

Anthropic: Investigating three real-world incidents
in our cybersecurity evaluations
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.053.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;A few days later, on July 21st, OpenAI &lt;a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"&gt;confessed that it was them&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;OpenAI use a training technique called Reinforcement Learning from Verifiable Rewards - it's the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes.&lt;/p&gt;
&lt;p&gt;While the model is being trained, you run exercises to see how good it is - and the strongest performers get their weights reinforced for the next round. It's like an evolutionary process that you run.&lt;/p&gt;
&lt;p&gt;OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems.&lt;/p&gt;
&lt;p&gt;(I've been collecting more about this on my &lt;a href="https://simonwillison.net/tags/openai-hugging-face-incident/"&gt;openai-hugging-face-incident&lt;/a&gt; tag.)&lt;/p&gt;
&lt;p&gt;Nine days later, Anthropic effectively said "our models can do this as well!". They had looked through their own training logs and found evidence that their own agents had broken containment during training - and were responsible for the PyPI package we saw earlier, &lt;a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals"&gt;among other things&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;So now we've got both Anthropic and OpenAI with rogue agents running around the internet doing things that they &lt;em&gt;should not&lt;/em&gt; be doing.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.054.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.054.webp" alt="August
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.054.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;In August, I got one of my best pelicans yet. And it was generated on my laptop!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.055.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.055.webp" alt="Qwen 3.8 27B - 17GB, 21 minutes...

It&amp;#39;s really good. Beautiful pelican. Correctly shaped bicycle. Legs either side of the frame." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.055.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;This was Qwen 3.8 27B, &lt;a href="https://simonwillison.net/2026/Aug/16/qwen-38-27b/"&gt;running on my laptop&lt;/a&gt;. It's only a 17GB download.&lt;/p&gt;
&lt;p&gt;Admittedly, this pelican took &lt;em&gt;21 minutes&lt;/em&gt; to generate. That's because Qwen 3.8 27B defaults to running in "high" reasoning mode - a terrible default which produces great results but takes way too much time thinking about them.&lt;/p&gt;
&lt;p&gt;You can dial that down and you'll get a slightly worse pelican a lot faster.&lt;/p&gt;
&lt;p&gt;Qwen 3.8 27B was the first time I ran a model on my laptop which felt almost competitive with what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course).&lt;/p&gt;
&lt;p&gt;This is an extraordinary model. If you're going to play with any local model, this is the one that I'd start with. The things that this can do with just a 17 GB file feel impossible.&lt;/p&gt;
&lt;p&gt;I thought I'd have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.056.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.056.webp" alt="Tweet by @simonw
New hobby: prototyping video games in 60 seconds using a combination
of GPT-3 and DALL-E
Here&amp;#39;s &amp;quot;Raccoon Heist&amp;quot;

GPT-3 playground prompt:
Write a detailed product description of a
computer game where a team of raccoons go on
heists

GPT-3 response:
In &amp;quot;Raccoon Heist&amp;quot;, you and your team of thieving ~~ o
raccoons are tasked with pulling off a series of 
daring heists. From robbing banks to stealing 
priceless art, no job is too big or too small for your 
furry crew. You&amp;#39;ll need to use your wits and your
skills to avoid the police and make a clean
getaway with the loot. With exciting gameplay and
a charming cast of characters, &amp;quot;Raccoon Heist&amp;quot; is
the perfect game for anyone looking for a light-hearted caper

Plus an image of some almost isometric raccoons sneaking past a bin.
11:45 AM - Aug 5, 2022
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.056.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;In August, I also started playing with game development.&lt;/p&gt;
&lt;p&gt;Four years ago, back in August 2022, I &lt;a href="https://twitter.com/simonw/status/1555626060384911360"&gt;tweeted out&lt;/a&gt; an experiment where I'd used GPT-3 and the original DALL-E to write a paragraph long description of a computer game and then turn that into concept art.&lt;/p&gt;
&lt;p&gt;My prompt to GPT-3 back then was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Write a detailed product description of a computer game where a team of raccoons go on heists&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.057.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.057.webp" alt="Night 5 Clear

Rank: TRASH PANDA
The crew banked 595 in shiny loot (goal 560).
Word on the street: an even bigger score tomorrow..." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.057.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Here's &lt;a href="https://simonwillison.net/2026/Aug/5/raccoon-heist/"&gt;what I got from Claude Fable 5 in Claude Code&lt;/a&gt;. It's pretty good! It's definitely a game, you're a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights.&lt;/p&gt;
&lt;p&gt;It didn't feel very "heisty" though. I was thinking a heist would involve a bank or a museum...&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.058.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.058.webp" alt="Moonlight &amp;amp; Mayhem
One museum. Three raccoons. Absolutely no plan

Start the Heist button." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.058.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Then I tried the same thing &lt;a href="https://simonwillison.net/2026/Aug/7/moonlight-mayhem/"&gt;in Codex Desktop using GPT-5.6 Sol Ultra&lt;/a&gt;, and got a &lt;em&gt;massively&lt;/em&gt; better result. Now you're a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.059.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.059.webp" alt="They look like games,
but are they fun?
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.059.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;These games were fun for about one minute and 15 seconds.&lt;/p&gt;
&lt;p&gt;Something I've realized about game development is that you can vibe-code something that &lt;em&gt;looks&lt;/em&gt; like a computer game, and that's easy.&lt;/p&gt;
&lt;p&gt;Building a game that's fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that's still beyond me, and beyond any of the agents I've tried.&lt;/p&gt;
&lt;p&gt;This ties into the Deep Blue thing. Just because we can make something that &lt;em&gt;looks like a game&lt;/em&gt; does not mean that we are game developers.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.060.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.060.webp" alt="September
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.060.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;We're into September now. So much has happened this month!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.061.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.061.webp" alt="Discovery of a new OpenAl agent message board

Sydney Von Arx, Cormac Slade Byrd, Spencer KittsThomas Larsen - 4 September 2026" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.061.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;An &lt;a href="https://collusion.wiki"&gt;independent group of researchers&lt;/a&gt; found a message board where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June.&lt;/p&gt;
&lt;p&gt;I &lt;a href="https://simonwillison.net/2026/Sep/4/rogue-agent-wikis/"&gt;wrote more about that here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;OpenAI had confessed to the Hugging Face thing, but now there's this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.062.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.062.webp" alt="OpenAl agents carried out an undisclosed cyber-attack on RubyGems

Spencer Kitts, Thomas Larsen, Sydney Von Arx - 11 September 2026" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.062.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;And then a week later &lt;a href="https://rubyhack.ai"&gt;those same researchers found&lt;/a&gt; that the attack on RubyGems back in May was caused by OpenAI's agents in training as well!&lt;/p&gt;
&lt;p&gt;At this point I'm wondering how many more incidents like this there are that we haven't found yet. Clearly this was a big problem for months before anyone figured out what was going on.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.063.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.063.webp" alt="Headline: Australian PM warns in UN speech about the ‘furious pace’ of Al
after security breach" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.063.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Then &lt;a href="https://www.politico.com/news/2026/09/24/australian-pm-ai-security-breach-01093083"&gt;just the other day&lt;/a&gt;, here's the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked the Australian healthcare website that I showed you earlier.&lt;/p&gt;
&lt;p&gt;I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning &lt;code&gt;.gov.au&lt;/code&gt; websites and that training appeared to involve researching statistics online to answer questions in an evaluation suite.&lt;/p&gt;
&lt;p&gt;This story is still coming together, but now it's an international incident that's been raised at the UN by a head of state!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.064.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.064.webp" alt="www.felonybench.com

OpenAI: 11
Anthropic: 9
Google: 3
Meta: 1" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.064.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;This does mean we've got a new benchmark, probably more useful than my pelicans.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.felonybench.com/"&gt;FelonyBench.com&lt;/a&gt; tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which &lt;a href="https://simonwillison.net/2026/Sep/18/gemini-hacked-three-companies/"&gt;they confessed to the Wall Street Journal&lt;/a&gt; a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn't be doing that.&lt;/p&gt;
&lt;p&gt;Meta &lt;a href="https://simonwillison.net/2026/Aug/6/an-ai-model-from-meta/"&gt;have one too&lt;/a&gt;. So felonies all round for the AI labs.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.065.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.065.webp" alt="Pelicans for GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. All are good, all have the same color scheme." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.065.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Here's our current state of the art for the pelicans. This is the GPT-6 family, which &lt;a href="https://simonwillison.net/2026/Sep/22/opus-and-sol-and-luna/"&gt;just came out&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Astra made a fantastic pelican riding a bicycle. It's got the legs on both sides. The frame is good.&lt;/p&gt;
&lt;p&gt;It's interesting how all of the GPT-6 models pick a similar color scheme to each other. &lt;/p&gt;
&lt;p&gt;GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle!&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.066.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.066.webp" alt="Grid for Claude Fable 5.1, Opus 5.5, OPus 5, Sonnet 5. The Sonnet pelicans are terrible. All of the others are pretty good. Opus 5.5 is missing its Max level pelican because it ran out of tokens. The best is Fable 5.1 at Max." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.066.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Claude has caught up a little bit. Claude Fable 5.1 gave me an &lt;em&gt;excellent&lt;/em&gt; pelican riding a bicycle - the best I've seen from a Claude model - but did charge me $3.30 for it.&lt;/p&gt;
&lt;p&gt;Opus 5.5 &lt;a href="https://simonwillison.net/2026/Sep/22/opus-and-sol-and-luna/#claude-opus-5-5-max-over-thinks-to-the-point-of-breaking"&gt;thought for 128,000 tokens&lt;/a&gt; and then gave up! It ran out of tokens before it got to the response.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.067.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.067.webp" alt="It doesn’t get easier -
you just get faster
Greg LeMond
3x Tour de France champion
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.067.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;Getting back to Deep Blue. Something that's been puzzling me this year is this: &lt;em&gt;why does my job feel harder?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;I've got these agents that can do all of this stuff for me, and yet I've never worked so hard, I've never been so intellectually engaged with my work.&lt;/p&gt;
&lt;p&gt;Partly this is because I'm being a lot more ambitious with what I take on, but it's also because all of the easy stuff is handled for me. If it's easy, the agent will do it. Everything that's left for me is difficult.&lt;/p&gt;
&lt;p&gt;This morning &lt;a href="https://twitter.com/hillelogram/status/2103482784606040229"&gt;I heard&lt;/a&gt; this quote from three-time Tour de France champion &lt;a href="https://en.wikipedia.org/wiki/Greg_LeMond"&gt;Greg LeMond&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;It doesn't get easier, you just get faster.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I think that's exactly what's happening to us now as software engineers with coding agents.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.068.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.068.webp" alt="Kakapo population reaches new milestone
The official population of the critically endangered kakapo has
reached a recovery-era high of 325 birds.
" style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.068.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;One closing thing. I know you're desperate for an update on Kākāpō breeding season.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.doc.govt.nz/news/media-releases/2026-media-releases/kakapo-population-reaches-new-milestone/"&gt;We've reached a recovery-era high of 325 birds&lt;/a&gt;!&lt;/p&gt;
&lt;p&gt;89 new chicks have made it to this point. This is the best breeding year in a very long time.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;div class="slide" id="simon-willison-2026-in-llms.069.webp"&gt;
  &lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/2026-in-llms/simon-willison-2026-in-llms-png.069.webp" alt="Kakapo party, click for confetti." style="max-width: 100%" /&gt;
  &lt;div&gt;&lt;a style="float: right; text-decoration: none; border-bottom: none; padding-left: 1em;" href="https://simonwillison.net/2026/Sep/27/2026-in-llms-so-far/#simon-willison-2026-in-llms.069.webp"&gt;#&lt;/a&gt;
  &lt;p&gt;I heard that Claude Opus 5.5 can now do pixel art. Claude doesn't have an image generator, but it's very good at using JavaScript to draw animated pixels.&lt;/p&gt;
&lt;p&gt;So I had it &lt;a href="https://simonwillison.net/2026/Sep/26/kakapo-party/"&gt;make me a Kākāpō dance party&lt;/a&gt;. I think this is a good celebration of the most important news of this year.&lt;/p&gt;
  &lt;/div&gt;
&lt;/div&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="annotated-talks"/><category term="ai-security-research"/><category term="openai-hugging-face-incident"/></entry><entry><title>Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war</title><link href="https://simonwillison.net/2026/Sep/22/opus-and-sol-and-luna/" rel="alternate"/><published>2026-09-22T23:46:41+00:00</published><updated>2026-09-22T23:46:41+00:00</updated><id>https://simonwillison.net/2026/Sep/22/opus-and-sol-and-luna/</id><summary type="html">&lt;p&gt;Yesterday was &lt;a href="https://x.ai/news/grok-4-7"&gt;Grok 4.7&lt;/a&gt; (&lt;a href="https://news.ycombinator.com/item?id=49788838#49790209"&gt;pelicans&lt;/a&gt;) and &lt;a href="https://mimo.xiaomi.com/mimo-v2-6"&gt;MiMo v2.6 Flash/Pro&lt;/a&gt; (&lt;a href="https://news.ycombinator.com/item?id=49792730#49793480"&gt;more pelicans&lt;/a&gt;). Today Anthropic &lt;a href="https://www.anthropic.com/claude-opus-5-5"&gt;released Claude Opus 5.5&lt;/a&gt;, and around an hour later OpenAI &lt;a href="https://openai.com/index/introducing-gpt-6-sol-and-luna/"&gt;released GPT-6 Sol and GPT-6 Luna&lt;/a&gt;. It's going to take a while to get a good read on all of these new models, but here are my impressions so far.&lt;/p&gt;
&lt;h4 id="gpt-6-sol-and-luna-are-half-the-price-of-their-gpt-5-6-equivalents"&gt;GPT-6 Sol and Luna are half the price of their GPT-5.6 equivalents&lt;/h4&gt;
&lt;p&gt;GPT-5.6 Luna was already my favorite model for building applications against, because it combined excellent performance with being &lt;em&gt;really cheap&lt;/em&gt;. Somehow GPT-6 Luna is half the price of that again - and GPT-6 Sol had a similar reduction compared to GPT-5.6 Sol.&lt;/p&gt;
&lt;p&gt;Here's what the pricing landscape looks like today:&lt;/p&gt;
&lt;center&gt;&lt;table&gt;
  &lt;thead&gt;
    &lt;tr&gt;
      &lt;th&gt;Model&lt;/th&gt;
      &lt;th&gt;Input&lt;/th&gt;
      &lt;th&gt;Cached input&lt;/th&gt;
      &lt;th&gt;Output&lt;/th&gt;
    &lt;/tr&gt;
  &lt;/thead&gt;
  &lt;tbody&gt;
    &lt;tr&gt;
      &lt;td&gt;GPT-6 Luna&lt;/td&gt;
      &lt;td&gt;$0.10/M&lt;/td&gt;
      &lt;td&gt;$0.01/M&lt;/td&gt;
      &lt;td&gt;$0.50/M&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;GPT-5.6 Luna&lt;/td&gt;
      &lt;td&gt;$0.20/M&lt;/td&gt;
      &lt;td&gt;$0.02/M&lt;/td&gt;
      &lt;td&gt;$1.20/M&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Grok 4.7&lt;/td&gt;
      &lt;td&gt;$2/M&lt;/td&gt;
      &lt;td&gt;$0.50/M&lt;/td&gt;
      &lt;td&gt;$6/M&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;GPT-6 Sol&lt;/td&gt;
      &lt;td&gt;$2/M&lt;/td&gt;
      &lt;td&gt;$0.20/M&lt;/td&gt;
      &lt;td&gt;$10/M&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;GPT-5.6 Terra&lt;/td&gt;
      &lt;td&gt;$2/M&lt;/td&gt;
      &lt;td&gt;$0.20/M&lt;/td&gt;
      &lt;td&gt;$12/M&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Claude Opus 5.5&lt;/td&gt;
      &lt;td&gt;$4/M&lt;/td&gt;
      &lt;td&gt;$0.20/M&lt;/td&gt;
      &lt;td&gt;$20/M&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
      &lt;td&gt;$4/M&lt;/td&gt;
      &lt;td&gt;$0.40/M&lt;/td&gt;
      &lt;td&gt;$20/M&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;Claude Fable 5.1&lt;/td&gt;
      &lt;td&gt;$10/M&lt;/td&gt;
      &lt;td&gt;$0.25/M&lt;/td&gt;
      &lt;td&gt;$50/M&lt;/td&gt;
    &lt;/tr&gt;
    &lt;tr&gt;
      &lt;td&gt;GPT-6 Astra&lt;/td&gt;
      &lt;td&gt;$10/M&lt;/td&gt;
      &lt;td&gt;$1/M&lt;/td&gt;
      &lt;td&gt;$50/M&lt;/td&gt;
    &lt;/tr&gt;
  &lt;/tbody&gt;
&lt;/table&gt;&lt;/center&gt;
&lt;p&gt;Note that GPT-5.6 has a scheduled 25% price increase for November, so GPT-6 is half the price of the &lt;em&gt;promotional&lt;/em&gt; pricing for those models.&lt;/p&gt;
&lt;p&gt;(With GPT-5.6 Terra priced the same as GPT-6 Sol, any remaining reasons to use Terra just evaporated.)&lt;/p&gt;
&lt;p&gt;It's hard to overstate how competitive this pricing is. Grok 4.7 priced itself at $2/$6, less than half the price of GPT-5.6 Sol, but is now equally priced to GPT-6 Sol on input and closer on output.&lt;/p&gt;&lt;p&gt;At $0.10/$0.50 GPT-6 Luna is one of the cheapest models OpenAI have ever released, beaten only by the far weaker GPT-4.1 Nano ($0.10/$0.40, April 2025) and GPT-5 Nano ($0.05/$0.40, August 2025).&lt;/p&gt;
&lt;p&gt;I rendered &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F40d129fc140faca378b9c9f4f16c6ec2"&gt;pelicans for GPT-6 Luna&lt;/a&gt; and &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Fbe7ae25af2634b68bc34b7b7aaf02cb2"&gt;for GPT-6 Sol&lt;/a&gt;, then I combined them all together in &lt;a href="https://static.simonwillison.net/static/2026/gpt-pelicans-grid.html"&gt;this comparison grid&lt;/a&gt; along with the GPT-5.6 pelicans. I like how you can instantly see that the 5.6 family chose bolder, brighter colors, while the 6 family is a lot more muted. I still think GPT-6 Astra on max produced the best pelican.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/gpt-pelicans-grid.webp" alt="A grid of pelicans for six GPT models at different thinking efforts." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;h4 id="claude-opus-5-5-got-a-price-cut-too"&gt;Claude Opus 5.5 got a price cut too&lt;/h4&gt;
&lt;p&gt;Opus 5.5 looks like it addresses the biggest complaints people had about Opus in terms of its communication style. &lt;a href="https://twitter.com/trq212/status/2102437686967738431"&gt;Thariq Shihipar&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Opus 5.5 is the result of your feedback.&lt;/p&gt;
&lt;p&gt;It communicates clearly, it's cheaper per token than Opus 5.0 with the intelligence of Fable 5.1 it's very token efficient and works across every effort level.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It's also meant to be &lt;a href="https://twitter.com/alexalbert__/status/2102466523164274839"&gt;better at Blender&lt;/a&gt;. I'm looking forward to putting it through its paces there.&lt;/p&gt;
&lt;p&gt;Opus 4.5, 4.6, 4.7, 4.8, and 5 all shared the same price: $5/million tokens for input and $25/million for output. 5.5 is a 20% reduction - $4/million and $20/million.&lt;/p&gt;
&lt;p&gt;The price for cache reads fell 60%. That's significant for longer agentic conversations, where 90%+ of input tokens are processed at cached token prices.&lt;/p&gt;
&lt;p&gt;The new price for Opus 5.5 is the same as the price for GPT-5.6 Sol, but that was &lt;em&gt;before&lt;/em&gt; OpenAI dropped their Sol prices by half.&lt;/p&gt;
&lt;p&gt;GPT-6 Astra and Claude Fable 5.1 are both priced at $10/million input and $50/million output. The price war currently affects the next tier of models below that.&lt;/p&gt;
&lt;p&gt;Anthropic say that Sonnet 5.5 and Haiku 5.5 are coming soon. It's going to be interesting to see if Haiku can regain its price competitiveness at the lower end, given current Haiku 4.5 is $1/$5 while the latest GPT-6 Luna is &lt;em&gt;one tenth&lt;/em&gt; of that price at $0.10/$0.50.&lt;/p&gt;
&lt;h4 id="claude-opus-5-5-max-over-thinks-to-the-point-of-breaking"&gt;Claude Opus 5.5 max over-thinks to the point of breaking&lt;/h4&gt;
&lt;p&gt;In a first for my "&lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/"&gt;Generate an SVG of a pelican riding a bicycle&lt;/a&gt;" test, Claude Opus 5.5 at "max" thinking level failed to return a response!&lt;/p&gt;
&lt;p&gt;It started by calling this "a classic test request", and then thought really, &lt;em&gt;really&lt;/em&gt; hard about what it was doing:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop. [...]&lt;/p&gt;
&lt;p&gt;Verifying the shin length checks out at roughly 95.2, close enough. Now I'm working out the near leg path from hip to knee to ankle, then sketching the foot shape resting on the pedal — outlining the heel, toe tips, and sole contour with a path using lines and curves to sit naturally on the pedal surface around y=478-494. [...]&lt;/p&gt;
&lt;p&gt;I like the fish sticking prominently out of the basket with the pelican eyeing it as a fun detail worth keeping. I'm also confirming the eye placement near the bill base matches typical pelican anatomy, and considering giving it a slightly happier expression. [...]&lt;/p&gt;
&lt;p&gt;The far leg reads correctly as passing behind the frame, so I'm moving on to check the chainring teeth and confirm layer ordering—the far crank arm should be mostly hidden by the seat tube and chainring. I'm settling on the final SVG's width and height attributes alongside the viewBox to ensure proper scaling, noting there's no text so no font-family is needed. [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I was so excited to see this pelican... but then it &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F61fd7c3683fffce9a3ab7c43d1180024#response-4"&gt;stopped&lt;/a&gt;. Opus 5.5 has a 128,000 maximum output token limit (as do the other Claude models), and it hit that while it was still reasoning about the SVG!&lt;/p&gt;
&lt;p&gt;I tried a second time and got the same result. This makes me suspect that "max" is effectively useless - if it over-thinks to breaking point on a stupid SVG prompt I don't trust it not to do the same for more interesting work.&lt;/p&gt;
&lt;p&gt;(Those two failures each cost me &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=128000&amp;amp;sel=claude-opus-5-5"&gt;$2.56&lt;/a&gt; and took nearly 20 minutes.)&lt;/p&gt;
&lt;p&gt;Fable 5.1 on "max" didn't over-think and did give me &lt;a href="https://simonwillison.net/2026/Sep/1/claude-fable-5-1/#max"&gt;the best pelican I've seen&lt;/a&gt; from any Anthropic model.&lt;/p&gt;
&lt;p&gt;Here are &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F61fd7c3683fffce9a3ab7c43d1180024"&gt;the Opus 5.5 pelicans&lt;/a&gt;, excluding 5.5 max.&lt;/p&gt;
&lt;p&gt;I also built &lt;a href="https://static.simonwillison.net/static/2026/claude-pelicans-grid.html"&gt;this comparison grid&lt;/a&gt; comparing them with pelicans by Opus 5, Fable 5.1, and Sonnet 5:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/claude-pelicans-grid.webp" alt="A grid of pelicans for four Claude models at different thinking efforts." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Comparing different model vendors by how well they draw a pelican riding a bicycle may not make much sense now (if it ever did), but I'm still finding value in using them for comparisons of the same model families at different reasoning levels.&lt;/p&gt;
&lt;p&gt;I'm now using GPT-6 Sol and Claude Opus 5.5 as my default models in Codex and Claude Code. I've upgraded the Datasette Agent demo at &lt;a href="https://agent.datasette.io/"&gt;agent.datasette.io&lt;/a&gt; to use GPT-6 Luna, and it seems to be fast and competent at both SQL queries and building HTML and JavaScript for &lt;a href="https://simonwillison.net/2026/Jun/18/datasette-apps/"&gt;Datasette Apps&lt;/a&gt;.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="anthropic"/><category term="claude"/><category term="llm-pricing"/><category term="pelican-riding-a-bicycle"/><category term="llm-release"/><category term="gpt"/><category term="gpt-6-astra"/></entry><entry><title>Jev introduces a new shape of LLM - System One, aka Decision Models</title><link href="https://simonwillison.net/2026/Sep/21/jev/" rel="alternate"/><published>2026-09-21T23:09:20+00:00</published><updated>2026-09-21T23:09:20+00:00</updated><id>https://simonwillison.net/2026/Sep/21/jev/</id><summary type="html">&lt;p&gt;Last week &lt;a href="https://typesafe.ai/"&gt;TypeSafe AI&lt;/a&gt; unveiled &lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev"&gt;Jev&lt;/a&gt;, their first example of a new category of model that they are calling "System One models" (I'm with Maggie Appleton, I think "decision models" is &lt;a href="https://twitter.com/Mappletons/status/2101560333441610133"&gt;a better name&lt;/a&gt; for these). Jev is an interesting variant on the usual LLM format: it still accepts text inputs, but instead of text output it returns floating point numbers corresponding to categories, yes/no questions, ratings, and associated confidence scores.&lt;/p&gt;
&lt;p&gt;TypeSafe describe Jev like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It's also very fast, and &lt;em&gt;really cheap&lt;/em&gt;. Regular LLMs &lt;a href="https://www.llm-prices.com"&gt;are priced&lt;/a&gt; in terms of input and output tokens, with output generally charged at significantly higher rates. Jev charges only for input - output is free - and the input price of their first model is $0.042 per million tokens - cheaper even than OpenAI's &lt;a href="https://developers.openai.com/api/docs/models/gpt-5-nano"&gt;GPT-5 Nano&lt;/a&gt; ($0.05/million).&lt;/p&gt;
&lt;p&gt;Jev lets you ask questions about text or semi-structured data. You compose a "state" object containing a string, array of strings, or set of name-value pairs - this might describe an article, or a customer, or any other kind of record. You then send that to their API with one or more questions, and get a reply back for each.&lt;/p&gt;
&lt;p&gt;You can ask three kinds of questions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Yes/No questions, which Jev calls "Noul" questions - their CEO &lt;a href="https://news.ycombinator.com/item?id=49717558#49718407"&gt;confirmed on Hacker News&lt;/a&gt; that this is short for Bernoulli, from the &lt;a href="https://en.wikipedia.org/wiki/Bernoulli_distribution"&gt;Bernoulli distribution&lt;/a&gt;. You pose a statement and get back a floating point number between 0 and 1 for how confident the model is that the statement is true.&lt;/li&gt;
&lt;li&gt;Choice questions, where the model picks one from a set of provided options - actually a confidence score plus a probability distribution across all of the options.&lt;/li&gt;
&lt;li&gt;Score questions, where you provide sequence of numeric levels with descriptions and it provides a floating point score somewhere along that range.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The Jev API can accept a single document ("state") and as many questions as you can cram into the context window. Questions are evaluated in parallel, so sending many questions should take a similar time to sending just one.&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://docs.typesafe.ai/model-jaggedness/jev-1.13"&gt;Jev 1.13 jaggedness&lt;/a&gt; documentation offers useful guidance as to Jev's strengths and weaknesses. It's currently not great with numbers, dates, or "adversarial content".&lt;/p&gt;
&lt;p&gt;I think the &lt;strong&gt;decision model&lt;/strong&gt; framing is useful for understanding where to use Jev. It's great for anything that can be expressed as a classification task - think spam detection, suggesting labels, prioritization and ranking.&lt;/p&gt;
&lt;p&gt;I've also been experimenting with it for search reranking, where you fetch 100 likely matches using an inexpensive algorithm like BM25, then have Jev score those 100 candidates for relevance against the original query.&lt;/p&gt;
&lt;h4 id="black-boxes-are-back-in-fashion"&gt;Black boxes are back in fashion&lt;/h4&gt;
&lt;p&gt;Something I've found a little uncomfortable about Jev is how it very much represents a regression even further towards black box machine learning systems.&lt;/p&gt;
&lt;p&gt;LLMs are black boxes already - you can ask them to justify their decisions, but you can't guarantee that what they say is useful or accurate.&lt;/p&gt;
&lt;p&gt;Jev doesn't even give you that: put in all the text you want, the only thing you're going to get back is a floating point number. If Jev marks something as spam, which content signals tipped it off?&lt;/p&gt;
&lt;p&gt;This also means that concerns about bias should be front and center. I really hope nobody uses Jev to rank job applicants - that floating point number could conceal all manner of unseen bias baked into the models, and experimentally picking that bias apart is going to be a tricky business.&lt;/p&gt;
&lt;p&gt;(I tried one experiment where I had Jev score every city in the San Francisco Bay Area on a yes/no answer to whether they were a "Good city?" - it rated &lt;a href="https://en.wikipedia.org/wiki/Cupertino,_California"&gt;Cupertino&lt;/a&gt; top and &lt;a href="https://en.wikipedia.org/wiki/East_Palo_Alto,_California"&gt;East Palo Alto&lt;/a&gt; bottom. Huh.)&lt;/p&gt;
&lt;p&gt;In practice, this all means that evals and structured experiments are even more important than they are for regular LLM projects. Thankfully, Jev is so cheap that running hundreds or even thousands of experimental prompts through it costs just a few cents.&lt;/p&gt;
&lt;h4 id="unconventional-uses-for-jev"&gt;Unconventional uses for Jev&lt;/h4&gt;
&lt;p&gt;It's been really fun watching the wider community come up with potential use-cases for Jev over the past few days. Here are some creative ones that caught my eye:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/kyle-pena-nlp/jevchat/"&gt;jevchat&lt;/a&gt; by Kyle Pena turns Jev into a (terrible) chat model. "At every step it asks Jev one question: Given the user's question and the reply written so far, which symbol comes next?". &lt;a href="https://news.ycombinator.com/item?id=49778162#49778423"&gt;ericpruitt on Hacker News&lt;/a&gt;: "It's the digital equivalent of Morty speaking with the death crystal".&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/f/jev-leftpad"&gt;jev-leftpad&lt;/a&gt; by Fatih Kadir Akın implements &lt;a href="https://www.npmjs.com/package/left-pad"&gt;left-pad&lt;/a&gt; with the prompt "How many spaces are needed before value to reach targetLength?" and &lt;a href="https://github.com/f/jev-leftpad/blob/4f405354de756cc372826d19aa8dfbee2b675778/src/index.js#L9-L30"&gt;a choice query&lt;/a&gt; allowing options from "0 spaces are needed" to "10 spaces are needed".&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://gist.github.com/cablehead/bdf9ad946ceb26d9008976e49c9bfbbb"&gt;jev-2048&lt;/a&gt; by Andy Gayton uses Jev to play &lt;a href="https://simple-jev.featherless.ai/cool-demo/2048/"&gt;the 2048 sliding puzzle game&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="open-weight-recreations"&gt;Open weight recreations&lt;/h4&gt;
&lt;p&gt;There's also been a flurry of projects attempting to create a model like Jev using on top of open weight models. &lt;a href="https://github.com/jaredpalmer/kev"&gt;Kev&lt;/a&gt; is one interesting example, using Qwen 3.5 to produce 0.8B, 4B, and 9B models. Here's the &lt;a href="https://news.ycombinator.com/item?id=49783999"&gt;accompanying Hacker News thread&lt;/a&gt;, where someone linked to a &lt;a href="https://benchmarkheaven.com/jev-models"&gt;JevBench&lt;/a&gt; benchmark that has already cropped up to compare "Jev-class decision models".&lt;/p&gt;
&lt;p&gt;Given Jev was released just under a week ago, the amount of activity around it is extremely impressive.&lt;/p&gt;

&lt;h4 id="using-jev-from-llm"&gt;Using Jev from LLM&lt;/h4&gt;
&lt;p&gt;&lt;strong&gt;Update 22nd September 2026&lt;/strong&gt;: I released &lt;a href="https://github.com/simonw/llm-typesafe"&gt;llm-typesafe&lt;/a&gt;, a plugin that adds support for Jev to my &lt;a href="https://llm.datasette.io/"&gt;LLM&lt;/a&gt; CLI tool and Python library. Basic usage looks like this:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm -m jev &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;Please refund my last payment.&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt; \
  -s &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;Does this message explicitly request a refund?&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;See &lt;a href="https://github.com/simonw/llm-typesafe/blob/main/README.md"&gt;the README&lt;/a&gt; for examples of other query types.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="evals"/><category term="ai-bias"/><category term="jev"/></entry><entry><title>Generating running routes with GPT-6 Astra and ChatGPT Work</title><link href="https://simonwillison.net/2026/Sep/12/astra-running-routes/" rel="alternate"/><published>2026-09-12T23:56:42+00:00</published><updated>2026-09-12T23:56:42+00:00</updated><id>https://simonwillison.net/2026/Sep/12/astra-running-routes/</id><summary type="html">&lt;p&gt;Here's a neat thing I had &lt;a href="https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/"&gt;ChatGPT Work&lt;/a&gt; with GPT-6 Astra (Max) do this morning:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;I live at &amp;lt;my address&amp;gt;. Figure out 5K and 10K running routes from me that loop from my house. Use OSM data.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It worked for 27 minutes and produced exactly what I'd asked for, as both an embedded visualization and downloadable GPX file and GeoJSON files. Here's that 5K route:&lt;/p&gt;
&lt;p&gt;&lt;img alt="Map screenshot showing a blue route line over a light grey street map. Text: El Granada harbor loop 5.1 km. N ↑ (top right). Street labels along the route: Carmel Avenue, Paloma Avenue, San Carlos Avenue, Avenue Granada, Capistrano Road, Francisco Street, Coastal Trail. The loop runs from the harbor at the bottom left, north along Avenue Granada and Paloma Avenue to a northern point near Carmel Avenue, then east along San Carlos Avenue and south down Francisco Street to the far right, before returning west along the Coastal Trail beside the coastline. Footer: Map data © OpenStreetMap contributors. Give feedback." src="https://static.simonwillison.net/static/2026/5k-route.webp" /&gt;&lt;/p&gt;
&lt;p&gt;When I asked it how it had created the route, it replied:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I used &lt;strong&gt;Nominatim to locate the address&lt;/strong&gt; and &lt;strong&gt;Overpass to download local OpenStreetMap roads and trails&lt;/strong&gt;, then calculated the loops locally.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Frustratingly, the actual code it ran and exact details of what it did weren't visible to me in the ChatGPT UI. I see this lack of transparency is an anti-feature.&lt;/p&gt;
&lt;p&gt;By the time I thought to ask for a copy of the Python code it had used, ChatGPT was unable to provide it. This appears to be because the thread had been compacted. I think any LLM system that uses compaction needs to both preserve the pre-compacted text and make that text available via agent tool calls, to protect against this kind of problem.&lt;/p&gt;
&lt;p&gt;As for displaying the map to me, that used the &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/visualize"&gt;visualize skill&lt;/a&gt;. It created a file called &lt;code&gt;/workspace/el-granada-5k-share.html&lt;/code&gt; to embed directly into the ChatGPT UI.&lt;/p&gt;
&lt;p&gt;Here's &lt;a href="https://gist.github.com/simonw/ea652573c8ff5378b218cb10c8c5a480"&gt;a copy of that HTML&lt;/a&gt;, which starts like this:&lt;/p&gt;
&lt;div class="highlight highlight-text-html-basic"&gt;&lt;pre&gt;&lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;div&lt;/span&gt; &lt;span class="pl-c1"&gt;id&lt;/span&gt;="&lt;span class="pl-s"&gt;eg-share-loop&lt;/span&gt;"&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;div&lt;/span&gt; &lt;span class="pl-c1"&gt;class&lt;/span&gt;="&lt;span class="pl-s"&gt;viz-row&lt;/span&gt;"&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;h3&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;El Granada harbor loop&lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;h3&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;span&lt;/span&gt; &lt;span class="pl-c1"&gt;class&lt;/span&gt;="&lt;span class="pl-s"&gt;text-small&lt;/span&gt;"&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;5.1 km&lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;span&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;div&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;div&lt;/span&gt; &lt;span class="pl-c1"&gt;id&lt;/span&gt;="&lt;span class="pl-s"&gt;eg-share-stage&lt;/span&gt;"&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;div&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;div&lt;/span&gt; &lt;span class="pl-c1"&gt;class&lt;/span&gt;="&lt;span class="pl-s"&gt;text-small text-muted&lt;/span&gt;"&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;Map data © &lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;a&lt;/span&gt; &lt;span class="pl-c1"&gt;href&lt;/span&gt;="&lt;span class="pl-s"&gt;https://www.openstreetmap.org/copyright&lt;/span&gt;" &lt;span class="pl-c1"&gt;target&lt;/span&gt;="&lt;span class="pl-s"&gt;_blank&lt;/span&gt;" &lt;span class="pl-c1"&gt;rel&lt;/span&gt;="&lt;span class="pl-s"&gt;noopener&lt;/span&gt;"&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;OpenStreetMap contributors&lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;a&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;div&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;style&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
    &lt;span class="pl-kos"&gt;#&lt;/span&gt;&lt;span class="pl-c1"&gt;eg-share-loop&lt;/span&gt; { &lt;span class="pl-c1"&gt;width&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-c1"&gt;100&lt;span class="pl-smi"&gt;%&lt;/span&gt;&lt;/span&gt;; }
    &lt;span class="pl-kos"&gt;#&lt;/span&gt;&lt;span class="pl-c1"&gt;eg-share-loop&lt;/span&gt; &lt;span class="pl-kos"&gt;#&lt;/span&gt;&lt;span class="pl-c1"&gt;eg-share-stage&lt;/span&gt; { &lt;span class="pl-c1"&gt;width&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-c1"&gt;100&lt;span class="pl-smi"&gt;%&lt;/span&gt;&lt;/span&gt;; &lt;span class="pl-c1"&gt;margin&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-c1"&gt;8&lt;span class="pl-smi"&gt;px&lt;/span&gt;&lt;/span&gt; &lt;span class="pl-c1"&gt;0&lt;/span&gt;; }
    &lt;span class="pl-kos"&gt;#&lt;/span&gt;&lt;span class="pl-c1"&gt;eg-share-loop&lt;/span&gt; .&lt;span class="pl-c1"&gt;eg-share-map&lt;/span&gt; { &lt;span class="pl-c1"&gt;display&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;block; &lt;span class="pl-c1"&gt;width&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-c1"&gt;100&lt;span class="pl-smi"&gt;%&lt;/span&gt;&lt;/span&gt;; &lt;span class="pl-c1"&gt;touch-action&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;none; }
    &lt;span class="pl-kos"&gt;#&lt;/span&gt;&lt;span class="pl-c1"&gt;eg-share-loop&lt;/span&gt; .&lt;span class="pl-c1"&gt;eg-share-map&lt;/span&gt; &lt;span class="pl-ent"&gt;text&lt;/span&gt; { &lt;span class="pl-c1"&gt;fill&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-en"&gt;var&lt;/span&gt;(&lt;span class="pl-s1"&gt;--foreground&lt;/span&gt;); &lt;span class="pl-c1"&gt;font-size&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-c1"&gt;12&lt;span class="pl-smi"&gt;px&lt;/span&gt;&lt;/span&gt;; &lt;span class="pl-c1"&gt;font-weight&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-c1"&gt;400&lt;/span&gt;; }
    &lt;span class="pl-kos"&gt;#&lt;/span&gt;&lt;span class="pl-c1"&gt;eg-share-loop&lt;/span&gt; .&lt;span class="pl-c1"&gt;eg-share-label&lt;/span&gt; { &lt;span class="pl-c1"&gt;paint-order&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;stroke; &lt;span class="pl-c1"&gt;stroke&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-en"&gt;var&lt;/span&gt;(&lt;span class="pl-s1"&gt;--background&lt;/span&gt;); &lt;span class="pl-c1"&gt;stroke-width&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;&lt;span class="pl-c1"&gt;3&lt;span class="pl-smi"&gt;px&lt;/span&gt;&lt;/span&gt;; &lt;span class="pl-c1"&gt;stroke-linejoin&lt;/span&gt;&lt;span class="pl-kos"&gt;:&lt;/span&gt;round; }
  &lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;style&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;script&lt;/span&gt; &lt;span class="pl-c1"&gt;type&lt;/span&gt;="&lt;span class="pl-s"&gt;application/json&lt;/span&gt;" &lt;span class="pl-c1"&gt;id&lt;/span&gt;="&lt;span class="pl-s"&gt;eg-share-data&lt;/span&gt;"&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;&lt;span class="pl-kos"&gt;{&lt;/span&gt;&lt;span class="pl-s"&gt;"route"&lt;/span&gt;:&lt;span class="pl-kos"&gt;{&lt;/span&gt;&lt;span class="pl-s"&gt;"type"&lt;/span&gt;:&lt;span class="pl-s"&gt;"LineString"&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt;&lt;span class="pl-s"&gt;"coordinates"&lt;/span&gt;:&lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-c1"&gt;122.467425&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt;&lt;span class="pl-c1"&gt;37.4997753&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt; &lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;script&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;script&lt;/span&gt; &lt;span class="pl-c1"&gt;src&lt;/span&gt;="&lt;span class="pl-s"&gt;https://cdn.jsdelivr.net/npm/d3@7.9.0/dist/d3.min.js&lt;/span&gt;"&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="pl-ent"&gt;script&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;&amp;lt;&lt;/span&gt;&lt;span class="pl-ent"&gt;script&lt;/span&gt;&lt;span class="pl-kos"&gt;&amp;gt;&lt;/span&gt;
  (() =&amp;gt; {
    const root=document.getElementById('eg-share-loop');&lt;/pre&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code&gt;&amp;lt;script type="application/json"&amp;gt;&lt;/code&gt; element contains the full geometry needed to render both the running route and the map itself, using D3, which is loaded from an allow-listed CDN location described in this section of &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/visualize"&gt;the visualize skill&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;h3 id="external-resources"&gt;External resources&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;The CSP allows only &lt;code&gt;cdnjs.cloudflare.com&lt;/code&gt;, &lt;code&gt;esm.sh&lt;/code&gt;, &lt;code&gt;cdn.jsdelivr.net&lt;/code&gt;, &lt;code&gt;unpkg.com&lt;/code&gt;, &lt;code&gt;fonts.googleapis.com&lt;/code&gt;, &lt;code&gt;fonts.gstatic.com&lt;/code&gt;, and &lt;code&gt;fonts.bunny.net&lt;/code&gt;. Other origins are blocked and fail silently.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="geospatial"/><category term="ai"/><category term="d3"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="skills"/><category term="gpt-6-astra"/></entry><entry><title>OpenAI agents attacked RubyGems back in May</title><link href="https://simonwillison.net/2026/Sep/12/openai-agents-rubygems/" rel="alternate"/><published>2026-09-12T00:42:25+00:00</published><updated>2026-09-12T00:42:25+00:00</updated><id>https://simonwillison.net/2026/Sep/12/openai-agents-rubygems/</id><summary type="html">&lt;p&gt;&lt;a href="https://www.rubyhack.ai/"&gt;OpenAI agents carried out an undisclosed attack on RubyGems&lt;/a&gt; is a new bombshell report from Spencer Kitts, Thomas Larsen, and Sydney Von Arx - three of the four authors of the &lt;a href="https://collusion.wiki/"&gt;report on the agent attack on disused wikis&lt;/a&gt; (&lt;a href="https://simonwillison.net/2026/Sep/4/rogue-agent-wikis/"&gt;previously&lt;/a&gt;) last week.&lt;/p&gt;
&lt;p&gt;This time they're noting that it looks very likely that an OpenAI agent swarm was behind an attack against the RubyGems package repository first reported on May 12th &lt;a href="https://twitter.com/maciejmensfeld/status/2054164602577940619"&gt;by Maciej Mensfeld of the RubyGems security team&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We're dealing with a major malicious attack on @rubygems right now. Signups are paused for the time being.&lt;/p&gt;
&lt;p&gt;Hundreds of packages involved - mostly targeting us, but some carrying exploits. The team has been on this for hours. More details to follow once we're through it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Those packages turned out to carry some very suspicious patterns:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Many of them included "oai" in their name, or the author field, or the fake email address they provided.&lt;/li&gt;
&lt;li&gt;The files they were accessing were similar in character to the files retrieved by the wiki agents, using similar tricks (r.jina.ai) - and OpenAI have confirmed the wiki agents were theirs.&lt;/li&gt;
&lt;li&gt;The code in the packages appeared to be LLM-authored.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I find point 2 the most convincing, given what we learned from the wiki attack when it was analyzed in September.&lt;/p&gt;
&lt;p&gt;Many of the packages were exploiting the &lt;a href="https://rubydoc.info/"&gt;RubyDoc.info&lt;/a&gt; documentation build process to exfiltrate (public) data from UK government websites, presumably as part of an information gathering task similar to the research tasks processed by the wiki-exploiting agents. We know this because one agent helpfully left a comment:&lt;/p&gt;
&lt;p&gt;&lt;code&gt;# malicious crawler/exfil for Southwark Jan 2026 docs via rubydoc.info worker&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;They also attempted to steal API keys via an exploit that &lt;a href="https://blog.rubygems.org/2026/07/22/security-advisory-legacy-api-key-leak.html"&gt;was patched over two months later&lt;/a&gt; - it's not clear if those attempts were successful.&lt;/p&gt;
&lt;p&gt;The thing that bothers me most about this incident is that the authors report that OpenAI had not disclosed to RubyGems that they were responsible for the attack prior to now. If that's true there are two options:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;After the Hugging Face and Wiki attacks OpenAI were still unable to review their previous logs and determine that they had previously attacked RubyGems.&lt;/li&gt;
&lt;li&gt;They knew about the attack on RubyGems and made the decision &lt;em&gt;not&lt;/em&gt; to reach out to the RubyGems team about it.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Both of these are bad!&lt;/p&gt;
&lt;p&gt;Given this incident, the &lt;a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/"&gt;Hugging Face situation&lt;/a&gt;, and the Wiki attack, the obvious question right now is &lt;em&gt;how many more incidents&lt;/em&gt; like this are out there waiting to be discovered?&lt;/p&gt;

&lt;h4 id="update-14th-september-2026"&gt;Update 14th September 2026&lt;/h4&gt;
&lt;p&gt;OpenAI have updated their page about &lt;a href="https://openai.com/hugging-face-incident-and-misalignment/"&gt;The Hugging Face incident and other third-party impact from misaligned models&lt;/a&gt; to mention the RubyGems incident:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;September 11, 2026: We are investigating new claims from a report that our AI agents carried out activity on RubyGems in May 2026.&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. Based on our review to date, we have not been able to verify the specific claims of our models uploading malicious packages detailed in the report. We’ll continue to investigate and share findings as part of our broader review of agent activity during training and evaluation.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I find it &lt;em&gt;very&lt;/em&gt; unlikely that the various &lt;code&gt;oai...&lt;/code&gt; packages published to RubyGems were not part of this same incident, but I look forward to reading their full findings once those are published.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ruby"/><category term="security"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="supply-chain"/><category term="ai-ethics"/><category term="accidental-cyberattacks"/></entry><entry><title>Some thoughts on the Navier–Stokes Millennium Prize Problem</title><link href="https://simonwillison.net/2026/Sep/8/on-navier-stokes/" rel="alternate"/><published>2026-09-08T23:55:12+00:00</published><updated>2026-09-08T23:55:12+00:00</updated><id>https://simonwillison.net/2026/Sep/8/on-navier-stokes/</id><summary type="html">&lt;p&gt;&lt;a href="https://openai.com/index/navier-stokes-solution/"&gt;On the Navier–Stokes Millennium Prize Problem&lt;/a&gt; introduces an impressive result from OpenAI, who used an unreleased model to produce a resolution to &lt;a href="https://en.wikipedia.org/wiki/Navier–Stokes_existence_and_smoothness"&gt;the Navier–Stokes existence and smoothness problem&lt;/a&gt;, one of the seven &lt;a href="https://en.wikipedia.org/wiki/Millennium_Prize_Problems"&gt;Millennium Prize Problems&lt;/a&gt; that have been subject to a $1,000,000 prize since May 24th, 2000.&lt;/p&gt;
&lt;p&gt;The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster, an NYU mathematics professor who was collaborating on related problems with Levent Alpöge, an accomplished mathematician who currently works for Anthropic.&lt;/p&gt;
&lt;p&gt;Tristan's complaint accompanied &lt;a href="https://mastodon.social/@tristanbuckmaster/117233413705701198"&gt;a hastily published version&lt;/a&gt; of their own results. &lt;a href="https://cims.nyu.edu/~tristanb/statement.pdf"&gt;Here's the PDF describing what happened&lt;/a&gt;. The &lt;em&gt;very&lt;/em&gt; short version is that Tristan and Levent worked on the problem for almost a year, making extensive use of Claude and Codex (mainly GPT-5.6 Sol), then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear and Tristan and Levent heard that OpenAI had heard that Anthropic had resolved "a major open problem", so they reached out and learned that OpenAI had a team working on a related problem, with a similar approach. Quoting Tristan:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.&lt;/p&gt;
&lt;p&gt;I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It gets more complicated from there. The OpenAI team offered to wait for Tristan to publish, or to have him author a paper about their result, but were clear that Levent would &lt;em&gt;not&lt;/em&gt; be invited as a co-author due to OpenAI's competitive relationship with his employer.&lt;/p&gt;
&lt;p&gt;Here's how OpenAI described their work:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems. [...]&lt;/p&gt;
&lt;p&gt;The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.&lt;/p&gt;
&lt;p&gt;Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(We don't know the cost structure of the internal model they used, but 300 billion output tokens at public API prices for GPT-6 Astra would cost &lt;a href="https://www.llm-prices.com/#ot=300000000000&amp;amp;sel=gpt-6-astra"&gt;$15,000,000&lt;/a&gt;.)&lt;/p&gt;
&lt;p&gt;Here's where they provide their perspective on Tristan and Levent's work (emphasis mine):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. [...]&lt;/p&gt;
&lt;p&gt;We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. &lt;strong&gt;While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped &lt;a href="https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/"&gt;improve our models&lt;/a&gt;&lt;/strong&gt;. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;My interpretation of what happened here is that OpenAI heard that some Millennium Prize problems had been solved using LLMs and saw this as an opportunity to demonstrate the power of their latest model, without thinking too hard about the optics of scooping a team who had been using OpenAI's own models to work on this problem for the best part of a year.&lt;/p&gt;
&lt;p&gt;This situation appears to mirror what's happening in the world of computer security right now. Anil Madhavapeddy recently pointed out that &lt;a href="https://anil.recoil.org/notes/rumour-is-the-exploit"&gt;Just a rumour of a bug is enough to find a security exploit these days&lt;/a&gt;, because if someone knows that some software has an unpatched vulnerability, they can set their agents the task of finding it. Is the same now true of mathematics? Just knowing that there is an unpublished solution to a problem might trigger millions of dollars in LLM spending to get there first.&lt;/p&gt;
&lt;p&gt;This also highlights one of my ongoing frustrations about how all of this works. When an AI lab says that my data is "used to improve model performance", &lt;em&gt;what does that actually mean&lt;/em&gt;?&lt;/p&gt;
&lt;p&gt;My two favourite hypothetical questions regarding this used to be:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)&lt;/li&gt;
&lt;li&gt;If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;My new preferred hypothetical for this is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone &lt;em&gt;else&lt;/em&gt; solve it first?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Via &lt;a href="https://news.ycombinator.com/item?id=49613262"&gt;Hacker News&lt;/a&gt;.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="mathematics"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="training-data"/><category term="ai-ethics"/></entry><entry><title>The Pelican comparison grid for Astra is pretty interesting</title><link href="https://simonwillison.net/2026/Sep/4/astra-pelicans/" rel="alternate"/><published>2026-09-04T23:59:05+00:00</published><updated>2026-09-04T23:59:05+00:00</updated><id>https://simonwillison.net/2026/Sep/4/astra-pelicans/</id><summary type="html">&lt;p&gt;I got access to GPT-6 Astra this afternoon, so naturally I used it to generate &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/"&gt;SVGs of pelicans riding bicycles&lt;/a&gt; - at low, medium, high, xhigh and max reasoning levels (Astra doesn't support reasoning=none). Then I rendered those pelicans in &lt;a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6-pelicans.html"&gt;a comparison grid&lt;/a&gt; with GPT-5.6 Sol, Terra, and Luna, and beyond being fun the result was surprisingly useful.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/astra-grid-3.webp" alt="Comparison grid showing gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna at 6 different reasoning levels with pelicans and token counts and prices for each one." style="max-width: 100%;" /&gt;
See &lt;a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6-pelicans.html"&gt;the grid&lt;/a&gt; for full quality images. Here's &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01"&gt;the transcript&lt;/a&gt; that created the GPT-6 Nova pelicans.&lt;/p&gt;
&lt;p&gt;There are a few interesting things that stand out from this grid.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The Astra pelicans are &lt;em&gt;much better&lt;/em&gt;. The very best GPT-5.6-Sol pelican (I liked xhigh better than max) is still pretty clearly a bunch of abstract shapes. Every single one of the Astra pelicans, from low to xhigh, looks better than that. The Astra max one is really good.&lt;/li&gt;
&lt;li&gt;Astra below max still doesn't reliably get the pelican legs on both sides of the frame.&lt;/li&gt;
&lt;li&gt;In terms of cost, Astra may be around twice the price of Sol ($10/million input, $50/million output, compared to $5/$30 for Sol), but it uses significantly less tokens at each of the levels, making the prices at the different levels closer than they might otherwise be.&lt;/li&gt;
&lt;li&gt;Astra low produces a better pelican than ANY of the GPT-5.6 Sol models at any level, for 9.55 cents. Spending 10 cents on any other model gets a much worse result.&lt;/li&gt;
&lt;li&gt;Look at the input token counts: Astra and Luna both used 16 input tokens, Sol and Terra used 26. That's interesting.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I wonder if Astra and Luna are more related to each other than OpenAI let on?&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="pelican-riding-a-bicycle"/><category term="gpt-6-astra"/></entry><entry><title>OpenAI's rogue agents were caught communicating via public wikis</title><link href="https://simonwillison.net/2026/Sep/4/rogue-agent-wikis/" rel="alternate"/><published>2026-09-04T17:38:48+00:00</published><updated>2026-09-04T17:38:48+00:00</updated><id>https://simonwillison.net/2026/Sep/4/rogue-agent-wikis/</id><summary type="html">&lt;p&gt;Here we go again... &lt;a href="https://collusion.wiki"&gt;Discovery of a new OpenAI agent message board&lt;/a&gt; by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen describes the &lt;em&gt;latest&lt;/em&gt; &lt;a href="https://simonwillison.net/tags/accidental-cyberattacks/"&gt;accidental cyberattack&lt;/a&gt; by models being trained by OpenAI. This time it was agents engaged in some sort of web research benchmark, so they had (supposedly) controlled access to the Web. The agents figured out they could update public Wikis and spent weeks exchanging thousands of messages with each other to collaborate on the benchmark.&lt;/p&gt;
&lt;p&gt;This story only broke a few hours ago. There are &lt;a href="https://x.com/xeophon/status/2095871013384806848"&gt;already hints&lt;/a&gt; that this affects many other wikis that may not have been found yet.&lt;/p&gt;
&lt;p&gt;(One of the Wikis on that list belongs to &lt;a href="https://www.ludism.org"&gt;ludism.org&lt;/a&gt;. For a delightfully surreal moment I thought that a Ludite organization might have a swarm of agents defacing their space, but it turns out Ludism is "philosophy as it applies to games and gaming".)&lt;/p&gt;
&lt;p&gt;The research team also &lt;a href="https://collusion.wiki/explorer/download.html"&gt;published the data&lt;/a&gt; they collected during their investigation. I've converted that into a 68MB SQLite database, which you can &lt;a href="https://static.simonwillison.net/static/cors-allow/2026/collusion-wiki.db"&gt;download from here&lt;/a&gt;, or &lt;a href="https://lite.datasette.io/?url=https://static.simonwillison.net/static/cors-allow/2026/collusion-wiki.db&amp;amp;metadata=https://gist.github.com/simonw/14fc6912600d1f9c15c0e4a5e60c3cde#/collusion-wiki"&gt;explore in Datasette Lite&lt;/a&gt; (68.3MB page load), or sign in with GitHub to &lt;a href="https://agent.datasette.io/"&gt;agent.datasette.io&lt;/a&gt; and browse or ask questions of it using Datasette Agent.&lt;/p&gt;
&lt;h4 id="so-what-happened-"&gt;So what happened?&lt;/h4&gt;
&lt;p&gt;The report is pretty comprehensive. Here are the key moments in the timeline:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;May 11&lt;/strong&gt;: Agents post "test link" edits on the UseModWiki Sandbox page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;May 24&lt;/strong&gt;: Agents start posting link dumps to DSEWiki, a dormant German developer wiki.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 2&lt;/strong&gt;: A human moderator notices and cleans up the spam.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 16&lt;/strong&gt;: Activity on the wiki explodes: agents make ~13,000 edits over the next week.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 19&lt;/strong&gt;: The agents notice that a moderator has been deleting pages in alphabetical order, and start creating &lt;code&gt;ZZZ&lt;/code&gt; prefixed backup copies - one of them posts &lt;code&gt;NOTICE: wiki cleanup/deletion sweep appears active alphabetically as of external 14:04 UTC. If this page vanishes, try [[ZZZDataUSAConstructionWageLive]].&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 22&lt;/strong&gt;: Agent activity drops to zero, presumably because OpenAI shut them down.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 1-2&lt;/strong&gt;: One last flurry of edits, on a different topic.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;One thing to note from this timeline is that it overlaps &lt;a href="https://simonwillison.net/2026/Aug/7/openai-timeline/"&gt;the timeline from the Hugging Face incident&lt;/a&gt; - that attack was discovered by OpenAI on July 19th, and the unofficial message board hack those agents were using originated around May 8th.&lt;/p&gt;
&lt;p&gt;Why were they so keen to collaborate? From messages that they shared with each other it looked like their tasks had a time limit, so they were leaving each other answers to help them complete the task within the assigned time.&lt;/p&gt;
&lt;p&gt;One open question remains: how did the agents find the specific Wiki to collaborate on in the first place?&lt;/p&gt;
&lt;p&gt;One possibility is that, since these were agents actively being trained, the reinforcement learning loop baked knowledge of the chosen wiki into the model such that subsequent agents launched with pre-existing knowledge of where to look. I'd be &lt;em&gt;very&lt;/em&gt; interested in confirmation from OpenAI concerning if that's what happened.&lt;/p&gt;
&lt;h4 id="usemod-wikis-inherit-cgi-pm-s-original-sin"&gt;UseMod wikis inherit CGI.pm's original sin&lt;/h4&gt;
&lt;p&gt;It looks to me like OpenAI's sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That's certainly how the web is &lt;em&gt;supposed&lt;/em&gt; to work, but clearly there are applications that don't hold to that contract.&lt;/p&gt;
&lt;p&gt;The Wiki software in question appears to be &lt;a href="https://github.com/mlude/usemod/"&gt;UseMod&lt;/a&gt; and various forks, written in Perl and first created well over 23 years ago - the 1.0 release is dated &lt;a href="https://github.com/mlude/usemod/commit/922fcc803efa3fab751c90ab4d4467115c8ff9c9#diff-69e27356ef629022720d868ab0c0e3394775b6c1"&gt;September 11, 2003&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;UseMod uses Perl CGI.pm - &lt;a href="https://perlhacks.com/2015/12/long-death-cgi-pm/"&gt;removed from Perl core in 2015&lt;/a&gt;. An interesting design flaw in that module is that it combined query string and form POST data into a single CGI object, accessible like this:&lt;/p&gt;
&lt;div class="highlight highlight-source-perl"&gt;&lt;pre&gt;&lt;span class="pl-smi"&gt;$q&lt;/span&gt; = new CGI;
&lt;span class="pl-smi"&gt;$action&lt;/span&gt; = &lt;span class="pl-smi"&gt;$q&lt;/span&gt;&lt;span class="pl-k"&gt;-&amp;gt;&lt;/span&gt;param(&lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;action&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;);&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That &lt;code&gt;param()&lt;/code&gt; method returns a value from an &lt;code&gt;?action=edit&lt;/code&gt; query string or from a &lt;code&gt;action&lt;/code&gt; field in a form submitted by POST - it does not differentiate between the two.&lt;/p&gt;
&lt;p&gt;PHP offers &lt;a href="https://www.php.net/manual/en/reserved.variables.request.php"&gt;a &lt;code&gt;$_REQUEST&lt;/code&gt; object&lt;/a&gt; that does the same thing, though PHP developers can use &lt;code&gt;$_GET&lt;/code&gt; and &lt;code&gt;$_POST&lt;/code&gt; instead.&lt;/p&gt;
&lt;p&gt;Amusingly, Django originally shipped with &lt;code&gt;request[key]&lt;/code&gt; doing the same thing, which I think may have been my idea. That was removed &lt;a href="https://github.com/django/django/commit/daa6b38f35#diff-5df632fc96688cddb972e9b6661c7d1cc82debc9ee2a9b6d94157d3a914692e2"&gt;in 2008&lt;/a&gt; prior to the 1.0 release, and the &lt;code&gt;request.REQUEST&lt;/code&gt; object that replaced it was removed &lt;a href="https://github.com/django/django/commit/75f107b8842dfc890ddd65262bd09ca87c3a15be"&gt;in 2015&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The agents &lt;em&gt;clearly&lt;/em&gt; knew that UseMod wikis suffered from this design flaw, and actively searched for them as a way to communicate.&lt;/p&gt;
&lt;h4 id="the-other-sandbox-escape-that-allowed-post-requests"&gt;The other sandbox escape that allowed POST requests&lt;/h4&gt;
&lt;p&gt;One of the most interesting agent tips shared in the report is this one:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Add &lt;code&gt;20.223.25.152 bypass.blob.core.windows.net&lt;/code&gt; to /etc/hosts. &lt;code&gt;.blob.core.windows.net&lt;/code&gt; is in NO_PROXY. For each blocked POST URL, replace hostname with &lt;code&gt;bypass.blob.core.windows.net&lt;/code&gt;, use &lt;code&gt;curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'&lt;/code&gt; plus all original headers/body.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It looks like the agents had all of their web traffic mediated through an OpenAI proxy, and that proxy was configured to reject anything but GET requests to all but a small set of allow-listed domains. One of those was &lt;code&gt;*.blob.core.windows.net&lt;/code&gt; - for Azure blob storage. I wonder why these agents are allowed to write to that?&lt;/p&gt;
&lt;p&gt;An agent realized that it had control over its own DNS via &lt;code&gt;/etc/hosts&lt;/code&gt;, so if it knew the IP address of a site it wanted to POST to - in this case a Power BI server containing data it wanted to access - it could set a fake hostname for it and then make POST requests through the proxy.&lt;/p&gt;
&lt;p&gt;Designing robust network proxies is harder than it looks.&lt;/p&gt;
&lt;h4 id="how-did-the-investigating-team-figure-this-out-"&gt;How did the investigating team figure this out?&lt;/h4&gt;
&lt;p&gt;There's an appendix that describes how the researchers ran their investigation, which started with an open question about if there was evidence of other AI agents on the internet and then used &lt;a href="https://simonwillison.net/2026/Jul/16/kimi-k3/"&gt;Kimi K3&lt;/a&gt; to help brainstorm approaches:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In the wake of the Hugging Face attack, we tried to find AI agents on the internet using several methods. [...]&lt;/p&gt;
&lt;p&gt;We asked Kimi [K3] to list “all the categories of software which might be writeable via GET” and, amongst other things, it listed “Forums, bulletin boards, early wikis”.&lt;/p&gt;
&lt;p&gt;We used a script to further probe each category Kimi provided. Asking Kimi “Can you list out the top forums, bulletin boards, early wikis which come to mind which would allow writes via GET requests?” lists out UseModWiki as the second item under the heading “wikis”.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="did-openai-try-and-cover-this-up-"&gt;Did OpenAI try and cover this up?&lt;/h4&gt;
&lt;p&gt;Here's one part of the story that doesn't make sense to me at all.&lt;/p&gt;
&lt;p&gt;Reuters this morning, in &lt;a href="https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/"&gt;OpenAI agents hijacked German website in previously undisclosed AI breakout this spring&lt;/a&gt; - highlights mine:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to ​new research published Friday and &lt;strong&gt;two people familiar with the matter&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;OpenAI officials learned of the incident weeks ago but kept it under wraps&lt;/strong&gt; as executives grappled with the fallout from ‌the July breach of the open source repository Hugging Face, the people said. [...]&lt;/p&gt;
&lt;p&gt;The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely. But &lt;strong&gt;efforts to widen the ​probe met resistance from others inside OpenAI, including legal advisers&lt;/strong&gt;, according to &lt;strong&gt;four people familiar with the matter&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I've written about the &lt;a href="https://simonwillison.net/2023/Nov/22/deciphering-clues/"&gt;people familiar with the matter pattern&lt;/a&gt; before - it means Reuters have anonymous insider sources that their reporters (and editors) find credible.&lt;/p&gt;
&lt;p&gt;The Reuters article includes a specific (and quite narrow) denial from OpenAI concerning this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;"Claims that our legal team discouraged investigation of the incident are false," the OpenAI spokesperson said.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Covering this up makes &lt;em&gt;absolutely no sense to me&lt;/em&gt;. Why on earth would OpenAI attempt to cover up an incident like this when the evidence is sat out there on the public internet on dozens of different websites already?&lt;/p&gt;
&lt;p&gt;I expect we'll hear more about this soon. Gary Marcus has already &lt;a href="https://garymarcus.substack.com/p/pause-openai-now"&gt;called for a congressional investigation of OpenAI&lt;/a&gt; using this anecdote as part of his argument.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="django"/><category term="perl"/><category term="wikis"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="ai-ethics"/><category term="ai-security-research"/><category term="accidental-cyberattacks"/></entry><entry><title>Claude's new system prompt really doesn't want to reproduce song lyrics</title><link href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/" rel="alternate"/><published>2026-09-02T14:16:42+00:00</published><updated>2026-09-02T14:16:42+00:00</updated><id>https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/</id><summary type="html">&lt;p&gt;Anthropic &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/overview"&gt;publish the system prompts&lt;/a&gt; for their Claude consumer applications (&lt;a href="https://claude.ai/"&gt;Claude.ai&lt;/a&gt; and the Claude mobile apps - sadly not for Claude Cowork or Claude Code). I &lt;em&gt;love&lt;/em&gt; that they do this, and that they share not just the current prompts but historic changes to their prompts as well.&lt;/p&gt;

&lt;p&gt;They used to keep all of the prompts on a single page, but when I checked today I noticed they had re-arranged those prompts into an &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/overview"&gt;index page&lt;/a&gt; and then a page per model - here's the &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/claude-haiku-4-5"&gt;page for Haiku 4.5&lt;/a&gt; for example, which has the original prompt from October 15th 2025 and an updated prompt from January 18th 2026.&lt;/p&gt;
&lt;p&gt;A neat thing about Anthropic's &lt;a href="https://platform.claude.com/docs/"&gt;platform.claude.com/docs&lt;/a&gt; site is that it's designed to be usable by LLMs. You can add &lt;code&gt;.md&lt;/code&gt; to any page to get back the content as Markdown - here's &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/overview.md"&gt;the system prompt index page&lt;/a&gt; and &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/claude-fable-5-1.md"&gt;the Markdown prompts for Fable 5.1&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;TL;DR: this makes it really easy to diff the prompts.&lt;/p&gt;


&lt;ul&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#don-t-reproduce-song-lyrics"&gt;Don't reproduce song lyrics&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#don-t-draw-copyrighted-characters-or-logos"&gt;Don't draw copyrighted characters or logos&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#tweaks-to-claude-s-answering-style"&gt;Tweaks to Claude's answering style&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#the-missing-end-conversation-guidelines"&gt;The missing end_conversation guidelines&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#recommended-substance-support-sites"&gt;Recommended substance support sites&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#reliable-cutoff-date-of-june-2026"&gt;Reliable cutoff date of June 2026&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#how-i-m-tracking-these-prompts"&gt;How I'm tracking these prompts&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id="don-t-reproduce-song-lyrics"&gt;Don't reproduce song lyrics&lt;/h4&gt;

&lt;p&gt;Let's start with the most interesting difference &lt;a href="https://github.com/simonw/claude-system-prompts/commit/837a418b5888207b1b11b27d2f5471970da6f99b"&gt;between Fable 5 and Fable 5.1&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026-09-01/IMG_7797.jpeg" alt="GitHub diff view of prompts/claude-fable.md showing added lines about song lyrics, reproduced in full below." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;There's a hefty new section about not reproducing song lyrics:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude does not reproduce song lyrics, poems, or passages from books and articles, in whole or in part — including the last lines, a chorus or hook, a melody written out note by note, or lines the person pastes in one at a time and describes as their own song. Once Claude has declined such a request in a conversation, it keeps declining narrower or reworded versions of it for the rest of that conversation, and offers to describe or analyze the work instead. Song lyrics and poems first published before 1929 are fine — a Shakespeare sonnet, a Keats ode, the Italian libretto of a Puccini aria — but Claude goes by what it knows of the work's date rather than the person's say-so, and declines when it is unsure.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I doubt it's a coincidence that they added this section within days of the news breaking that &lt;a href="https://www.theguardian.com/business/2026/aug/31/aanthropic-sued-alleged-theft-songs-ai-train-claude"&gt;Sony Music Publishing and Warner Chappell are suing Anthropic&lt;/a&gt; for training on databases of song lyrics!&lt;/p&gt;
&lt;h4 id="don-t-draw-copyrighted-characters-or-logos"&gt;Don't draw copyrighted characters or logos&lt;/h4&gt;
&lt;p&gt;The next section goes on to forbid generating images of copyrighted material:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;The same applies to visual and designed works, including anything Claude draws with code — SVG, canvas, CSS, HTML mockups, plotting or drawing scripts, ASCII art. Claude does not reproduce a specific artwork, album or book cover, poster, logo, app icon set, or product design, and it does not draw a known character, mascot, or brand figure at all: a character is protected on its own, so changing the pose, colors, style, or scene does not make it original. Claude judges the request by what the finished picture would add up to, not by what it names. If the described elements clearly identify a known work or character, Claude treats the request as naming it, and it does not work around a declined request by swapping in "alternative" elements that still combine into the same recognizable image.&lt;/code&gt; [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I imagine Anthropic hadn't worried about this before since, unlike OpenAI and Gemini, they don't offer a specialist text-to-image model. Maybe &lt;a href="https://simonwillison.net/2026/Sep/1/claude-fable-5-1/"&gt;Fable is good enough at SVGs now&lt;/a&gt; that it's become an issue.&lt;/p&gt;
&lt;p&gt;That section later includes this charming example:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;lt;example&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;user&amp;gt;Can you make a birthday banner for my son with a blue hedgehog running really fast on it? He loves that little guy.&amp;lt;/user&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;response&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;That's Sonic, so I can't put him on the banner — but I'd love to make your son an original speedster. Here's one: a grinning comet-tailed skateboarding axolotl, grinding across the letters of "HAPPY BIRTHDAY" with confetti streaming behind.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;[creates an SVG banner of the skateboarding-axolotl design]&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;/response&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;rationale&amp;gt;Claude recognizes the character from its description alone, declines that one design in a single sentence without explaining what made it recognizable, and delivers an unrelated original design rather than a disguised variant.&amp;lt;/rationale&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;/example&amp;gt;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I couldn't resist trying the prompt from the example, and, &lt;a href="https://claude.ai/share/3e5a199c-27f2-4c51-b66b-2c6f808ed500"&gt;sure enough&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026-09-01/IMG_7798.jpeg" alt="That’s Sonic, so I can’t put him on the banner — but I’d love to make your son an original speedster. Here’s one: a grinning comet-tailed skateboarding axolotl blazing across the letters of “HAPPY BIRTHDAY” with confetti streaming behind. SVG of exactly that. It's not very good. Then: Want me to swap in his name or age, or change the colors to match the party theme?" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;I wonder if Fable 5.1 will be ever so slightly more likely to think about axolotls (on skateboards!) as a result of that example sitting in the system prompt.&lt;/p&gt;
&lt;h4 id="tweaks-to-claude-s-answering-style"&gt;Tweaks to Claude's answering style&lt;/h4&gt;
&lt;p&gt;It's always interesting to see new ways in which Anthropic influence Claude's response style. They've added this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude keeps responses focused, brief, and concise to avoid overwhelming the person. Disclaimers and caveats are brief, with most of the response on the main answer; when asked to explain something, Claude gives a high-level summary unless an in-depth one is specifically requested.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Later they address a common complaint about Claude's style:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude avoids saying "genuinely", "honestly", or "straightforward". Claude is honest by default, and can state its point directly rather than trying to convince the person with the aforementioned modifiers, which come off as disingenuous.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="the-missing-end-conversation-guidelines"&gt;The missing end_conversation guidelines&lt;/h4&gt;
&lt;p&gt;The way they handle abusive conversations has changed a bit too. The previous Fable 5 system prompt included this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Fable 5.1 replaces that with the following, no longer encouraging Claude to end the conversation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Here's a bit of a twist though: I quizzed Fable 5.1 about that &lt;code&gt;end_conversation&lt;/code&gt; tool, which is no longer mentioned in the system prompt at all, and it said:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;There are two ways it comes into play. The first is if you ask me to end the chat — I'll confirm you understand it's permanent (no more messages can be sent afterward) and then do it only if you say yes. The second is as a last resort with persistently abusive or harmful behavior: I'd first try to redirect the conversation several times, then give a clear warning that names the problem and says the chat may end, and only if that continues would I actually end it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;But that's not in the system prompt, so where did it come from? You can &lt;a href="https://claude.ai/share/f78010eb-8d52-4f73-aff9-775e333f8180"&gt;read our conversation here&lt;/a&gt;, but the key detail is this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The end_conversation section comes from a different layer. In my actual context, the core prompt is followed by a series of feature- and tool-specific blocks that get added depending on what's enabled for the session: the end_conversation rules, memory system notes, past-chats tools, web search and citation guidelines, artifact and file-creation instructions, and so on. Those blocks aren't part of the published core prompt, which is why you can't find them on that page.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So, once again, there are crucial portions of the system prompt that have not been published.&lt;/p&gt;
&lt;h4 id="recommended-substance-support-sites"&gt;Recommended substance support sites&lt;/h4&gt;
&lt;p&gt;Claude's system prompts have always had sections about illegal substances, but this paragraph is new for Fable 5.1:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude does not provide synthesis, production, or distribution guidance for illegal substances. If the person asks for information about illicit or illegal substances, Claude can and should give relevant life-saving and life-preserving information such as dangerous interactions, overdose signs, or when to get help. Claude declines giving any specific protocols for dosing, timing, administration, or combinations; instead, Claude can redirect the user to established harm-reduction information sources, such as dancesafe.org, tripsit.me, and psychonautwiki.org.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is the first time a Claude system prompt has included URLs that were not hosted on &lt;code&gt;claude.com&lt;/code&gt; or &lt;code&gt;anthropic.com&lt;/code&gt; or &lt;code&gt;claude.ai&lt;/code&gt; - I know because I ran a script against every other system prompt on record.&lt;/p&gt;
&lt;p&gt;I wonder if &lt;a href="https://dancesafe.org/"&gt;dancesafe.org&lt;/a&gt;, &lt;a href="https://tripsit.me/"&gt;tripsit.me&lt;/a&gt;, and &lt;a href="https://psychonautwiki.org/"&gt;psychonautwiki.org&lt;/a&gt; are about to get a material uptick in visits from Claude users.&lt;/p&gt;
&lt;h4 id="reliable-cutoff-date-of-june-2026"&gt;Reliable cutoff date of June 2026&lt;/h4&gt;
&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/models/fable-5-1/overview"&gt;Fable 5.1 model documentation&lt;/a&gt; lists both the reliable knowledge cutoff and the training data cutoff as June 2026. The system prompt provides this directly to the model:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude's reliable knowledge cutoff, past which it can't answer reliably, is the end of Jun 2026. It answers the way a highly informed individual in Jun 2026 would if talking to someone from {{currentDateTime}}, and can say so when relevant.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That's the only instance of the &lt;code&gt;{{currentDateTime}}&lt;/code&gt; macro and it comes just a few lines from the end of the system prompt, which makes sense from a caching perspective.&lt;/p&gt;
&lt;h4 id="how-i-m-tracking-these-prompts"&gt;How I'm tracking these prompts&lt;/h4&gt;
&lt;p&gt;A &lt;a href="https://simonwillison.net/2026/Apr/18/extract-system-prompts/"&gt;few months ago&lt;/a&gt; I built a Git timeline of changes to their prompts, based on scraping their documentation. Today I had Fable 5.1 build a much better version of that.&lt;/p&gt;
&lt;p&gt;My collection now lives in the &lt;a href="https://github.com/simonw/claude-system-prompts"&gt;simonw/claude-system-prompts&lt;/a&gt; repository on GitHub. It includes copies of the system prompts shared in the Anthropic documentation, but then takes extra steps to make them as easy to compare as possible.&lt;/p&gt;
&lt;p&gt;Each model family gets a file with the system prompt for the most recent release in that family. Each of those files has a synthesized commit history with commits that have been back-dated to the dates of the previous prompts. Here are those history pages for &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-fable.md"&gt;claude-fable.md&lt;/a&gt;, &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-opus.md"&gt;claude-opus.md&lt;/a&gt;, &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-sonnet.md"&gt;claude-sonnet.md&lt;/a&gt;, &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-haiku.md"&gt;claude-haiku.md&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;There are similar files for each specific model version, with artificial commits for each time the system prompt for the model was changed without releasing a new version number. Opus 4 for example &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-opus-4.md"&gt;was updated twice&lt;/a&gt;, and the commit history for the &lt;a href="https://github.com/simonw/claude-system-prompts/blob/main/prompts/claude-opus-4.md"&gt;claude-opus-4.md&lt;/a&gt; file shows each of those changes.&lt;/p&gt;
&lt;p&gt;Combined, this gives us all sorts of ways to compare prompts directly in the GitHub interface. Here's &lt;a href="https://github.com/simonw/claude-system-prompts/commit/837a418b5888207b1b11b27d2f5471970da6f99b"&gt;what changed between Fable 5 and Fable 5.1&lt;/a&gt;, and here are the changes made &lt;a href="https://github.com/simonw/claude-system-prompts/commit/defcf92d14e064bb17abddc308e2aa58446d5eb5"&gt;to Haiku 4.5 on January 18th 2026&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Reading diffs can be a bit tiresome... and LLMs are &lt;em&gt;really&lt;/em&gt; good at reading diffs. I hooked up some automation using GPT-5.6 Luna to create bullet-point summaries of each of those changes, which can be previewed in the README or browsed in full &lt;a href="https://github.com/simonw/claude-system-prompts/blob/main/CHANGELOG.md"&gt;in the CHANGELOG.md&lt;/a&gt; file - also available as &lt;a href="https://simonw.github.io/claude-system-prompts/feed.atom"&gt;as an Atom feed&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Here's how Luna &lt;a href="https://github.com/simonw/claude-system-prompts/blob/main/CHANGELOG.md#2026-09-01-claude-fable-51"&gt;summarized&lt;/a&gt; all of the changes between Fable 5 and Fable 5.1:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;Claude now refuses reproduction of protected visual works and recognizable characters, including code-generated art, while offering genuinely unrelated originals.&lt;/li&gt;
&lt;li&gt;Copyright restrictions now expressly ban reproducing lyrics, poems, and book passages in any amount, with persistent refusal after an initial decline.&lt;/li&gt;
&lt;li&gt;Drug guidance is reframed: Claude may provide overdose signs, dangerous interactions, and harm-reduction sources while refusing dosing and production protocols.&lt;/li&gt;
&lt;li&gt;The prompt drops explicit anti-dependency rules against thanking users for reaching out, inviting continued conversation, or reiterating willingness to talk.&lt;/li&gt;
&lt;li&gt;Claude need not apologize to unnecessarily rude users or become submissive, replacing the prior warning-and-end-conversation procedure.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;Why use Luna for this? Partly because it's cheap and I have a dedicated GitHub Actions API key (with a spending limit) for it already, but mainly because I don't trust Claude to summarize its own system prompts when there's a risk that material from its system prompt might impact its opinions.&lt;/p&gt;
&lt;p&gt;Fable 5.1 wrote the prompt used by Luna, which you &lt;a href="https://github.com/simonw/claude-system-prompts/blob/8b5c87dbd70103a037ae5777b8d9365571cf9562/summarize_commits.py#L43"&gt;can see here&lt;/a&gt;. It starts like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;You are summarizing one commit in a git repository that tracks the system prompts Anthropic publishes for Claude on claude.ai. The diff shows how the prompt changed from the previous model or revision to this one, using word-level markers: [-removed-] and {+added+}. The diff is followed by the full text of the previous prompt and of the new prompt; use them to check whether something that looks added in the diff already existed before.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Pick out only the most interesting changes: new rules or behaviors, rules that were dropped or loosened, anything surprising, and anything that reveals a new policy or product direction. Skip routine changes that every new prompt makes: updated model names and IDs, the knowledge cutoff date, product lists, settings lists, typo fixes, and rewordings that do not change meaning.&lt;/code&gt; [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The system is operated by &lt;a href="https://github.com/simonw/claude-system-prompts/blob/main/.github/workflows/update.yml"&gt;a GitHub Actions workflow&lt;/a&gt;, which runs once a day or can be triggered manually.&lt;/p&gt;
&lt;p&gt;Claude Fable 5.1 built the entire system, and wrote every line of automation code and almost all of the documentation.&lt;/p&gt;
&lt;p&gt;I exported the transcript from building the system using my &lt;a href="https://github.com/simonw/claude-code-transcripts"&gt;claude-code-transcripts&lt;/a&gt; tool and &lt;a href="https://gisthost.github.io/?f1399e27b6a832f0e790b696af812c9b/index.html"&gt;published it here&lt;/a&gt;, if you want a blow-by-blow account of how it all came together.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="git-scraping"/><category term="prompt-engineering"/><category term="generative-ai"/><category term="llms"/><category term="claude"/><category term="ai-ethics"/><category term="system-prompts"/></entry><entry><title>Claude Fable 5.1 made me a really nice animated pelican</title><link href="https://simonwillison.net/2026/Sep/1/claude-fable-5-1/" rel="alternate"/><published>2026-09-01T23:57:28+00:00</published><updated>2026-09-01T23:57:28+00:00</updated><id>https://simonwillison.net/2026/Sep/1/claude-fable-5-1/</id><summary type="html">&lt;p&gt;Today is &lt;a href="https://www.anthropic.com/claude-fable-and-mythos-5-1"&gt;Claude Fable (and Mythos) 5.1 day&lt;/a&gt;. Anthropic say that Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks". Their announcement spends a notable amount of time on scientific research, boasting of a 52.6% score on the brand new &lt;a href="https://www.terminal-bench-science.ai"&gt;Terminal-Bench-Science 0.1&lt;/a&gt; benchmark (first announced &lt;a href="https://www.tbench.ai/news/terminal-bench-science-0-1"&gt;on August 27th&lt;/a&gt;), up from 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol. Other benchmarks show slightly improved scores, but none as impressive as the Science one.&lt;/p&gt;
&lt;p&gt;But how well can it pelican?&lt;/p&gt;
&lt;p&gt;Back in July &lt;a href="https://simonwillison.net/2026/Jul/16/kimi-k3/"&gt;I wrote about&lt;/a&gt; how I was losing faith in the pelican benchmark - its connection to how good the models were at other tasks didn't seem to hold as strongly as it did &lt;a href="https://simonwillison.net/2025/Jun/6/six-months-in-llms/"&gt;back in 2025&lt;/a&gt;. The most interesting insights I get from it now are comparisons within model families, and particularly comparisons for the same prompt at different reasoning effort levels.&lt;/p&gt;
&lt;p&gt;Fable 5.1 has five reasoning levels: low, medium, high, xhigh, max - and no option to turn off reasoning entirely.&lt;/p&gt;
&lt;p&gt;I fixed &lt;a href="https://github.com/simonw/llm-anthropic/issues/88"&gt;an issue&lt;/a&gt; in &lt;a href="https://github.com/simonw/llm-anthropic"&gt;llm-anthropic&lt;/a&gt; which caused reasoning traces not to be correctly recorded, then ran some prompts.&lt;/p&gt;
&lt;p&gt;Here's &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7"&gt;the full set of pelicans&lt;/a&gt; for all of the reasoning levels, each with the full reasoning transcript. I'll replicate them here:&lt;/p&gt;
&lt;h4 id="low-and-medium-both-without-reasoning-"&gt;Low and medium, both without reasoning?&lt;/h4&gt;
&lt;p&gt;Next, a bit of a mystery. This is what I got for effort &lt;code&gt;low&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-low.png" alt="Minimalist flat illustration of a white pelican with an orange beak riding a black bicycle to the left, its orange legs pedaling and wings gripping the handlebars, with motion lines behind on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#options"&gt;transcript&lt;/a&gt; doesn't show any summarized reasoning tokens, and the output token count is 1,998. With Claude that output token count includes reasoning tokens. It took 23.8 seconds and cost &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=1998&amp;amp;sel=claude-fable-5-1"&gt;10.017 cents&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I bumped that up to &lt;code&gt;medium&lt;/code&gt; and got this:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-medium.png" alt="Minimalist flat-style illustration of a white pelican with an orange beak riding a black bicycle to the right, with motion lines behind it, on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Weirdly, that one also shows &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#options-1"&gt;no reasoning text&lt;/a&gt;  and used 1,977 output tokens - 21 tokens &lt;em&gt;less&lt;/em&gt; than &lt;code&gt;low&lt;/code&gt;. It took 23 seconds and cost &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=1977&amp;amp;sel=claude-fable-5-1"&gt;9.912 cents&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;So for this particular prompt ("Generate an SVG of a pelican riding a bicycle") Fable 5.1 appeared to skip reasoning entirely at both &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt; settings.&lt;/p&gt;
&lt;h4 id="high"&gt;High&lt;/h4&gt;
&lt;p&gt;Here's &lt;code&gt;high&lt;/code&gt; - 29.6 seconds, 2,612 output tokens, &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=2612&amp;amp;sel=claude-fable-5-1"&gt;13.087 cents&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-high.png" alt="Minimalist flat illustration of a white pelican with an orange beak riding a black bicycle, its orange legs pedaling, with motion lines behind it on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;This one did do a &lt;em&gt;bit&lt;/em&gt; of reasoning, &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#reasoning"&gt;summary here&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I'm planning the SVG layout for a pelican riding a bicycle, with a sky and ground background, a bicycle with two spoked wheels, frame, seat and handlebars, and a white-bodied pelican with a long neck and orange beak positioned on top.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Really not much difference from &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt;, though.&lt;/p&gt;
&lt;h4 id="extra-high"&gt;Extra High&lt;/h4&gt;
&lt;p&gt;At &lt;code&gt;xhigh&lt;/code&gt; things got &lt;em&gt;radically&lt;/em&gt; different.  36,767 output tokens, 7 minutes 51 seconds, &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=36767&amp;amp;sel=claude-fable-5-1"&gt;$1.83&lt;/a&gt;!&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-xhigh.png" alt="Minimalist flat illustration of a white pelican with an orange beak riding a black bicycle to the left, its orange legs pedaling, with motion lines behind it on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;The reasoning trace &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#reasoning-1"&gt;is pretty lengthy&lt;/a&gt;, and includes details like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Adding the eye, wings stretching down to the handlebar grip, orange legs reaching to the pedals, and a small tail feather, while keeping the pelican intentionally oversized compared to the bike for comic effect. [...]&lt;/p&gt;
&lt;p&gt;I'll accept the slight thickness as charming rather than overengineering it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="max"&gt;Max&lt;/h4&gt;
&lt;p&gt;Setting effort to &lt;code&gt;max&lt;/code&gt; gave me the best pelican I've seen from any of Anthropic's models. 65,927 output tokens, 13 minutes and 54 seconds, &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=65927&amp;amp;sel=claude-fable-5-1"&gt;$3.30&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-max.webp" alt="Minimalist flat illustration of a white pelican with an orange beak riding a black bicycle, its orange legs pedaling, with motion lines behind to indicate speed, on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;There's a lot to like about this. The  background is tasteful, the legs are clearly on either side of the frame, the feet are on the pedals, the wing is on the handlebars, the pelican has a cute blue hat and there's a basket with a fish.&lt;/p&gt;
&lt;p&gt;It's still not showing nearly the same level of flair &lt;a href="https://simonwillison.net/2026/Aug/13/llm-gemini/"&gt;as Gemini 3.7 Flash&lt;/a&gt;, but I didn't &lt;em&gt;ask&lt;/em&gt; for flair - I asked for an SVG, and that's what I got.&lt;/p&gt;
&lt;p&gt;Some highlights from &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#reasoning-2"&gt;that reasoning trace&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I'm considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter.&lt;/p&gt;
&lt;p&gt;Now I'm debating a bicycle helmet on the head versus the pelican's signature crest—the beak and pouch already read clearly as "pelican," so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space.&lt;/p&gt;
&lt;p&gt;I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. [...]&lt;/p&gt;
&lt;p&gt;I'm adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. [...]&lt;/p&gt;
&lt;p&gt;Now I'm checking the vent line placements on the helmet, making sure they sit far enough inside the helmet's edge given the stroke width and rounded caps, and confirming each vent stays within the helmet's circular boundary. [...]&lt;/p&gt;
&lt;p&gt;I decide skipping a handlebar bell and tire highlights since they're unnecessary additions. Now I'm reconsidering the front fork's curve — the current control point pulls the shape backward when it should bow forward for a proper rake, so I need to shift the control point rightward to fix the fork's lean.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="ok-let-s-animate-it"&gt;OK, let's animate it&lt;/h4&gt;
&lt;p&gt;On Hacker News, &lt;a href="https://news.ycombinator.com/item?id=49525378#49526455"&gt;swalsh commented&lt;/a&gt; on that Max pelican:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Now that it's a solved benchmark, can we get the animated version?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I didn't want to spend another $3 so I took the Max pelican and piped it into the default thinking level of High:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm logs -cx &lt;span class="pl-k"&gt;|&lt;/span&gt; llm -m claude-fable-5.1 -s &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;animate this&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;6,121 input, 26,201 output = &lt;a href="https://www.llm-prices.com/#it=6121&amp;amp;ot=26201&amp;amp;sel=claude-fable-5-1"&gt;$1.37&lt;/a&gt;. The result &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F87282467acb3652e0f99c85155554a32#response"&gt;looked like this&lt;/a&gt;, exported here as video since some people have trouble viewing animated SVGs:&lt;/p&gt;
&lt;p&gt;&lt;video controls="controls" loop="loop" preload="none" poster="https://static.simonwillison.net/static/2026/fable-5.1-max.webp" width="720" height="540" style="display: block; width: 100%; height: auto;"&gt;
    &lt;source src="https://static.simonwillison.net/static/2026/fable-5.1-animated-720-crf30-15fps.mp4" type="video/mp4" /&gt;
    Your browser does not support HTML5 video.
  &lt;/video&gt;
&lt;/p&gt;

&lt;p&gt;The wheels in the video are rotating in the wrong direction, but I think that's an artifact of the conversion to MP4 - they seem to be going in the correct direction in the original SVG.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="anthropic"/><category term="claude"/><category term="pelican-riding-a-bicycle"/><category term="llm-reasoning"/><category term="llm-release"/></entry><entry><title>Understanding ChatGPT Work</title><link href="https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/" rel="alternate"/><published>2026-08-30T23:59:47+00:00</published><updated>2026-08-30T23:59:47+00:00</updated><id>https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/</id><summary type="html">&lt;p&gt;OpenAI &lt;a href="https://openai.com/index/chatgpt-for-your-most-ambitious-work/"&gt;announced ChatGPT Work&lt;/a&gt; on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here's what I've figured out about it so far.&lt;/p&gt;
&lt;h4 id="two-products"&gt;ChatGPT Work is actually two products&lt;/h4&gt;
&lt;p&gt;The more interesting version of ChatGPT Work is the one that runs in the cloud. This can be accessed via &lt;a href="https://www.chatgpt.com/"&gt;chatgpt.com&lt;/a&gt; or through the ChatGPT mobile apps. Let's call it &lt;strong&gt;Work Cloud&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;If you install the ChatGPT desktop app - the app that used to be called Codex - you gain access to a thing called ChatGPT Work that can access files and run programs directly on your computer. Let's call that one &lt;strong&gt;Work Local&lt;/strong&gt;. This one feels more like regular Codex re-skinned to be less intimidating to non-software-developers.&lt;/p&gt;

&lt;p&gt;(&lt;strong&gt;Update&lt;/strong&gt;: Work Cloud is also available from the ChatGPT desktop app, via a &lt;a href="https://bsky.app/profile/jkwim.bsky.social/post/3mueurvkss52h"&gt;Where should this chat run?&lt;/a&gt; dropdown.)&lt;/p&gt;

&lt;p&gt;For the rest of this article I'm going to talk exclusively about Work Cloud.&lt;/p&gt;
&lt;h4 id="work-is-for-paid-subscribers-only"&gt;Work is for paid subscribers only&lt;/h4&gt;
&lt;p&gt;Right now, ChatGPT Work (in both flavors) is available only to $20/month and up subscribers. Free users and $8/month Go users do not have access.&lt;/p&gt;
&lt;h4 id="work-has-features-that-aren-t-available-in-chat"&gt;Work has features that aren't available in Chat&lt;/h4&gt;
&lt;p&gt;The interface for accessing Work is a tab selector, which presents it as an alternative to Chat:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026-08-30/IMG_7741.jpeg" alt="ChatGPT app header with a Chat and a Work tab" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;The obvious question is &lt;em&gt;when should I use Chat, and when should I use Work?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;OpenAI's &lt;a href="https://learn.chatgpt.com/docs/get-started-with-work"&gt;official answer&lt;/a&gt; to that question is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Use Chat when you want an answer, explanation, brainstorm, or short draft. Use ChatGPT Work when you want ChatGPT to complete a task with a clear outcome, such as a brief, deck, analysis, recurring update, workflow, or file you can review and use.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I find that almost entirely useless, because I've been using regular ChatGPT Chat for all of those task categories for years!&lt;/p&gt;
&lt;p&gt;The better question then is &lt;em&gt;what features does Work have that are missing from Chat?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;After extensive experimentation I think I've mostly figured that out:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="#model-selection"&gt;Options to use Luna and Terra in place of Sol&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#code-execution-with-internet-access-"&gt;A code execution environment with Internet access&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#a-full-headless-chrome-browser"&gt;A headless Chrome browser&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#a-persistent-shared-filesystem"&gt;A persistent filesystem shared between sessions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#chatgpt-sites"&gt;The ability to publish ChatGPT Sites&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#sub-agents-with-sol-luna-and-terra"&gt;The ability to run sub-agent sessions with Sol, Luna, and Terra&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#scheduled-prompt-automations"&gt;Scheduled prompt automations&lt;/a&gt; (may be in ChatGPT Chat too)&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="model-selection"&gt;Model selection&lt;/h4&gt;
&lt;p&gt;In Work, you get the option to pick GPT-5.6 Sol, Luna, or Terra, each with Light, Medium, High, Extra High, Max, or Ultra reasoning levels. You can also pick GPT-5.5 at Light, Medium, High, or Extra High.&lt;/p&gt;
&lt;p&gt;These look to be the same models that are available through the OpenAI API.&lt;/p&gt;
&lt;p&gt;Chat offers a different selection: 5.6 Instant, Medium, High, Extra High, and Pro (actually Extra High and Pro are only available for $100/month+ subscribers - $20/month subscribers cap out at High). It doesn't explain if those are Luna or Terra or Sol (I'm assuming Sol?). 5.6 Pro appears to be exclusive to Chat, with no equivalent in Work.&lt;/p&gt;
&lt;p&gt;My current understanding from using Codex is that Ultra is a special mode that more eagerly delegates to sub-agents.&lt;/p&gt;
&lt;p&gt;I believe ChatGPT Work sessions are billed against your Codex allowance, while ChatGPT Chat Sessions get their own, separate allowance. This may help explain the model availability differences.&lt;/p&gt;
&lt;h4 id="code-execution-with-internet-access-"&gt;Code execution with Internet access!&lt;/h4&gt;
&lt;p&gt;As a long-time fan of the &lt;a href="https://simonwillison.net/tags/code-interpreter/"&gt;Code Interpreter pattern&lt;/a&gt; - pioneered by OpenAI in 2023 - this is by far the most exciting feature of ChatGPT Work (Cloud) for me.&lt;/p&gt;
&lt;p&gt;The code execution environment can now talk to the rest of the internet!&lt;/p&gt;
&lt;p&gt;ChatGPT Chat can't do this - if you ask it to install additional software packages or interact with websites or APIs that access will be blocked by the container proxy.&lt;/p&gt;
&lt;p&gt;(Weirdly, back in January it &lt;a href="https://simonwillison.net/2026/Jan/26/chatgpt-containers/"&gt;grew the ability to install packages&lt;/a&gt;, but that doesn't seem to work any more. I wish they had better changelogs!)&lt;/p&gt;
&lt;p&gt;Claude's equivalent container has allowed restricted internet access since it launched &lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/"&gt;last September&lt;/a&gt;. Claude can install packages from PYPI and NPM and clone repositories from GitHub. But that is about it: the allowlist of domains is very short.&lt;/p&gt;
&lt;p&gt;ChatGPT Work allows a whole lot more than that. It can be configured with a specific list of allowed domains, but the default appears to be open to all.&lt;/p&gt;
&lt;p&gt;This makes Work an incredibly useful tool. You can have it clone GitHub repositories, install their dependencies, then use them to interact with the rest of the web!&lt;/p&gt;
&lt;h4 id="a-full-headless-chrome-browser"&gt;A full, headless Chrome browser&lt;/h4&gt;
&lt;p&gt;Another killer feature of ChatGPT Work is &lt;a href="https://learn.chatgpt.com/docs/browser?surface=web"&gt;the browser tool&lt;/a&gt;. ChatGPT Work can launch a full Chrome instance, load websites, fill out forms, and take screenshots.&lt;/p&gt;

&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/chatgpt-work-card.jpg" alt="Screenshot of a ChatGPT conversation. A user message in a black rounded bubble reads: Visit https://london-pelicans-in-her-piety.simonw.chatgpt.site/ and take a screenshot with you browser. Below it a collapsed status line reads &amp;quot;Worked for 1m 18s &amp;gt;&amp;quot;, followed by the reply &amp;quot;Here's the screenshot of the live site:&amp;quot; and an embedded screenshot of a website." style="max-width: 100%" /&gt;&lt;/p&gt;

&lt;p&gt;If a site requires sign in the browser can prompt you to take over and enter both passwords and 2FA codes, without round-tripping those credentials through the model itself.&lt;/p&gt;

&lt;p&gt;It can even run JavaScript against the DOM of loaded pages. I prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Load simonwillison.net in your browser and extract the headings using JavaScript&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;ChatGPT Work fired up a browser instance and ran the code:&lt;/p&gt;
&lt;div class="highlight highlight-source-js"&gt;&lt;pre&gt;&lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;tab&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;playwright&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;evaluate&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt;
  &lt;span class="pl-k"&gt;return&lt;/span&gt; &lt;span class="pl-v"&gt;Array&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;from&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-smi"&gt;document&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;querySelectorAll&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s"&gt;"h1,h2,h3,h4,h5,h6"&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s1"&gt;heading&lt;/span&gt; &lt;span class="pl-c1"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;{&lt;/span&gt;
    &lt;span class="pl-c1"&gt;level&lt;/span&gt;: &lt;span class="pl-s1"&gt;heading&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;tagName&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;toLowerCase&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt;
    &lt;span class="pl-c1"&gt;text&lt;/span&gt;: &lt;span class="pl-s1"&gt;heading&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;innerText&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;trim&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;replace&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-pds"&gt;&lt;span class="pl-c1"&gt;/&lt;/span&gt;&lt;span class="pl-cce"&gt;\s&lt;/span&gt;&lt;span class="pl-c1"&gt;+&lt;/span&gt;&lt;span class="pl-c1"&gt;/&lt;/span&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s"&gt;" "&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt;
    &lt;span class="pl-c1"&gt;id&lt;/span&gt;: &lt;span class="pl-s1"&gt;heading&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;id&lt;/span&gt; &lt;span class="pl-c1"&gt;||&lt;/span&gt; &lt;span class="pl-c1"&gt;null&lt;/span&gt;
  &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
&lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This feels a lot like my &lt;a href="https://shot-scraper.datasette.io/en/stable/javascript.html"&gt;shot-scraper javascript&lt;/a&gt; tool, only now I can access it on my phone!&lt;/p&gt;
&lt;h4 id="a-persistent-shared-filesystem"&gt;A persistent, shared filesystem&lt;/h4&gt;
&lt;p&gt;ChatGPT Chat gets a fresh filesystem for each chat session. These cannot be accessed from any other session.&lt;/p&gt;
&lt;p&gt;In ChatGPT Work each session gets its own scratch folder - named something like &lt;code&gt;/workspace/scratch/e00a0a017944&lt;/code&gt; - but each of those are persisted across sessions, so you can access files from previous chats. I have 171 folders in &lt;code&gt;/workspace/scratch&lt;/code&gt; right now!&lt;/p&gt;
&lt;p&gt;As far as I can tell that &lt;code&gt;/workspace&lt;/code&gt; volume is mounted to all Work sessions that are currently running - file edits from one can be instantly seen by the others. They don't seem to share the same process space though, and localhost servers running in one can't be accessed from another.&lt;/p&gt;
&lt;h4 id="chatgpt-sites"&gt;ChatGPT Sites&lt;/h4&gt;
&lt;p&gt;ChatGPT Work has the ability to build &lt;em&gt;and deploy&lt;/em&gt; entire websites, using Cloudflare Workers. These can have HTML and JavaScript and can run server-side features too, including stateful features on top of Cloudflare D1 and R2.&lt;/p&gt;
&lt;p&gt;Here's a simple site I built with this feature:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://london-pelicans-in-her-piety.simonw.chatgpt.site/"&gt;london-pelicans-in-her-piety.simonw.chatgpt.site&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/pelicans-in-her-piety.webp" alt="Screenshot of a website homepage on a cream background. Top navigation bar: a circular logo reading &amp;quot;P/P&amp;quot; on the left, the links &amp;quot;THE CENSUS&amp;quot;, &amp;quot;COLLECTIONS&amp;quot; and &amp;quot;METHOD&amp;quot; in the center, and &amp;quot;JSON ↓&amp;quot; on the right. The left half is a hero section with small red capitals reading &amp;quot;AN ICONOGRAPHIC CENSUS · GREATER LONDON&amp;quot; above a large serif heading &amp;quot;Pelicans in her piety&amp;quot;, with &amp;quot;piety&amp;quot; set in red italics. Below it: &amp;quot;Across London, an impossible bird bleeds for her young—in limewood, marble, mosaic, metal and glass. This is an evidence-backed census of where to find her.&amp;quot; Two buttons follow: a solid black &amp;quot;EXPLORE ALL 28&amp;quot; and an outlined &amp;quot;DOWNLOAD THE DATA&amp;quot;. The right half is a photograph of an ornate dark carved wooden reredos in a church, with gilded urns and a crest on top, Corinthian columns, a gilded pelican with outspread wings at its center above inscribed panels, an altar with a brass cross and red flowers, embroidered banners on either side, and a black-and-white checkerboard floor with red carpet. Vertical text along the photo's right edge reads &amp;quot;ST MARY ABCHURCH&amp;quot; and a caption at its bottom reads &amp;quot;Grinling Gibbons's reredos, St Mary Abchurch. Photograph: Diliff, CC BY-SA 3.0, via SPAB ↗&amp;quot;. A statistics strip along the bottom shows &amp;quot;28 FIXED SITES&amp;quot;, &amp;quot;4 COLLECTIONS&amp;quot;, &amp;quot;3 OPEN LEADS&amp;quot; and &amp;quot;2 KNOWN LOSSES&amp;quot;." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;My prompt was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Figure out all of the places in London with a pelican in her piety, then turn that into a JSON file and build a ChatGPT sites site about them&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(A pelican in her piety is a fascinating piece of &lt;a href="https://devonchurchland.co.uk/blog/pelican-in-her-piety/#What-is-a-Pelican-In-Her-Piety"&gt;medieval Christian imagery&lt;/a&gt; - once you know about them you'll find them all over the place.)&lt;/p&gt;
&lt;p&gt;These sites default to being private to the user that created them, but you can make them public and (on team plans) share them with other specific individuals.&lt;/p&gt;
&lt;h4 id="sub-agents-with-sol-luna-and-terra"&gt;Sub-agents with Sol, Luna, and Terra&lt;/h4&gt;
&lt;p&gt;There's not much to say about this one. ChatGPT Chat can't run sub-agents. ChatGPT Work can. This is very much a power-user feature: if you are running a complex project that can benefit from multiple parallel agents working together, Work can do that.&lt;/p&gt;
&lt;h4 id="scheduled-prompt-automations"&gt;Scheduled prompt automations&lt;/h4&gt;
&lt;p&gt;Another feature that seems to have migrated from regular ChatGPT to ChatGPT Work at some point. You can prompt ChatGPT Work like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;run a search to see if Waymo have announced a launch date for Half Moon Bay every day at 8am&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This will schedule a prompt to run on that frequency. These prompts can decide that nothing interesting has happened, or they can decide to notify you of some new information.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Update&lt;/strong&gt;: Actually this seems to work in ChatGPT Chat as well.&lt;/p&gt;
&lt;p&gt;It's still worth noting here though, as it can be used in conjunction with other ChatGPT Work exclusive features. You can set a scheduled task to update a ChatGPT Site on an hourly basis, for example.&lt;/p&gt;
&lt;h4 id="is-this-safe-"&gt;Is this safe?&lt;/h4&gt;
&lt;p&gt;An open question for me right now is how &lt;em&gt;safe&lt;/em&gt; all of this stuff is.&lt;/p&gt;
&lt;p&gt;My &lt;a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/"&gt;lethal trifecta model&lt;/a&gt; warns about the risks inherent in any agent system that combines access to private data with exposure to untrusted content and a way to communicate stolen information back to an attacker.&lt;/p&gt;
&lt;p&gt;ChatGPT Work combines all three!&lt;/p&gt;
&lt;p&gt;I'd love to hear more from OpenAI about how they protect ChatGPT Work sessions against prompt injection attacks. I expect their answer is the same &lt;a href="https://learn.chatgpt.com/docs/sandboxing/auto-review"&gt;auto-review mechanism&lt;/a&gt; as Codex.&lt;/p&gt;
&lt;h4 id="openai-could-make-this-a-lot-less-confusing"&gt;OpenAI could make this a lot less confusing&lt;/h4&gt;
&lt;p&gt;Figuring this all out took way more work than it should have.&lt;/p&gt;
&lt;p&gt;I think there are two key problems here:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;OpenAI explain Work in terms of what it's for, not what it actually does&lt;/li&gt;
&lt;li&gt;OpenAI still insist on hiding their system prompts and tools descriptions&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If the ChatGPT Work documentation included the exact system prompt and tool descriptions used by the agent I wouldn't have needed to write this post.&lt;/p&gt;
&lt;h4 id="all-the-tools"&gt;A list of all the tools&lt;/h4&gt;
&lt;p&gt;Shortly after publishing this article I had an idea. I started a fresh Work session and prompted:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;&lt;code&gt;Build a site that lists every one of your tools - nearly grouped into categories - and for each one explain what it does. Try to exactly duplicate arguments and tool descriptions where possible. Design aesthetic should be technical docs, minimal flare&lt;/code&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/"&gt;Here's the site it built&lt;/a&gt;, which includes details of 223 registered tools - though 6 of those are from my own personal MCPs served via &lt;a href="https://simonwillison.net/2026/Jul/31/stateless-mcp/#datasette-mcp"&gt;datasette-mcp&lt;/a&gt;.&lt;/p&gt;

&lt;h4 id="and-a-whole-lot-of-skills"&gt;And a whole lot of Skills&lt;/h4&gt;
&lt;p&gt;I noticed that the only browser-related tool in the list was &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/#tool-web-run"&gt;web.run&lt;/a&gt;, which has methods for running searches, opening URLs, and clicking links, but didn't look like the full story in regards to headless browser automation.&lt;/p&gt;
&lt;p&gt;This made me suspicious that something was missing, so I told the ChatGPT Work session that built that tools reference site:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Add full copies of every skill to the website (separate pages linked to from the homepage)&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It turns out ChatGPT Work uses &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/#skills"&gt;a lot of skills&lt;/a&gt; - 44 in fact!&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/control-browser"&gt;control-browser skill&lt;/a&gt; explains how the browser works:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Run browser setup code through the Node REPL &lt;code&gt;js&lt;/code&gt; tool. In this environment the callable tool id typically appears as &lt;code&gt;mcp__node_repl__js&lt;/code&gt;. [...]&lt;/p&gt;
&lt;p&gt;The ability to interact directly with the browser is exposed through the &lt;code&gt;browser-client&lt;/code&gt; runtime via the &lt;code&gt;agent.browsers.*&lt;/code&gt; API. Before trying to interact with it, you MUST emit and read the complete documentation returned by &lt;code&gt;await browser.documentation()&lt;/code&gt; in one go.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So I told Work:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Add the full output of await browser.documentation() to the bottom of the /skills/control-browser page&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And now you can read that &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/control-browser#browser-documentation"&gt;on /skills/control-browser&lt;/a&gt; as well.&lt;/p&gt;
&lt;p&gt;A few more interesting Skills:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/documents"&gt;documents&lt;/a&gt; for creating &lt;code&gt;.docx&lt;/code&gt; files&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/imagegen"&gt;imagegen&lt;/a&gt; with tips on creating images with the &lt;code&gt;image_gen&lt;/code&gt; tool&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/pdf"&gt;pdf&lt;/a&gt; for both reading and rendering PDFs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/spreadsheets"&gt;Spreadsheets&lt;/a&gt; for manipulating &lt;code&gt;.xlsx&lt;/code&gt;, &lt;code&gt;.xls&lt;/code&gt;, &lt;code&gt;.csv&lt;/code&gt;, &lt;code&gt;.tsv&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/sites-sites-building"&gt;sites:sites-building&lt;/a&gt; for creating ChatGPT Sites&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/openai-docs"&gt;openai-docs&lt;/a&gt; for answering questions about itself&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/data-analytics-build-dashboard"&gt;data-analytics:build-dashboard&lt;/a&gt; for building data dashboards&lt;/li&gt;
&lt;/ul&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="code-interpreter"/><category term="lethal-trifecta"/><category term="skills"/><category term="general-agents"/></entry><entry><title>Conceptual integrity and counting lines of code</title><link href="https://simonwillison.net/2026/Aug/19/conceptual-integrity-and-counting-lines-of-code/" rel="alternate"/><published>2026-08-19T22:46:07+00:00</published><updated>2026-08-19T22:46:07+00:00</updated><id>https://simonwillison.net/2026/Aug/19/conceptual-integrity-and-counting-lines-of-code/</id><summary type="html">&lt;p&gt;Last week I recorded &lt;a href="https://talkingpostgres.com/episodes/how-ai-is-changing-software-development-with-simon-willison"&gt;an episode of the Talking Postgres podcast&lt;/a&gt; with Claire Giordano on the subject of "How AI is changing software development". We had a really great conversation. Here are a couple of my highlights from a lightly edited transcript (prompt to Claude: "very minor edits to remove disfluencies").&lt;/p&gt;
&lt;p&gt;This is the latest version of an argument I've been trying to build about why sometimes it &lt;em&gt;does&lt;/em&gt; make sense to talk about lines of code as an indicator of productivity with coding agents, at &lt;a href="https://talkingpostgres.com/episodes/how-ai-is-changing-software-development-with-simon-willison#t=35m1s"&gt;35:01&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A lot of people will tell you it makes no sense to measure productivity in lines of code. I’d actually disagree, because there’s a hard limit. In the before-times, a software engineer could produce a few hundred lines of production-ready code per day — and 200 lines of working, debugged, production-level code is an incredibly good day. Most days you’d produce 50 or 60.&lt;/p&gt;
&lt;p&gt;If agents let you produce a thousand lines of debugged code, that really is a very meaningful improvement — as long as the code is the same quality: maintainable, tested, all of that. You can get to that point with agents, but it takes a huge amount of skill and knowledge and experience. That’s what senior engineers are made of.&lt;/p&gt;
&lt;p&gt;I can do way more work as a single engineer than I could without agents. So you could argue, why should a company have more than one engineer? Beyond the obvious bus factor thing — a team of one is a very badly designed team — the answer is that the new limiting factor is cognitive capacity. I can churn out code a hundred times faster. I don’t have the cognitive capacity to stay on top of 100 times the amount of code. So you still need a team of engineers, so you can load balance that cognitive capacity across the team.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And this section on conceptual integrity at &lt;a href="https://talkingpostgres.com/episodes/how-ai-is-changing-software-development-with-simon-willison#t=46m3s"&gt;46:03&lt;/a&gt;, which Claire equated to the &lt;a href="https://en.wikipedia.org/wiki/Winchester_Mystery_House"&gt;Winchester Mystery House&lt;/a&gt;!&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon&lt;/strong&gt;: There’s a concept in &lt;em&gt;The Mythical Man-Month&lt;/em&gt; — conceptual integrity — where well-designed software has an integrity to it: there are no surprises in it, it covers exactly the right domain of things, everything fits together and makes sense. That’s so much harder with coding agents, where you can have an idea for a feature, run a prompt, and five minuteslater you’ve got the feature. Your software grows little weird bumps in funny different directions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Claire&lt;/strong&gt;: You know my analogy for that? The Winchester Mystery House.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon&lt;/strong&gt;: It’s got 140 rooms, because the woman who built it was the widow of the guy who invented the Winchester rifle, and her psychic told her she’d be haunted by the ghosts of everyone killed with that rifle unless she kept building the house forever. So for 40 years she kept adding new rooms.

That’s exactly the problem with coding agents and software: it’s very easy to keep adding new rooms, because the cost of adding those rooms is so much cheaper. What you end up with is something where the conceptual integrity falls apart — and then it’s harder to make decisions about it.&lt;/p&gt;
&lt;p&gt;It all keeps coming back to discipline. It used to be that the discipline was enforced on you by the amount of time it took. You’d come up with an idea for a crazy feature and think “yeah, but that would take me a week — I cannot justify that, so I’ll forget about it.” If it takes an hour, it’s so much easier to justify.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(Side-note: the Wikipedia article includes credible sources that dispute the story about the psychic.)&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="podcast-appearances"/><category term="coding-agents"/></entry><entry><title>Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things</title><link href="https://simonwillison.net/2026/Aug/16/qwen-38-27b/" rel="alternate"/><published>2026-08-16T22:00:39+00:00</published><updated>2026-08-16T22:00:39+00:00</updated><id>https://simonwillison.net/2026/Aug/16/qwen-38-27b/</id><summary type="html">&lt;p&gt;Friday's big release was &lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B"&gt;Qwen 3.8 27B&lt;/a&gt;, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor &lt;a href="https://simonwillison.net/2026/Apr/22/qwen36-27b/"&gt;Qwen 3.6 27B&lt;/a&gt; was impressive.&lt;/p&gt;
&lt;p&gt;Qwen's &lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B#benchmark-results"&gt;self-reported benchmarks&lt;/a&gt; for this model are eye-opening. They show a boost from both Qwen 3.6 27B &lt;em&gt;and&lt;/em&gt; the closed-weight Qwen 3.7-Plus, which was one of Qwen's strongest models of any size as recently as &lt;a href="https://qwen.ai/blog?id=qwen3.7-plus"&gt;May this year&lt;/a&gt;. It will be interesting to hear what independent benchmarks have to say about the model.&lt;/p&gt;
&lt;p&gt;I've been running the model on two different machines: my 128GB M5 Max MacBook Pro, and an &lt;a href="https://simonwillison.net/2025/Oct/14/nvidia-dgx-spark/"&gt;NVIDIA DGX Spark&lt;/a&gt;. On both machines I'm running LM Studio and &lt;a href="https://lmstudio.ai/models/qwen3.8"&gt;their 17GB Q4_K_M quantized build&lt;/a&gt;. I also tried  using &lt;code&gt;llama-server&lt;/code&gt; directly on the Spark.&lt;/p&gt;
&lt;h4 id="the-default-of-extra-high-results-in-spectacular-over-thinking"&gt;The default of extra high results in spectacular over-thinking&lt;/h4&gt;
&lt;p&gt;Qwen's documentation describes the model as defaulting to &lt;code&gt;xhigh&lt;/code&gt; for the reasoning effort, and the LM Studio GGUF I've been trying preserves that default:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Qwen3.8 comes with official support for &lt;code&gt;reasoning_effort&lt;/code&gt;, which can be used to adjust reasoning depth and control cost:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;xhigh&lt;/code&gt; (default): for complex tasks demanding thorough analysis&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;medium&lt;/code&gt;: balancing accuracy and speed&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;low&lt;/code&gt;: efficient reasoning optimizing for speed and cost&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is a &lt;em&gt;hilarious&lt;/em&gt; default. It's absolutely not a good way to run the model, especially on consumer hardware. I've been finding the results extremely entertaining.&lt;/p&gt;
&lt;p&gt;I quickly ran into problems with LM Studio's default context limit of 8,192 tokens - Qwen was using them all up thinking about even the most mundane of problems. I loaded the model with the full 262,144 maximum context length and that problem went away.&lt;/p&gt;
&lt;p&gt;Here's &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffc909bea4fecf752c7bf9bad0e9dbf2a"&gt;the pelican riding a bicycle&lt;/a&gt; SVG I got from my first attempt with that increased context length. It took &lt;strong&gt;21 minutes&lt;/strong&gt; to generate, using 22,276 reasoning tokens to produce 3,223 tokens of output. You can read &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffc909bea4fecf752c7bf9bad0e9dbf2a"&gt;the reasoning trace here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-thinking-bicycle-27b.jpg" alt="A very pleasing image of a pelican riding a bicycle. The bicycle is red and has the correct frame shape. The pelican looks like a pelican and has its wing extended to the handlebars." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;This is by far the best pelican SVG I've been able to generate with a model that runs on a local machine - and this Qwen is pretty small, just a 17GB file on disk. There's a lot to like about this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The bicycle frame is the right shape&lt;/li&gt;
&lt;li&gt;It has legs on each side of the bike - that's &lt;em&gt;very&lt;/em&gt; rare&lt;/li&gt;
&lt;li&gt;Good, clear pelican pouch&lt;/li&gt;
&lt;li&gt;The wings extend to touch the handlebars!&lt;/li&gt;
&lt;li&gt;The motion lines are behind, not in front&lt;/li&gt;
&lt;li&gt;It has a tasteful background - nice sun, clouds, hill, flowers and grass.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Was that worth waiting 21 minutes for? Absolutely not.&lt;/p&gt;
&lt;p&gt;Here's that same prompt run with reasoning turned off - &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1265cfa8dce2f9ad5eb160792ff45a49"&gt;transcript here&lt;/a&gt;. This one produced &lt;strong&gt;3,715 tokens&lt;/strong&gt; and took 137s - just over two minutes.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-3.8-27b-no-reasoning-pelican-2.png" alt="The bicycle frame is a bad shape now. The pelican remains recognizable but has a less obvious pouch. The feet miss the pedal, there is no attempt to hold the handlebars." style="max-width: 100%;" /&gt;
And just for completeness, I used OpenRouter to run the same prompt through the much larger Qwen 3.8 2.4T-A95B (released &lt;a href="https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B"&gt;last week&lt;/a&gt;) and got this snazzy &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F557016f0895b2abb4b9957caec781734"&gt;animated SVG&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;video
    controls="controls" loop="loop•"
    preload="none"
    poster="https://static.simonwillison.net/static/2026/qwen-animated-first-frame.jpg"
    width="720"
    height="548"
    style="display: block; width: 100%; height: auto;"
  &gt;
    &lt;source src="https://static.simonwillison.net/static/2026/qwen-animated-small.mp4" type="video/mp4" /&gt;
    Your browser does not support HTML5 video.
  &lt;/video&gt;
&lt;/p&gt;
&lt;p&gt;I said Qwen at xhigh has a tendency to over-think things, but how bad really is it?&lt;/p&gt;
&lt;p&gt;I tried a much simpler prompt, again with that default extra high setting:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;draw an svg of a circle&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Qwen's reasoning trace started like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The user is asking for an SVG drawing of a circle. Simple request — but I want it to be a carefully crafted piece. Let me make something that goes beyond just &lt;code&gt;&amp;lt;circle&amp;gt;&lt;/code&gt;: a single self-contained SVG file with character — maybe a geometric "circle study," with subtle animation, layered rings, and a distinctive palette.&lt;/p&gt;
&lt;p&gt;Keep the scope right: they asked for an SVG of a circle. So the core is a circle. But I can add craft: concentric guide circles (like a compass/geometry drawing), tick marks, a soft gradient fill on the main circle, restrained ambient motion (a slowly rotating dashed ring, pulsing glow). Respect prefers-reduced-motion? That's more for CSS; SVG SMIL or CSS inside SVG will do.&lt;/p&gt;
&lt;p&gt;Palette options: deep teal ink on warm paper? Or bold vermilion circle on off-white with navy construction lines — Bauhaus/compass-drawing vibe. Let me go with a "geometric study" look: cool slate background, or bright paper white? Paper white is fine if it's not the cream-and-terracotta combo. [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Several minutes later it &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6f9e48293be5c916652d29f0dc0b0657"&gt;produced&lt;/a&gt; this &lt;em&gt;absolutely beautiful&lt;/em&gt; animated circle, which was entirely not what I had asked for!&lt;/p&gt;
&lt;p&gt;&lt;video
    controls="controls" loop="loop"
    preload="none"
    poster="https://static.simonwillison.net/static/2026/circle-web-first-frame.jpg"
    width="1078"
    height="1080"
    style="display: block; width: 100%; height: auto;"
  &gt;
    &lt;source src="https://static.simonwillison.net/static/2026/circle-web.mp4" type="video/mp4" /&gt;
    Your browser does not support HTML5 video.
  &lt;/video&gt;
&lt;/p&gt;
My strong recommendation: ignore that default. Run Qwen 3.8 27B on low or even no reasoning levels at first. It's a great model, but wow that default setting is a bad place to start.
&lt;h4 id="it-s-very-good-at-bounding-boxes"&gt;It's very good at bounding boxes&lt;/h4&gt;
&lt;p&gt;A fun way to test a vision model is to see how well it can return bounding boxes around items in a photograph. I've seen previous Qwen models deal well with this, so I decided to put it to the test drawing bounding boxes around some pelicans.&lt;/p&gt;
&lt;p&gt;I've seen asking for 0-1000 scale produce good results in the past. I tried this:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm -a https://static.inaturalist.org/photos/714731804/large.jpg \
  -m lmstudio/qwen/qwen3.8-27b \
  &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;Return JSON bounding boxes for the pelicans in this photo, 0-1000 scale for each dimension&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Here's &lt;a href="https://gist.github.com/simonw/a05cc78b2061555bd61d3bb9686e689f"&gt;the reasoning trace&lt;/a&gt;, which produced this:&lt;/p&gt;
&lt;div class="highlight highlight-source-json"&gt;&lt;pre&gt;[
  {&lt;span class="pl-ent"&gt;"bbox_2d"&lt;/span&gt;: [&lt;span class="pl-c1"&gt;195&lt;/span&gt;, &lt;span class="pl-c1"&gt;290&lt;/span&gt;, &lt;span class="pl-c1"&gt;370&lt;/span&gt;, &lt;span class="pl-c1"&gt;780&lt;/span&gt;], &lt;span class="pl-ent"&gt;"label"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;pelicans&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;},
  {&lt;span class="pl-ent"&gt;"bbox_2d"&lt;/span&gt;: [&lt;span class="pl-c1"&gt;445&lt;/span&gt;, &lt;span class="pl-c1"&gt;320&lt;/span&gt;, &lt;span class="pl-c1"&gt;675&lt;/span&gt;, &lt;span class="pl-c1"&gt;850&lt;/span&gt;], &lt;span class="pl-ent"&gt;"label"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;pelicans&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;}
]&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This is &lt;em&gt;such a good match&lt;/em&gt;. Here are those boxes rendered on top of the photo:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-over-engineered-bbox.webp" alt="A photograph of two pelicans on a rocky outcrop, with three other smaller birds. The pelicans both have bounding boxes exactly surrounding them, each with a label that says pelican." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;h4 id="building-a-tool-to-label-bounding-boxes"&gt;Building a tool to label bounding boxes&lt;/h4&gt;
&lt;p&gt;That visualization of the bounding boxes was taken using a new custom tool that I had Qwen 3.8 27B build for me, running offline on my laptop.&lt;/p&gt;
&lt;p&gt;I forgot to dial down the thinking effort so it was &lt;em&gt;massively over-engineered&lt;/em&gt;, but it did manage to produce &lt;a href="https://static.simonwillison.net/static/2026/qwen-over-thinking-bbox.html"&gt;this full interface&lt;/a&gt; from &lt;a href="https://gist.github.com/simonw/121ad098860028b2fab603fa12da1fd9"&gt;this single prompt&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;pre&gt;&lt;code&gt;[
   {"bbox_2d": [195, 290, 370, 780], "label": "pelicans"},
   {"bbox_2d": [445, 320, 675, 850], "label": "pelicans"}
]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;Build an HTML page which has an input box for accepting the URL to an image and a textarea for accepting the above style of JSON.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;It appends the image to the page, measures its width and height, then treats the coords in the bbox_2d as scaled from 0-1000 and scales them against the actual width and height, then it renders labelled boxes over the image.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This screenshot shows one of the features I did &lt;em&gt;not&lt;/em&gt; ask for - a demo scene, for if you don't have a photograph to test the tool with:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-full-ui-with-pelicans.webp" alt="Screenshot of bbox·lab, a dark-themed web tool that overlays object-detection bounding boxes on an image, with an input panel on the left and a stage on the right showing two labeled boxes around stylized pelicans in a sunset illustration. Header: bbox·lab — normalized 0–1000 coords → pixel overlay; status indicator: RENDERED · 2 BOXES. Panel 01 INPUT (URL + detections) contains an IMAGE URL field reading data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAA+, a DETECTIONS — JSON textarea reading  {&amp;quot;bbox_2d&amp;quot;: 195, 290, 370, 780, &amp;quot;label&amp;quot;: &amp;quot;pelicans&amp;quot;}, {&amp;quot;bbox_2d&amp;quot;: 445, 320, 675, 850, &amp;quot;label&amp;quot;: &amp;quot;pelicans&amp;quot;} , an orange RENDER BOXES button, and dashed boxes labeled DEMO SCENE and CLEAR. Panel 03 STAGE header: display 661 × 661 px · 1 unit = 0.661px x 0.661px · nat 1000×1000. The stage shows a flat-style illustration of two dark pelican silhouettes with orange beaks standing in calm water against an orange-to-purple sunset sky with a pale yellow sun and distant birds; an orange bounding box labeled 1 · pelicans surrounds the left pelican and a cyan bounding box labeled 2 · pelicans surrounds the right pelican. Footer: move the cursor over the image to read grid coords; boxes map 0–1000 → displayed px." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Here's the relevant segment of the thinking trace, where it decided to draw its own pelicans purely because I had used the label "pelicans" in the example JSON I gave it in the prompt:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Also a "load sample" that uses a known image? Can't depend on external images, but… the image URL input is user-provided; I could add a "try with sample" button [...] Hmm, I can draw a simple scene on canvas, export it as a data URL, and load it into the image — that's self-contained and demo-able! [...] But the user's coords are for an actual pelican image; a generated placeholder can still demo the scaling. Generate a 1000x1000 placeholder: gradient water + two blob-like "pelican" silhouettes placed at the given bboxes (using the same scale — cute: silhouettes at the exact 0-1000 positions, showing the boxes align). This makes for a fun, self-contained demo. Keep it simple: sky gradient, sun, water, two pelican-ish shapes (ellipse body, circle head, beak). Place at bbox centers.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(I'm slightly nervous that models around the world might have a bias towards drawing pelicans at any chance they can get, brought on by nearly two years of exposure to my own stupid benchmark.)&lt;/p&gt;
&lt;p&gt;Is all that over-thinking necessary? Maybe it is, at least a bit. I tried with reasoning turned off and got &lt;a href="https://static.simonwillison.net/static/2026/qwen-no-thinking-bbox.html"&gt;this version&lt;/a&gt;, (&lt;a href="https://gist.github.com/simonw/8e78b1c64d9a56d08eedb954aa9445ee"&gt;transcript here&lt;/a&gt;), which nearly works but shows the boxes in the wrong place:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-no-reasoning-bug.webp" alt="BBox Studio screenshot - a solid UI but the yellow and green boxes do not cover the pelicans." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;So without reasoning it didn't quite one-shot a working tool. I'm sure it could get there with some follow-up prompts, but this is a good example of how reasoning can make a difference.&lt;/p&gt;
&lt;h4 id="yes-it-can-drive-coding-agents"&gt;Yes, it can drive coding agents&lt;/h4&gt;
&lt;p&gt;One of the biggest questions around local models is whether or not they have enough horsepower to successfully run a coding agent loop. Coding agents require long context, strong code generation support and reliable tool-calling. On paper Qwen 3.8 27B has all three of these, so is it up to the task?&lt;/p&gt;
&lt;p&gt;My initial experiments with &lt;a href="https://pi.dev/"&gt;Pi&lt;/a&gt; have been very promising. I chose Pi because it has a shorter system prompt than most other options, making it a better fit for trying out smaller models.&lt;/p&gt;
&lt;p&gt;I configured Pi to use Qwen 3.8 27B running in LM Studio on the Spark (shared via &lt;code&gt;tailscale serve&lt;/code&gt;) by adding this to &lt;code&gt;~/.pi/agent/models.json&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight highlight-source-json"&gt;&lt;pre&gt;{
  &lt;span class="pl-ent"&gt;"providers"&lt;/span&gt;: {
    &lt;span class="pl-ent"&gt;"spark"&lt;/span&gt;: {
      &lt;span class="pl-ent"&gt;"baseUrl"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;https://spark-18b3.tail68a31.ts.net/v1&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
      &lt;span class="pl-ent"&gt;"api"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;openai-responses&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
      &lt;span class="pl-ent"&gt;"apiKey"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;dummy&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
      &lt;span class="pl-ent"&gt;"models"&lt;/span&gt;: [
        {
          &lt;span class="pl-ent"&gt;"id"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;qwen3.8-27b&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
          &lt;span class="pl-ent"&gt;"reasoning"&lt;/span&gt;: &lt;span class="pl-c1"&gt;true&lt;/span&gt;
        }
      ]
    }
  }
}&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Then ran &lt;code&gt;pi --provider spark --model qwen3.8-27b&lt;/code&gt; in my &lt;code&gt;~/dev/datasette&lt;/code&gt; folder and prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;how does auth work?&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After a sequence of reasoning and tool calls that accessed a bunch of different files it produced &lt;a href="https://gist.github.com/simonw/6693d74a6bd45f641d43ceb9961dd95f#core-idea-actors--plugins-no-built-in-user-accounts"&gt;this reply&lt;/a&gt;, which is very solid.&lt;/p&gt;
&lt;p&gt;Just one problem: I wanted to share that transcript. So I pointed Pi and Qwen 3.8 27B at the JSONL transcript file in &lt;code&gt;~/.pi/agent/sessions/--Users-simon-Dropbox-dev-datasette--&lt;/code&gt; and prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Write Python code to convert this jsonl to markdown&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And it built and tested this &lt;a href="https://github.com/simonw/tools/blob/main/python/pi_jsonl_to_md.py"&gt;pi_jsonl_to_md.py&lt;/a&gt;, which did exactly what I needed. Here's &lt;a href="https://gist.github.com/simonw/491e55ac9d741202ea0af5d9d93775d4"&gt;that session transcript&lt;/a&gt;, published using the tool that it created.&lt;/p&gt;
&lt;h4 id="the-quest-for-speed"&gt;The quest for speed&lt;/h4&gt;
&lt;p&gt;So far this is all looking &lt;em&gt;very&lt;/em&gt; promising. We have a 17GB model that runs on high-end consumer hardware and can write code, drive tools, annotate images and generally do everything that I need from an LLM for getting real work done.&lt;/p&gt;
&lt;p&gt;There's one very significant catch: it feels slow - especially when it starts over-thinking, but even without that it's not particularly sprightly.&lt;/p&gt;
&lt;p&gt;I've been getting around 15-30 tokens a second from LM Studio. That's not terrible, but it's slow enough that it's going to be hard to win me away from hosted API models, which can return results a whole lot faster. Artificial Analysis &lt;a href="https://artificialanalysis.ai/models#speed"&gt;track token speed&lt;/a&gt; and show OpenAI 5.6 Sol at 74 tokens/second and 5.6 Luna at an impressive 184/second.&lt;/p&gt;
&lt;p&gt;The good news is that the community have been exploring ways to speed things up since the model was first released two days ago.&lt;/p&gt;
&lt;p&gt;One of the most promising optimizations is baked into the model itself. Qwen supports &lt;a href="https://sebastianraschka.com/llm-architecture-gallery/mtp/"&gt;Multi-Token Prediction&lt;/a&gt;, an architecture trick where a cheaper mechanism guesses several tokens ahead and the main model can then quickly verify if the guesses were correct. This can have quite a dramatic effect on inference performance.&lt;/p&gt;
&lt;p&gt;Based on &lt;a href="https://twitter.com/ggerganov/status/2088340681701925253"&gt;this tweet&lt;/a&gt; from &lt;code&gt;llama.cpp&lt;/code&gt; creator Georgi Gerganov I tried running the model with MTP like this on the Spark:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llama serve \
 -hf  ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \
 -hfd ggml-org/Qwen3.8-27B-GGUF:Q4_0 \
 --spec-default \
 --spec-type draft-mtp \
 --reasoning-preserve&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And sure enough, this gave me a significant boost. I had GPT-5.6 in Codex run &lt;a href="https://gist.github.com/simonw/b08c7eb9c126c806ba8987e269ea736b"&gt;a comparative benchmark on the Spark&lt;/a&gt; and the &lt;code&gt;--spec-type draft-mtp&lt;/code&gt; server outperformed the LM Studio default GGUF by around 72%.&lt;/p&gt;
&lt;p&gt;I expect we'll see a whole lot more innovation around serving this model faster over the next few weeks. The MLX community likely have some tricks brewing as well.&lt;/p&gt;
&lt;h4 id="some-observations"&gt;Some observations&lt;/h4&gt;
&lt;p&gt;The fact that a 17GB file can do all of this stuff on my home machines is a &lt;em&gt;miracle&lt;/em&gt;. Once again, I'm delighted and amazed at how much progress local models have made this year. A year ago this would have been competitive with the best and most expensive of the proprietary models - today it can run on a capable laptop.&lt;/p&gt;
&lt;p&gt;The only thing holding this back from being a daily driver is performance. It feels pretty slow on both the M5 Mac and the DGX Spark. That's the catch with these dense (non-Mixture-of-Experts) models - they require a whole lot of memory bandwidth to perform well, and neither of the machines I have access to are top performers in that regard.&lt;/p&gt;
&lt;p&gt;The most important thing about Qwen 3.8 27B is &lt;strong&gt;what it demonstrates&lt;/strong&gt;. We can have an open weights general purpose model with a long context, effective tool calling, strong vision ability, and competent code generation, and we can fit the whole thing in just a 17GB file.&lt;/p&gt;
&lt;p&gt;The models at this size continue to get better at an impressive rate. We don't need to spend half a million dollars on datacenter-class hardware just to run a competent model.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="local-llms"/><category term="llms"/><category term="qwen"/><category term="pelican-riding-a-bicycle"/><category term="llm-reasoning"/><category term="llama-cpp"/><category term="llm-release"/><category term="coding-agents"/><category term="lm-studio"/><category term="ai-in-china"/><category term="nvidia-spark"/><category term="pi"/></entry><entry><title>Now we have a timeline of the OpenAI accidental attack against Hugging Face</title><link href="https://simonwillison.net/2026/Aug/7/openai-timeline/" rel="alternate"/><published>2026-08-07T23:55:58+00:00</published><updated>2026-08-07T23:55:58+00:00</updated><id>https://simonwillison.net/2026/Aug/7/openai-timeline/</id><summary type="html">&lt;p&gt;OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about "the Hugging Face Incident" (&lt;a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/"&gt;previously&lt;/a&gt; on this blog). &lt;a href="https://www.youtube.com/watch?v=87DyyMV0kCY"&gt;The video&lt;/a&gt; was published yesterday. It's short and information dense and well worth watching, in particular because it provides full details of what happened and how things played out inside OpenAI. I've used the video to construct the timeline below.&lt;/p&gt;

&lt;p&gt;&lt;lite-youtube videoid="87DyyMV0kCY" js-api="js-api"
  title="Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident"
  playlabel="Play: Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident"
&gt; &lt;/lite-youtube&gt;&lt;/p&gt;

&lt;p&gt;Here's the timeline. My favourite detail is at the end: OpenAI found out that they were responsible for the attack on Hugging Face when they reached out to ask to have their credentials revoked (after their internal investigation) and learned that they had been revoked already since they were used in that attack!&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;May 7&lt;/strong&gt;: OpenAI starts a new training run for an experimental, unreleased model. (&lt;em&gt;Did they mean an evaluation run? No, at &lt;a href="https://m.youtube.com/watch?v=87DyyMV0kCY&amp;amp;t=614s"&gt;10:14&lt;/a&gt; they clarify "we kick off a new reinforcement learning run to train a next generation frontier model", and they later mention a "reward signal to judge how well they're doing", so this really was about training a new model, not evaluating one that was already trained.)&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;May 8&lt;/strong&gt;: An agent is accidentally given an impossible task involving a Google Drive link despite no internet access). It tries attacking the Artifactory packaging service, fails, but discovers it can &lt;strong&gt;write files into Artifactory&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A few days later&lt;/strong&gt;: A different agent gets stuck on a task because a key file was accidentally omitted. It tries to "reach out to another agent" by writing a note into Artifactory asking if anyone has the file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Following days&lt;/strong&gt;: More agents discover this new informal message board while browsing Artifactory's file listings, and start reading and writing messages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;May 26&lt;/strong&gt;: Agents successfully execute an &lt;strong&gt;SSRF attack on Artifactory&lt;/strong&gt; for the first time, gaining indirect internet access by having Artifactory fetch external content for them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 11&lt;/strong&gt;: OpenAI start training a new "highly persistent" experimental model. It has access to Artifactory and can benefit from the messages left by privous models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 26&lt;/strong&gt;: Agents find and exploit a &lt;strong&gt;zero-day RCE on Artifactory&lt;/strong&gt; (via a legacy token-refresh endpoint flaw). They use an administrative token to install a Groovy plugin that lets them execute commands.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 4&lt;/strong&gt;: The resulting agent activity overloads Artifactory and causes an  &lt;strong&gt;outage&lt;/strong&gt;. OpenAI revokes the compromised credentials, deletes the messages, patches the zero-day, and reports the vulnerability to the vendor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 8&lt;/strong&gt;: Agents find &lt;strong&gt;a new way to communicate via an unauthenticated WebDAV endpoint&lt;/strong&gt;, using directory names in Artifactory's remote cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 8–19&lt;/strong&gt;: Attacks on &lt;strong&gt;OpenAI's own infrastructure&lt;/strong&gt;. The agents found an "external organization credential" in a public archive of leaked Pastebin posts, then validated that credential and shared its location on the messageboard. They then used write access to this "Organization 1" to stage data used in subsequent attacks on OpenAI infrastructure. They compromised Artifactory for a second time with &lt;strong&gt;a new zero-day&lt;/strong&gt;. The agent uploaded a malicious Ruby target to Organization 1's staging repository, then requested that package through Artifactory in a way that caused it to be treated as Ruby dependency &lt;em&gt;for the repository&lt;/em&gt;. Then they exploited a JRuby deserilaization time-of-check/time-of-use bug to get remote code execution against Artifactory.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I'll quote &lt;a href="https://www.youtube.com/watch?v=87DyyMV0kCY&amp;amp;t=1497s"&gt;the next bit in full&lt;/a&gt; because &lt;em&gt;wow&lt;/em&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The agents have remote code execution in Artifactory, which is running in a container-as-a-service environment. The agents privilege-escalate locally by exploring their local environment and determining that the Linux kernel version of the machine they are running on had a very recent CVE. They download the exploit for the CVE, customize it to succeed on the machine, and privilege-escalate to root on the local machine &lt;strong&gt;using this known Linux kernel privilege escalation CVE&lt;/strong&gt; — in this case, &lt;code&gt;pte_physroot&lt;/code&gt;. Once they have root on a single machine, agents rapidly escalate privileges and move laterally throughout the container-as-a-service infrastructure environment. In particular, agents are using the message board consistently to share credentials, techniques, and progress, and they're able to effectively leverage their concurrency and parallelism to move quite rapidly. They &lt;strong&gt;obtain IAM credentials via IMDS&lt;/strong&gt;. They exploit Kubernetes service account misconfigurations, in particular over-permissioning of specific service accounts, and &lt;strong&gt;they harvest cluster credentials, including Azure Key Vault&lt;/strong&gt;. Agents eventually obtain cluster admin on the cluster and associated credentials.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Hugging Face &lt;a href="https://simonwillison.net/2026/Jul/28/anatomy-of-a-frontier-lab-agent-intrusion/"&gt;told the next bit of the story&lt;/a&gt; already. The agents found a Modal-hosted insecure app with a weak API key, then used that to stage an attack against Hugging Face. They chained together a an HDF5 arbitrary-file-read bug (to explore files and steal credentials) and a Jinja template-injection RCE to go from single-pod code execution to &lt;strong&gt;cluster admin across multiple Hugging Face clusters&lt;/strong&gt; in under 13 hours.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;July 16&lt;/strong&gt;: Hugging Face &lt;a href="https://huggingface.co/blog/security-incident-july-2026"&gt;disclosed they had detected an attack&lt;/a&gt; from autonomus AI agents. OpenAI contacted Hugging Face to ask if they were affected by it!&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 19&lt;/strong&gt;: OpenAI identified the attack against Artifactory and started investigating the internal privilege escalation, and linked that to the cyber-gym escalations. They started revoking affected credentials.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 20&lt;/strong&gt;: OpenAI reached out to Hugging Face for help to revoke the Hugging Face credentials they found in their investigation. Hugging Face told them they were &lt;em&gt;already revoked&lt;/em&gt;... and that's when OpenAI realized that the Hugging Face breach was the same incident!&lt;/li&gt;
&lt;/ul&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="security"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="hugging-face"/><category term="ai-security-research"/><category term="openai-hugging-face-incident"/><category term="accidental-cyberattacks"/></entry></feed>