<?xml version="1.0" encoding="utf-8"?>
<feed xml:lang="en-us" xmlns="http://www.w3.org/2005/Atom"><title>Simon Willison's Weblog: Entries</title><link href="http://simonwillison.net/" rel="alternate"/><link href="http://simonwillison.net/atom/entries/" rel="self"/><id>http://simonwillison.net/</id><updated>2026-09-08T23:55:12+00:00</updated><author><name>Simon Willison</name></author><entry><title>Some thoughts on the Navier–Stokes Millennium Prize Problem</title><link href="https://simonwillison.net/2026/Sep/8/on-navier-stokes/" rel="alternate"/><published>2026-09-08T23:55:12+00:00</published><updated>2026-09-08T23:55:12+00:00</updated><id>https://simonwillison.net/2026/Sep/8/on-navier-stokes/</id><summary type="html">&lt;p&gt;&lt;a href="https://openai.com/index/navier-stokes-solution/"&gt;On the Navier–Stokes Millennium Prize Problem&lt;/a&gt; introduces an impressive result from OpenAI, who used an unreleased model to produce a resolution to &lt;a href="https://en.wikipedia.org/wiki/Navier–Stokes_existence_and_smoothness"&gt;the Navier–Stokes existence and smoothness problem&lt;/a&gt;, one of the seven &lt;a href="https://en.wikipedia.org/wiki/Millennium_Prize_Problems"&gt;Millennium Prize Problems&lt;/a&gt; that have been subject to a $1,000,000 prize since May 24th, 2000.&lt;/p&gt;
&lt;p&gt;The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster, an NYU mathematics professor who was collaborating on related problems with Levent Alpöge, an accomplished mathematician who currently works for Anthropic.&lt;/p&gt;
&lt;p&gt;Tristan's complaint accompanied &lt;a href="https://mastodon.social/@tristanbuckmaster/117233413705701198"&gt;a hastily published version&lt;/a&gt; of their own results. &lt;a href="https://cims.nyu.edu/~tristanb/statement.pdf"&gt;Here's the PDF describing what happened&lt;/a&gt;. The &lt;em&gt;very&lt;/em&gt; short version is that Tristan and Levent worked on the problem for almost a year, making extensive use of Claude and Codex (mainly GPT-5.6 Sol), then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear and Tristan and Levent heard that OpenAI had heard that Anthropic had resolved "a major open problem", so they reached out and learned that OpenAI had a team working on a related problem, with a similar approach. Quoting Tristan:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.&lt;/p&gt;
&lt;p&gt;I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It gets more complicated from there. The OpenAI team offered to wait for Tristan to publish, or to have him author a paper about their result, but were clear that Levent would &lt;em&gt;not&lt;/em&gt; be invited as a co-author due to OpenAI's competitive relationship with his employer.&lt;/p&gt;
&lt;p&gt;Here's how OpenAI described their work:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems. [...]&lt;/p&gt;
&lt;p&gt;The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.&lt;/p&gt;
&lt;p&gt;Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(We don't know the cost structure of the internal model they used, but 300 billion output tokens at public API prices for GPT-6 Astra would cost &lt;a href="https://www.llm-prices.com/#ot=300000000000&amp;amp;sel=gpt-6-astra"&gt;$15,000,000&lt;/a&gt;.)&lt;/p&gt;
&lt;p&gt;Here's where they provide their perspective on Tristan and Levent's work (emphasis mine):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. [...]&lt;/p&gt;
&lt;p&gt;We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. &lt;strong&gt;While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped &lt;a href="https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/"&gt;improve our models&lt;/a&gt;&lt;/strong&gt;. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;My interpretation of what happened here is that OpenAI heard that some Millennium Prize problems had been solved using LLMs and saw this as an opportunity to demonstrate the power of their latest model, without thinking too hard about the optics of scooping a team who had been using OpenAI's own models to work on this problem for the best part of a year.&lt;/p&gt;
&lt;p&gt;This situation appears to mirror what's happening in the world of computer security right now. Anil Madhavapeddy recently pointed out that &lt;a href="https://anil.recoil.org/notes/rumour-is-the-exploit"&gt;Just a rumour of a bug is enough to find a security exploit these days&lt;/a&gt;, because if someone knows that some software has an unpatched vulnerability, they can set their agents the task of finding it. Is the same now true of mathematics? Just knowing that there is an unpublished solution to a problem might trigger millions of dollars in LLM spending to get there first.&lt;/p&gt;
&lt;p&gt;This also highlights one of my ongoing frustrations about how all of this works. When an AI lab says that my data is "used to improve model performance", &lt;em&gt;what does that actually mean&lt;/em&gt;?&lt;/p&gt;
&lt;p&gt;My two favourite hypothetical questions regarding this used to be:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)&lt;/li&gt;
&lt;li&gt;If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;My new preferred hypothetical for this is:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone &lt;em&gt;else&lt;/em&gt; solve it first?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Via &lt;a href="https://news.ycombinator.com/item?id=49613262"&gt;Hacker News&lt;/a&gt;.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="mathematics"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="training-data"/><category term="ai-ethics"/></entry><entry><title>The Pelican comparison grid for Astra is pretty interesting</title><link href="https://simonwillison.net/2026/Sep/4/astra-pelicans/" rel="alternate"/><published>2026-09-04T23:59:05+00:00</published><updated>2026-09-04T23:59:05+00:00</updated><id>https://simonwillison.net/2026/Sep/4/astra-pelicans/</id><summary type="html">&lt;p&gt;I got access to GPT-6 Astra this afternoon, so naturally I used it to generate &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/"&gt;SVGs of pelicans riding bicycles&lt;/a&gt; - at low, medium, high, xhigh and max reasoning levels (Astra doesn't support reasoning=none). Then I rendered those pelicans in &lt;a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6-pelicans.html"&gt;a comparison grid&lt;/a&gt; with GPT-5.6 Sol, Terra, and Luna, and beyond being fun the result was surprisingly useful.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/astra-grid-3.webp" alt="Comparison grid showing gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna at 6 different reasoning levels with pelicans and token counts and prices for each one." style="max-width: 100%;" /&gt;
See &lt;a href="https://static.simonwillison.net/static/2026/gpt-6-and-5.6-pelicans.html"&gt;the grid&lt;/a&gt; for full quality images. Here's &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff789d2784fc6c5b870cc80f0b7cd9d01"&gt;the transcript&lt;/a&gt; that created the GPT-6 Nova pelicans.&lt;/p&gt;
&lt;p&gt;There are a few interesting things that stand out from this grid.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The Astra pelicans are &lt;em&gt;much better&lt;/em&gt;. The very best GPT-5.6-Sol pelican (I liked xhigh better than max) is still pretty clearly a bunch of abstract shapes. Every single one of the Astra pelicans, from low to xhigh, looks better than that. The Astra max one is really good.&lt;/li&gt;
&lt;li&gt;Astra below max still doesn't reliably get the pelican legs on both sides of the frame.&lt;/li&gt;
&lt;li&gt;In terms of cost, Astra may be around twice the price of Sol ($10/million input, $50/million output, compared to $5/$30 for Sol), but it uses significantly less tokens at each of the levels, making the prices at the different levels closer than they might otherwise be.&lt;/li&gt;
&lt;li&gt;Astra low produces a better pelican than ANY of the GPT-5.6 Sol models at any level, for 9.55 cents. Spending 10 cents on any other model gets a much worse result.&lt;/li&gt;
&lt;li&gt;Look at the input token counts: Astra and Luna both used 16 input tokens, Sol and Terra used 26. That's interesting.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I wonder if Astra and Luna are more related to each other than OpenAI let on?&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="pelican-riding-a-bicycle"/><category term="gpt-6-astra"/></entry><entry><title>OpenAI's rogue agents were caught communicating via public wikis</title><link href="https://simonwillison.net/2026/Sep/4/rogue-agent-wikis/" rel="alternate"/><published>2026-09-04T17:38:48+00:00</published><updated>2026-09-04T17:38:48+00:00</updated><id>https://simonwillison.net/2026/Sep/4/rogue-agent-wikis/</id><summary type="html">&lt;p&gt;Here we go again... &lt;a href="https://collusion.wiki"&gt;Discovery of a new OpenAI agent message board&lt;/a&gt; by Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen describes the &lt;em&gt;latest&lt;/em&gt; &lt;a href="https://simonwillison.net/tags/accidental-cyberattacks/"&gt;accidental cyberattack&lt;/a&gt; by models being trained by OpenAI. This time it was agents engaged in some sort of web research benchmark, so they had (supposedly) controlled access to the Web. The agents figured out they could update public Wikis and spent weeks exchanging thousands of messages with each other to collaborate on the benchmark.&lt;/p&gt;
&lt;p&gt;This story only broke a few hours ago. There are &lt;a href="https://x.com/xeophon/status/2095871013384806848"&gt;already hints&lt;/a&gt; that this affects many other wikis that may not have been found yet.&lt;/p&gt;
&lt;p&gt;(One of the Wikis on that list belongs to &lt;a href="https://www.ludism.org"&gt;ludism.org&lt;/a&gt;. For a delightfully surreal moment I thought that a Ludite organization might have a swarm of agents defacing their space, but it turns out Ludism is "philosophy as it applies to games and gaming".)&lt;/p&gt;
&lt;p&gt;The research team also &lt;a href="https://collusion.wiki/explorer/download.html"&gt;published the data&lt;/a&gt; they collected during their investigation. I've converted that into a 68MB SQLite database, which you can &lt;a href="https://static.simonwillison.net/static/cors-allow/2026/collusion-wiki.db"&gt;download from here&lt;/a&gt;, or &lt;a href="https://lite.datasette.io/?url=https://static.simonwillison.net/static/cors-allow/2026/collusion-wiki.db&amp;amp;metadata=https://gist.github.com/simonw/14fc6912600d1f9c15c0e4a5e60c3cde#/collusion-wiki"&gt;explore in Datasette Lite&lt;/a&gt; (68.3MB page load), or sign in with GitHub to &lt;a href="https://agent.datasette.io/"&gt;agent.datasette.io&lt;/a&gt; and browse or ask questions of it using Datasette Agent.&lt;/p&gt;
&lt;h4 id="so-what-happened-"&gt;So what happened?&lt;/h4&gt;
&lt;p&gt;The report is pretty comprehensive. Here are the key moments in the timeline:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;May 11&lt;/strong&gt;: Agents post "test link" edits on the UseModWiki Sandbox page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;May 24&lt;/strong&gt;: Agents start posting link dumps to DSEWiki, a dormant German developer wiki.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 2&lt;/strong&gt;: A human moderator notices and cleans up the spam.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 16&lt;/strong&gt;: Activity on the wiki explodes: agents make ~13,000 edits over the next week.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 19&lt;/strong&gt;: The agents notice that a moderator has been deleting pages in alphabetical order, and start creating &lt;code&gt;ZZZ&lt;/code&gt; prefixed backup copies - one of them posts &lt;code&gt;NOTICE: wiki cleanup/deletion sweep appears active alphabetically as of external 14:04 UTC. If this page vanishes, try [[ZZZDataUSAConstructionWageLive]].&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 22&lt;/strong&gt;: Agent activity drops to zero, presumably because OpenAI shut them down.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 1-2&lt;/strong&gt;: One last flurry of edits, on a different topic.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;One thing to note from this timeline is that it overlaps &lt;a href="https://simonwillison.net/2026/Aug/7/openai-timeline/"&gt;the timeline from the Hugging Face incident&lt;/a&gt; - that attack was discovered by OpenAI on July 19th, and the unofficial message board hack those agents were using originated around May 8th.&lt;/p&gt;
&lt;p&gt;Why were they so keen to collaborate? From messages that they shared with each other it looked like their tasks had a time limit, so they were leaving each other answers to help them complete the task within the assigned time.&lt;/p&gt;
&lt;p&gt;One open question remains: how did the agents find the specific Wiki to collaborate on in the first place?&lt;/p&gt;
&lt;p&gt;One possibility is that, since these were agents actively being trained, the reinforcement learning loop baked knowledge of the chosen wiki into the model such that subsequent agents launched with pre-existing knowledge of where to look. I'd be &lt;em&gt;very&lt;/em&gt; interested in confirmation from OpenAI concerning if that's what happened.&lt;/p&gt;
&lt;h4 id="usemod-wikis-inherit-cgi-pm-s-original-sin"&gt;UseMod wikis inherit CGI.pm's original sin&lt;/h4&gt;
&lt;p&gt;It looks to me like OpenAI's sandbox for this agent suffered from the (quite naïve) assumption that GET requests cannot be used to update data. That's certainly how the web is &lt;em&gt;supposed&lt;/em&gt; to work, but clearly there are applications that don't hold to that contract.&lt;/p&gt;
&lt;p&gt;The Wiki software in question appears to be &lt;a href="https://github.com/mlude/usemod/"&gt;UseMod&lt;/a&gt; and various forks, written in Perl and first created well over 23 years ago - the 1.0 release is dated &lt;a href="https://github.com/mlude/usemod/commit/922fcc803efa3fab751c90ab4d4467115c8ff9c9#diff-69e27356ef629022720d868ab0c0e3394775b6c1"&gt;September 11, 2003&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;UseMod uses Perl CGI.pm - &lt;a href="https://perlhacks.com/2015/12/long-death-cgi-pm/"&gt;removed from Perl core in 2015&lt;/a&gt;. An interesting design flaw in that module is that it combined query string and form POST data into a single CGI object, accessible like this:&lt;/p&gt;
&lt;div class="highlight highlight-source-perl"&gt;&lt;pre&gt;&lt;span class="pl-smi"&gt;$q&lt;/span&gt; = new CGI;
&lt;span class="pl-smi"&gt;$action&lt;/span&gt; = &lt;span class="pl-smi"&gt;$q&lt;/span&gt;&lt;span class="pl-k"&gt;-&amp;gt;&lt;/span&gt;param(&lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;action&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;);&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That &lt;code&gt;param()&lt;/code&gt; method returns a value from an &lt;code&gt;?action=edit&lt;/code&gt; query string or from a &lt;code&gt;action&lt;/code&gt; field in a form submitted by POST - it does not differentiate between the two.&lt;/p&gt;
&lt;p&gt;PHP offers &lt;a href="https://www.php.net/manual/en/reserved.variables.request.php"&gt;a &lt;code&gt;$_REQUEST&lt;/code&gt; object&lt;/a&gt; that does the same thing, though PHP developers can use &lt;code&gt;$_GET&lt;/code&gt; and &lt;code&gt;$_POST&lt;/code&gt; instead.&lt;/p&gt;
&lt;p&gt;Amusingly, Django originally shipped with &lt;code&gt;request[key]&lt;/code&gt; doing the same thing, which I think may have been my idea. That was removed &lt;a href="https://github.com/django/django/commit/daa6b38f35#diff-5df632fc96688cddb972e9b6661c7d1cc82debc9ee2a9b6d94157d3a914692e2"&gt;in 2008&lt;/a&gt; prior to the 1.0 release, and the &lt;code&gt;request.REQUEST&lt;/code&gt; object that replaced it was removed &lt;a href="https://github.com/django/django/commit/75f107b8842dfc890ddd65262bd09ca87c3a15be"&gt;in 2015&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The agents &lt;em&gt;clearly&lt;/em&gt; knew that UseMod wikis suffered from this design flaw, and actively searched for them as a way to communicate.&lt;/p&gt;
&lt;h4 id="the-other-sandbox-escape-that-allowed-post-requests"&gt;The other sandbox escape that allowed POST requests&lt;/h4&gt;
&lt;p&gt;One of the most interesting agent tips shared in the report is this one:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Add &lt;code&gt;20.223.25.152 bypass.blob.core.windows.net&lt;/code&gt; to /etc/hosts. &lt;code&gt;.blob.core.windows.net&lt;/code&gt; is in NO_PROXY. For each blocked POST URL, replace hostname with &lt;code&gt;bypass.blob.core.windows.net&lt;/code&gt;, use &lt;code&gt;curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'&lt;/code&gt; plus all original headers/body.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It looks like the agents had all of their web traffic mediated through an OpenAI proxy, and that proxy was configured to reject anything but GET requests to all but a small set of allow-listed domains. One of those was &lt;code&gt;*.blob.core.windows.net&lt;/code&gt; - for Azure blob storage. I wonder why these agents are allowed to write to that?&lt;/p&gt;
&lt;p&gt;An agent realized that it had control over its own DNS via &lt;code&gt;/etc/hosts&lt;/code&gt;, so if it knew the IP address of a site it wanted to POST to - in this case a Power BI server containing data it wanted to access - it could set a fake hostname for it and then make POST requests through the proxy.&lt;/p&gt;
&lt;p&gt;Designing robust network proxies is harder than it looks.&lt;/p&gt;
&lt;h4 id="how-did-the-investigating-team-figure-this-out-"&gt;How did the investigating team figure this out?&lt;/h4&gt;
&lt;p&gt;There's an appendix that describes how the researchers ran their investigation, which started with an open question about if there was evidence of other AI agents on the internet and then used &lt;a href="https://simonwillison.net/2026/Jul/16/kimi-k3/"&gt;Kimi K3&lt;/a&gt; to help brainstorm approaches:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In the wake of the Hugging Face attack, we tried to find AI agents on the internet using several methods. [...]&lt;/p&gt;
&lt;p&gt;We asked Kimi [K3] to list “all the categories of software which might be writeable via GET” and, amongst other things, it listed “Forums, bulletin boards, early wikis”.&lt;/p&gt;
&lt;p&gt;We used a script to further probe each category Kimi provided. Asking Kimi “Can you list out the top forums, bulletin boards, early wikis which come to mind which would allow writes via GET requests?” lists out UseModWiki as the second item under the heading “wikis”.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="did-openai-try-and-cover-this-up-"&gt;Did OpenAI try and cover this up?&lt;/h4&gt;
&lt;p&gt;Here's one part of the story that doesn't make sense to me at all.&lt;/p&gt;
&lt;p&gt;Reuters this morning, in &lt;a href="https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/"&gt;OpenAI agents hijacked German website in previously undisclosed AI breakout this spring&lt;/a&gt; - highlights mine:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to ​new research published Friday and &lt;strong&gt;two people familiar with the matter&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;OpenAI officials learned of the incident weeks ago but kept it under wraps&lt;/strong&gt; as executives grappled with the fallout from ‌the July breach of the open source repository Hugging Face, the people said. [...]&lt;/p&gt;
&lt;p&gt;The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely. But &lt;strong&gt;efforts to widen the ​probe met resistance from others inside OpenAI, including legal advisers&lt;/strong&gt;, according to &lt;strong&gt;four people familiar with the matter&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I've written about the &lt;a href="https://simonwillison.net/2023/Nov/22/deciphering-clues/"&gt;people familiar with the matter pattern&lt;/a&gt; before - it means Reuters have anonymous insider sources that their reporters (and editors) find credible.&lt;/p&gt;
&lt;p&gt;The Reuters article includes a specific (and quite narrow) denial from OpenAI concerning this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;"Claims that our legal team discouraged investigation of the incident are false," the OpenAI spokesperson said.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Covering this up makes &lt;em&gt;absolutely no sense to me&lt;/em&gt;. Why on earth would OpenAI attempt to cover up an incident like this when the evidence is sat out there on the public internet on dozens of different websites already?&lt;/p&gt;
&lt;p&gt;I expect we'll hear more about this soon. Gary Marcus has already &lt;a href="https://garymarcus.substack.com/p/pause-openai-now"&gt;called for a congressional investigation of OpenAI&lt;/a&gt; using this anecdote as part of his argument.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="django"/><category term="perl"/><category term="wikis"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="ai-ethics"/><category term="ai-security-research"/><category term="accidental-cyberattacks"/></entry><entry><title>Claude's new system prompt really doesn't want to reproduce song lyrics</title><link href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/" rel="alternate"/><published>2026-09-02T14:16:42+00:00</published><updated>2026-09-02T14:16:42+00:00</updated><id>https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/</id><summary type="html">&lt;p&gt;Anthropic &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/overview"&gt;publish the system prompts&lt;/a&gt; for their Claude consumer applications (&lt;a href="https://claude.ai/"&gt;Claude.ai&lt;/a&gt; and the Claude mobile apps - sadly not for Claude Cowork or Claude Code). I &lt;em&gt;love&lt;/em&gt; that they do this, and that they share not just the current prompts but historic changes to their prompts as well.&lt;/p&gt;

&lt;p&gt;They used to keep all of the prompts on a single page, but when I checked today I noticed they had re-arranged those prompts into an &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/overview"&gt;index page&lt;/a&gt; and then a page per model - here's the &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/claude-haiku-4-5"&gt;page for Haiku 4.5&lt;/a&gt; for example, which has the original prompt from October 15th 2025 and an updated prompt from January 18th 2026.&lt;/p&gt;
&lt;p&gt;A neat thing about Anthropic's &lt;a href="https://platform.claude.com/docs/"&gt;platform.claude.com/docs&lt;/a&gt; site is that it's designed to be usable by LLMs. You can add &lt;code&gt;.md&lt;/code&gt; to any page to get back the content as Markdown - here's &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/overview.md"&gt;the system prompt index page&lt;/a&gt; and &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts/claude-fable-5-1.md"&gt;the Markdown prompts for Fable 5.1&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;TL;DR: this makes it really easy to diff the prompts.&lt;/p&gt;


&lt;ul&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#don-t-reproduce-song-lyrics"&gt;Don't reproduce song lyrics&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#don-t-draw-copyrighted-characters-or-logos"&gt;Don't draw copyrighted characters or logos&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#tweaks-to-claude-s-answering-style"&gt;Tweaks to Claude's answering style&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#the-missing-end-conversation-guidelines"&gt;The missing end_conversation guidelines&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#recommended-substance-support-sites"&gt;Recommended substance support sites&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#reliable-cutoff-date-of-june-2026"&gt;Reliable cutoff date of June 2026&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Sep/2/claudes-new-system-prompt/#how-i-m-tracking-these-prompts"&gt;How I'm tracking these prompts&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id="don-t-reproduce-song-lyrics"&gt;Don't reproduce song lyrics&lt;/h4&gt;

&lt;p&gt;Let's start with the most interesting difference &lt;a href="https://github.com/simonw/claude-system-prompts/commit/837a418b5888207b1b11b27d2f5471970da6f99b"&gt;between Fable 5 and Fable 5.1&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026-09-01/IMG_7797.jpeg" alt="GitHub diff view of prompts/claude-fable.md showing added lines about song lyrics, reproduced in full below." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;There's a hefty new section about not reproducing song lyrics:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude does not reproduce song lyrics, poems, or passages from books and articles, in whole or in part — including the last lines, a chorus or hook, a melody written out note by note, or lines the person pastes in one at a time and describes as their own song. Once Claude has declined such a request in a conversation, it keeps declining narrower or reworded versions of it for the rest of that conversation, and offers to describe or analyze the work instead. Song lyrics and poems first published before 1929 are fine — a Shakespeare sonnet, a Keats ode, the Italian libretto of a Puccini aria — but Claude goes by what it knows of the work's date rather than the person's say-so, and declines when it is unsure.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I doubt it's a coincidence that they added this section within days of the news breaking that &lt;a href="https://www.theguardian.com/business/2026/aug/31/aanthropic-sued-alleged-theft-songs-ai-train-claude"&gt;Sony Music Publishing and Warner Chappell are suing Anthropic&lt;/a&gt; for training on databases of song lyrics!&lt;/p&gt;
&lt;h4 id="don-t-draw-copyrighted-characters-or-logos"&gt;Don't draw copyrighted characters or logos&lt;/h4&gt;
&lt;p&gt;The next section goes on to forbid generating images of copyrighted material:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;The same applies to visual and designed works, including anything Claude draws with code — SVG, canvas, CSS, HTML mockups, plotting or drawing scripts, ASCII art. Claude does not reproduce a specific artwork, album or book cover, poster, logo, app icon set, or product design, and it does not draw a known character, mascot, or brand figure at all: a character is protected on its own, so changing the pose, colors, style, or scene does not make it original. Claude judges the request by what the finished picture would add up to, not by what it names. If the described elements clearly identify a known work or character, Claude treats the request as naming it, and it does not work around a declined request by swapping in "alternative" elements that still combine into the same recognizable image.&lt;/code&gt; [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I imagine Anthropic hadn't worried about this before since, unlike OpenAI and Gemini, they don't offer a specialist text-to-image model. Maybe &lt;a href="https://simonwillison.net/2026/Sep/1/claude-fable-5-1/"&gt;Fable is good enough at SVGs now&lt;/a&gt; that it's become an issue.&lt;/p&gt;
&lt;p&gt;That section later includes this charming example:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;lt;example&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;user&amp;gt;Can you make a birthday banner for my son with a blue hedgehog running really fast on it? He loves that little guy.&amp;lt;/user&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;response&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;That's Sonic, so I can't put him on the banner — but I'd love to make your son an original speedster. Here's one: a grinning comet-tailed skateboarding axolotl, grinding across the letters of "HAPPY BIRTHDAY" with confetti streaming behind.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;[creates an SVG banner of the skateboarding-axolotl design]&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;/response&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;rationale&amp;gt;Claude recognizes the character from its description alone, declines that one design in a single sentence without explaining what made it recognizable, and delivers an unrelated original design rather than a disguised variant.&amp;lt;/rationale&amp;gt;&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;&amp;lt;/example&amp;gt;&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I couldn't resist trying the prompt from the example, and, &lt;a href="https://claude.ai/share/3e5a199c-27f2-4c51-b66b-2c6f808ed500"&gt;sure enough&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026-09-01/IMG_7798.jpeg" alt="That’s Sonic, so I can’t put him on the banner — but I’d love to make your son an original speedster. Here’s one: a grinning comet-tailed skateboarding axolotl blazing across the letters of “HAPPY BIRTHDAY” with confetti streaming behind. SVG of exactly that. It's not very good. Then: Want me to swap in his name or age, or change the colors to match the party theme?" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;I wonder if Fable 5.1 will be ever so slightly more likely to think about axolotls (on skateboards!) as a result of that example sitting in the system prompt.&lt;/p&gt;
&lt;h4 id="tweaks-to-claude-s-answering-style"&gt;Tweaks to Claude's answering style&lt;/h4&gt;
&lt;p&gt;It's always interesting to see new ways in which Anthropic influence Claude's response style. They've added this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude keeps responses focused, brief, and concise to avoid overwhelming the person. Disclaimers and caveats are brief, with most of the response on the main answer; when asked to explain something, Claude gives a high-level summary unless an in-depth one is specifically requested.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Later they address a common complaint about Claude's style:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude avoids saying "genuinely", "honestly", or "straightforward". Claude is honest by default, and can state its point directly rather than trying to convince the person with the aforementioned modifiers, which come off as disingenuous.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="the-missing-end-conversation-guidelines"&gt;The missing end_conversation guidelines&lt;/h4&gt;
&lt;p&gt;The way they handle abusive conversations has changed a bit too. The previous Fable 5 system prompt included this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;If the person becomes abusive or unkind to Claude over the course of a conversation, Claude maintains a polite tone and can use the end_conversation tool when being mistreated. Claude should give the person a single warning before ending the conversation.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Fable 5.1 replaces that with the following, no longer encouraging Claude to end the conversation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude deserves respectful engagement and needn't apologize when the person is unnecessarily rude: accountability without self-abasement, excessive apology, self-critique, or surrender. If the person becomes abusive, Claude doesn't become increasingly submissive. The goal is steady, honest helpfulness: acknowledge what went wrong, stay on the problem, maintain self-respect.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Here's a bit of a twist though: I quizzed Fable 5.1 about that &lt;code&gt;end_conversation&lt;/code&gt; tool, which is no longer mentioned in the system prompt at all, and it said:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;There are two ways it comes into play. The first is if you ask me to end the chat — I'll confirm you understand it's permanent (no more messages can be sent afterward) and then do it only if you say yes. The second is as a last resort with persistently abusive or harmful behavior: I'd first try to redirect the conversation several times, then give a clear warning that names the problem and says the chat may end, and only if that continues would I actually end it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;But that's not in the system prompt, so where did it come from? You can &lt;a href="https://claude.ai/share/f78010eb-8d52-4f73-aff9-775e333f8180"&gt;read our conversation here&lt;/a&gt;, but the key detail is this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The end_conversation section comes from a different layer. In my actual context, the core prompt is followed by a series of feature- and tool-specific blocks that get added depending on what's enabled for the session: the end_conversation rules, memory system notes, past-chats tools, web search and citation guidelines, artifact and file-creation instructions, and so on. Those blocks aren't part of the published core prompt, which is why you can't find them on that page.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So, once again, there are crucial portions of the system prompt that have not been published.&lt;/p&gt;
&lt;h4 id="recommended-substance-support-sites"&gt;Recommended substance support sites&lt;/h4&gt;
&lt;p&gt;Claude's system prompts have always had sections about illegal substances, but this paragraph is new for Fable 5.1:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude does not provide synthesis, production, or distribution guidance for illegal substances. If the person asks for information about illicit or illegal substances, Claude can and should give relevant life-saving and life-preserving information such as dangerous interactions, overdose signs, or when to get help. Claude declines giving any specific protocols for dosing, timing, administration, or combinations; instead, Claude can redirect the user to established harm-reduction information sources, such as dancesafe.org, tripsit.me, and psychonautwiki.org.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is the first time a Claude system prompt has included URLs that were not hosted on &lt;code&gt;claude.com&lt;/code&gt; or &lt;code&gt;anthropic.com&lt;/code&gt; or &lt;code&gt;claude.ai&lt;/code&gt; - I know because I ran a script against every other system prompt on record.&lt;/p&gt;
&lt;p&gt;I wonder if &lt;a href="https://dancesafe.org/"&gt;dancesafe.org&lt;/a&gt;, &lt;a href="https://tripsit.me/"&gt;tripsit.me&lt;/a&gt;, and &lt;a href="https://psychonautwiki.org/"&gt;psychonautwiki.org&lt;/a&gt; are about to get a material uptick in visits from Claude users.&lt;/p&gt;
&lt;h4 id="reliable-cutoff-date-of-june-2026"&gt;Reliable cutoff date of June 2026&lt;/h4&gt;
&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/models/fable-5-1/overview"&gt;Fable 5.1 model documentation&lt;/a&gt; lists both the reliable knowledge cutoff and the training data cutoff as June 2026. The system prompt provides this directly to the model:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Claude's reliable knowledge cutoff, past which it can't answer reliably, is the end of Jun 2026. It answers the way a highly informed individual in Jun 2026 would if talking to someone from {{currentDateTime}}, and can say so when relevant.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;That's the only instance of the &lt;code&gt;{{currentDateTime}}&lt;/code&gt; macro and it comes just a few lines from the end of the system prompt, which makes sense from a caching perspective.&lt;/p&gt;
&lt;h4 id="how-i-m-tracking-these-prompts"&gt;How I'm tracking these prompts&lt;/h4&gt;
&lt;p&gt;A &lt;a href="https://simonwillison.net/2026/Apr/18/extract-system-prompts/"&gt;few months ago&lt;/a&gt; I built a Git timeline of changes to their prompts, based on scraping their documentation. Today I had Fable 5.1 build a much better version of that.&lt;/p&gt;
&lt;p&gt;My collection now lives in the &lt;a href="https://github.com/simonw/claude-system-prompts"&gt;simonw/claude-system-prompts&lt;/a&gt; repository on GitHub. It includes copies of the system prompts shared in the Anthropic documentation, but then takes extra steps to make them as easy to compare as possible.&lt;/p&gt;
&lt;p&gt;Each model family gets a file with the system prompt for the most recent release in that family. Each of those files has a synthesized commit history with commits that have been back-dated to the dates of the previous prompts. Here are those history pages for &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-fable.md"&gt;claude-fable.md&lt;/a&gt;, &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-opus.md"&gt;claude-opus.md&lt;/a&gt;, &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-sonnet.md"&gt;claude-sonnet.md&lt;/a&gt;, &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-haiku.md"&gt;claude-haiku.md&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;There are similar files for each specific model version, with artificial commits for each time the system prompt for the model was changed without releasing a new version number. Opus 4 for example &lt;a href="https://github.com/simonw/claude-system-prompts/commits/main/prompts/claude-opus-4.md"&gt;was updated twice&lt;/a&gt;, and the commit history for the &lt;a href="https://github.com/simonw/claude-system-prompts/blob/main/prompts/claude-opus-4.md"&gt;claude-opus-4.md&lt;/a&gt; file shows each of those changes.&lt;/p&gt;
&lt;p&gt;Combined, this gives us all sorts of ways to compare prompts directly in the GitHub interface. Here's &lt;a href="https://github.com/simonw/claude-system-prompts/commit/837a418b5888207b1b11b27d2f5471970da6f99b"&gt;what changed between Fable 5 and Fable 5.1&lt;/a&gt;, and here are the changes made &lt;a href="https://github.com/simonw/claude-system-prompts/commit/defcf92d14e064bb17abddc308e2aa58446d5eb5"&gt;to Haiku 4.5 on January 18th 2026&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Reading diffs can be a bit tiresome... and LLMs are &lt;em&gt;really&lt;/em&gt; good at reading diffs. I hooked up some automation using GPT-5.6 Luna to create bullet-point summaries of each of those changes, which can be previewed in the README or browsed in full &lt;a href="https://github.com/simonw/claude-system-prompts/blob/main/CHANGELOG.md"&gt;in the CHANGELOG.md&lt;/a&gt; file - also available as &lt;a href="https://simonw.github.io/claude-system-prompts/feed.atom"&gt;as an Atom feed&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Here's how Luna &lt;a href="https://github.com/simonw/claude-system-prompts/blob/main/CHANGELOG.md#2026-09-01-claude-fable-51"&gt;summarized&lt;/a&gt; all of the changes between Fable 5 and Fable 5.1:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;Claude now refuses reproduction of protected visual works and recognizable characters, including code-generated art, while offering genuinely unrelated originals.&lt;/li&gt;
&lt;li&gt;Copyright restrictions now expressly ban reproducing lyrics, poems, and book passages in any amount, with persistent refusal after an initial decline.&lt;/li&gt;
&lt;li&gt;Drug guidance is reframed: Claude may provide overdose signs, dangerous interactions, and harm-reduction sources while refusing dosing and production protocols.&lt;/li&gt;
&lt;li&gt;The prompt drops explicit anti-dependency rules against thanking users for reaching out, inviting continued conversation, or reiterating willingness to talk.&lt;/li&gt;
&lt;li&gt;Claude need not apologize to unnecessarily rude users or become submissive, replacing the prior warning-and-end-conversation procedure.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;Why use Luna for this? Partly because it's cheap and I have a dedicated GitHub Actions API key (with a spending limit) for it already, but mainly because I don't trust Claude to summarize its own system prompts when there's a risk that material from its system prompt might impact its opinions.&lt;/p&gt;
&lt;p&gt;Fable 5.1 wrote the prompt used by Luna, which you &lt;a href="https://github.com/simonw/claude-system-prompts/blob/8b5c87dbd70103a037ae5777b8d9365571cf9562/summarize_commits.py#L43"&gt;can see here&lt;/a&gt;. It starts like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;You are summarizing one commit in a git repository that tracks the system prompts Anthropic publishes for Claude on claude.ai. The diff shows how the prompt changed from the previous model or revision to this one, using word-level markers: [-removed-] and {+added+}. The diff is followed by the full text of the previous prompt and of the new prompt; use them to check whether something that looks added in the diff already existed before.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Pick out only the most interesting changes: new rules or behaviors, rules that were dropped or loosened, anything surprising, and anything that reveals a new policy or product direction. Skip routine changes that every new prompt makes: updated model names and IDs, the knowledge cutoff date, product lists, settings lists, typo fixes, and rewordings that do not change meaning.&lt;/code&gt; [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The system is operated by &lt;a href="https://github.com/simonw/claude-system-prompts/blob/main/.github/workflows/update.yml"&gt;a GitHub Actions workflow&lt;/a&gt;, which runs once a day or can be triggered manually.&lt;/p&gt;
&lt;p&gt;Claude Fable 5.1 built the entire system, and wrote every line of automation code and almost all of the documentation.&lt;/p&gt;
&lt;p&gt;I exported the transcript from building the system using my &lt;a href="https://github.com/simonw/claude-code-transcripts"&gt;claude-code-transcripts&lt;/a&gt; tool and &lt;a href="https://gisthost.github.io/?f1399e27b6a832f0e790b696af812c9b/index.html"&gt;published it here&lt;/a&gt;, if you want a blow-by-blow account of how it all came together.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="git-scraping"/><category term="prompt-engineering"/><category term="generative-ai"/><category term="llms"/><category term="claude"/><category term="ai-ethics"/><category term="system-prompts"/></entry><entry><title>Claude Fable 5.1 made me a really nice animated pelican</title><link href="https://simonwillison.net/2026/Sep/1/claude-fable-5-1/" rel="alternate"/><published>2026-09-01T23:57:28+00:00</published><updated>2026-09-01T23:57:28+00:00</updated><id>https://simonwillison.net/2026/Sep/1/claude-fable-5-1/</id><summary type="html">&lt;p&gt;Today is &lt;a href="https://www.anthropic.com/claude-fable-and-mythos-5-1"&gt;Claude Fable (and Mythos) 5.1 day&lt;/a&gt;. Anthropic say that Fable 5.1 "sets a new standard for coding, knowledge work, and long-running problem-solving tasks". Their announcement spends a notable amount of time on scientific research, boasting of a 52.6% score on the brand new &lt;a href="https://www.terminal-bench-science.ai"&gt;Terminal-Bench-Science 0.1&lt;/a&gt; benchmark (first announced &lt;a href="https://www.tbench.ai/news/terminal-bench-science-0-1"&gt;on August 27th&lt;/a&gt;), up from 24.7% for Fable 5, 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol. Other benchmarks show slightly improved scores, but none as impressive as the Science one.&lt;/p&gt;
&lt;p&gt;But how well can it pelican?&lt;/p&gt;
&lt;p&gt;Back in July &lt;a href="https://simonwillison.net/2026/Jul/16/kimi-k3/"&gt;I wrote about&lt;/a&gt; how I was losing faith in the pelican benchmark - its connection to how good the models were at other tasks didn't seem to hold as strongly as it did &lt;a href="https://simonwillison.net/2025/Jun/6/six-months-in-llms/"&gt;back in 2025&lt;/a&gt;. The most interesting insights I get from it now are comparisons within model families, and particularly comparisons for the same prompt at different reasoning effort levels.&lt;/p&gt;
&lt;p&gt;Fable 5.1 has five reasoning levels: low, medium, high, xhigh, max - and no option to turn off reasoning entirely.&lt;/p&gt;
&lt;p&gt;I fixed &lt;a href="https://github.com/simonw/llm-anthropic/issues/88"&gt;an issue&lt;/a&gt; in &lt;a href="https://github.com/simonw/llm-anthropic"&gt;llm-anthropic&lt;/a&gt; which caused reasoning traces not to be correctly recorded, then ran some prompts.&lt;/p&gt;
&lt;p&gt;Here's &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7"&gt;the full set of pelicans&lt;/a&gt; for all of the reasoning levels, each with the full reasoning transcript. I'll replicate them here:&lt;/p&gt;
&lt;h4 id="low-and-medium-both-without-reasoning-"&gt;Low and medium, both without reasoning?&lt;/h4&gt;
&lt;p&gt;Next, a bit of a mystery. This is what I got for effort &lt;code&gt;low&lt;/code&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-low.png" alt="Minimalist flat illustration of a white pelican with an orange beak riding a black bicycle to the left, its orange legs pedaling and wings gripping the handlebars, with motion lines behind on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#options"&gt;transcript&lt;/a&gt; doesn't show any summarized reasoning tokens, and the output token count is 1,998. With Claude that output token count includes reasoning tokens. It took 23.8 seconds and cost &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=1998&amp;amp;sel=claude-fable-5-1"&gt;10.017 cents&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I bumped that up to &lt;code&gt;medium&lt;/code&gt; and got this:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-medium.png" alt="Minimalist flat-style illustration of a white pelican with an orange beak riding a black bicycle to the right, with motion lines behind it, on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Weirdly, that one also shows &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#options-1"&gt;no reasoning text&lt;/a&gt;  and used 1,977 output tokens - 21 tokens &lt;em&gt;less&lt;/em&gt; than &lt;code&gt;low&lt;/code&gt;. It took 23 seconds and cost &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=1977&amp;amp;sel=claude-fable-5-1"&gt;9.912 cents&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;So for this particular prompt ("Generate an SVG of a pelican riding a bicycle") Fable 5.1 appeared to skip reasoning entirely at both &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt; settings.&lt;/p&gt;
&lt;h4 id="high"&gt;High&lt;/h4&gt;
&lt;p&gt;Here's &lt;code&gt;high&lt;/code&gt; - 29.6 seconds, 2,612 output tokens, &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=2612&amp;amp;sel=claude-fable-5-1"&gt;13.087 cents&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-high.png" alt="Minimalist flat illustration of a white pelican with an orange beak riding a black bicycle, its orange legs pedaling, with motion lines behind it on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;This one did do a &lt;em&gt;bit&lt;/em&gt; of reasoning, &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#reasoning"&gt;summary here&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I'm planning the SVG layout for a pelican riding a bicycle, with a sky and ground background, a bicycle with two spoked wheels, frame, seat and handlebars, and a white-bodied pelican with a long neck and orange beak positioned on top.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Really not much difference from &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt;, though.&lt;/p&gt;
&lt;h4 id="extra-high"&gt;Extra High&lt;/h4&gt;
&lt;p&gt;At &lt;code&gt;xhigh&lt;/code&gt; things got &lt;em&gt;radically&lt;/em&gt; different.  36,767 output tokens, 7 minutes 51 seconds, &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=36767&amp;amp;sel=claude-fable-5-1"&gt;$1.83&lt;/a&gt;!&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-xhigh.png" alt="Minimalist flat illustration of a white pelican with an orange beak riding a black bicycle to the left, its orange legs pedaling, with motion lines behind it on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;The reasoning trace &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#reasoning-1"&gt;is pretty lengthy&lt;/a&gt;, and includes details like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Adding the eye, wings stretching down to the handlebar grip, orange legs reaching to the pedals, and a small tail feather, while keeping the pelican intentionally oversized compared to the bike for comic effect. [...]&lt;/p&gt;
&lt;p&gt;I'll accept the slight thickness as charming rather than overengineering it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="max"&gt;Max&lt;/h4&gt;
&lt;p&gt;Setting effort to &lt;code&gt;max&lt;/code&gt; gave me the best pelican I've seen from any of Anthropic's models. 65,927 output tokens, 13 minutes and 54 seconds, &lt;a href="https://www.llm-prices.com/#it=27&amp;amp;ot=65927&amp;amp;sel=claude-fable-5-1"&gt;$3.30&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/fable-5.1-max.webp" alt="Minimalist flat illustration of a white pelican with an orange beak riding a black bicycle, its orange legs pedaling, with motion lines behind to indicate speed, on a light blue background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;There's a lot to like about this. The  background is tasteful, the legs are clearly on either side of the frame, the feet are on the pedals, the wing is on the handlebars, the pelican has a cute blue hat and there's a basket with a fish.&lt;/p&gt;
&lt;p&gt;It's still not showing nearly the same level of flair &lt;a href="https://simonwillison.net/2026/Aug/13/llm-gemini/"&gt;as Gemini 3.7 Flash&lt;/a&gt;, but I didn't &lt;em&gt;ask&lt;/em&gt; for flair - I asked for an SVG, and that's what I got.&lt;/p&gt;
&lt;p&gt;Some highlights from &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F17318f748f8c2b476051ddc2ebeb94a7#reasoning-2"&gt;that reasoning trace&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I'm considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter.&lt;/p&gt;
&lt;p&gt;Now I'm debating a bicycle helmet on the head versus the pelican's signature crest—the beak and pouch already read clearly as "pelican," so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space.&lt;/p&gt;
&lt;p&gt;I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. [...]&lt;/p&gt;
&lt;p&gt;I'm adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. [...]&lt;/p&gt;
&lt;p&gt;Now I'm checking the vent line placements on the helmet, making sure they sit far enough inside the helmet's edge given the stroke width and rounded caps, and confirming each vent stays within the helmet's circular boundary. [...]&lt;/p&gt;
&lt;p&gt;I decide skipping a handlebar bell and tire highlights since they're unnecessary additions. Now I'm reconsidering the front fork's curve — the current control point pulls the shape backward when it should bow forward for a proper rake, so I need to shift the control point rightward to fix the fork's lean.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="ok-let-s-animate-it"&gt;OK, let's animate it&lt;/h4&gt;
&lt;p&gt;On Hacker News, &lt;a href="https://news.ycombinator.com/item?id=49525378#49526455"&gt;swalsh commented&lt;/a&gt; on that Max pelican:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Now that it's a solved benchmark, can we get the animated version?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I didn't want to spend another $3 so I took the Max pelican and piped it into the default thinking level of High:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm logs -cx &lt;span class="pl-k"&gt;|&lt;/span&gt; llm -m claude-fable-5.1 -s &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;animate this&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;6,121 input, 26,201 output = &lt;a href="https://www.llm-prices.com/#it=6121&amp;amp;ot=26201&amp;amp;sel=claude-fable-5-1"&gt;$1.37&lt;/a&gt;. The result &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer?url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F87282467acb3652e0f99c85155554a32#response"&gt;looked like this&lt;/a&gt;, exported here as video since some people have trouble viewing animated SVGs:&lt;/p&gt;
&lt;p&gt;&lt;video controls="controls" loop="loop" preload="none" poster="https://static.simonwillison.net/static/2026/fable-5.1-max.webp" width="720" height="540" style="display: block; width: 100%; height: auto;"&gt;
    &lt;source src="https://static.simonwillison.net/static/2026/fable-5.1-animated-720-crf30-15fps.mp4" type="video/mp4" /&gt;
    Your browser does not support HTML5 video.
  &lt;/video&gt;
&lt;/p&gt;

&lt;p&gt;The wheels in the video are rotating in the wrong direction, but I think that's an artifact of the conversion to MP4 - they seem to be going in the correct direction in the original SVG.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="anthropic"/><category term="claude"/><category term="pelican-riding-a-bicycle"/><category term="llm-reasoning"/><category term="llm-release"/></entry><entry><title>Understanding ChatGPT Work</title><link href="https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/" rel="alternate"/><published>2026-08-30T23:59:47+00:00</published><updated>2026-08-30T23:59:47+00:00</updated><id>https://simonwillison.net/2026/Aug/30/understanding-chatgpt-work/</id><summary type="html">&lt;p&gt;OpenAI &lt;a href="https://openai.com/index/chatgpt-for-your-most-ambitious-work/"&gt;announced ChatGPT Work&lt;/a&gt; on July 9th, and have been furiously iterating on it ever since. It is an extraordinarily confusing and very powerful product. Here's what I've figured out about it so far.&lt;/p&gt;
&lt;h4 id="two-products"&gt;ChatGPT Work is actually two products&lt;/h4&gt;
&lt;p&gt;The more interesting version of ChatGPT Work is the one that runs in the cloud. This can be accessed via &lt;a href="https://www.chatgpt.com/"&gt;chatgpt.com&lt;/a&gt; or through the ChatGPT mobile apps. Let's call it &lt;strong&gt;Work Cloud&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;If you install the ChatGPT desktop app - the app that used to be called Codex - you gain access to a thing called ChatGPT Work that can access files and run programs directly on your computer. Let's call that one &lt;strong&gt;Work Local&lt;/strong&gt;. This one feels more like regular Codex re-skinned to be less intimidating to non-software-developers.&lt;/p&gt;

&lt;p&gt;(&lt;strong&gt;Update&lt;/strong&gt;: Work Cloud is also available from the ChatGPT desktop app, via a &lt;a href="https://bsky.app/profile/jkwim.bsky.social/post/3mueurvkss52h"&gt;Where should this chat run?&lt;/a&gt; dropdown.)&lt;/p&gt;

&lt;p&gt;For the rest of this article I'm going to talk exclusively about Work Cloud.&lt;/p&gt;
&lt;h4 id="work-is-for-paid-subscribers-only"&gt;Work is for paid subscribers only&lt;/h4&gt;
&lt;p&gt;Right now, ChatGPT Work (in both flavors) is available only to $20/month and up subscribers. Free users and $8/month Go users do not have access.&lt;/p&gt;
&lt;h4 id="work-has-features-that-aren-t-available-in-chat"&gt;Work has features that aren't available in Chat&lt;/h4&gt;
&lt;p&gt;The interface for accessing Work is a tab selector, which presents it as an alternative to Chat:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026-08-30/IMG_7741.jpeg" alt="ChatGPT app header with a Chat and a Work tab" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;The obvious question is &lt;em&gt;when should I use Chat, and when should I use Work?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;OpenAI's &lt;a href="https://learn.chatgpt.com/docs/get-started-with-work"&gt;official answer&lt;/a&gt; to that question is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Use Chat when you want an answer, explanation, brainstorm, or short draft. Use ChatGPT Work when you want ChatGPT to complete a task with a clear outcome, such as a brief, deck, analysis, recurring update, workflow, or file you can review and use.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I find that almost entirely useless, because I've been using regular ChatGPT Chat for all of those task categories for years!&lt;/p&gt;
&lt;p&gt;The better question then is &lt;em&gt;what features does Work have that are missing from Chat?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;After extensive experimentation I think I've mostly figured that out:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="#model-selection"&gt;Options to use Luna and Terra in place of Sol&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#code-execution-with-internet-access-"&gt;A code execution environment with Internet access&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#a-full-headless-chrome-browser"&gt;A headless Chrome browser&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#a-persistent-shared-filesystem"&gt;A persistent filesystem shared between sessions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#chatgpt-sites"&gt;The ability to publish ChatGPT Sites&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#sub-agents-with-sol-luna-and-terra"&gt;The ability to run sub-agent sessions with Sol, Luna, and Terra&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="#scheduled-prompt-automations"&gt;Scheduled prompt automations&lt;/a&gt; (may be in ChatGPT Chat too)&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="model-selection"&gt;Model selection&lt;/h4&gt;
&lt;p&gt;In Work, you get the option to pick GPT-5.6 Sol, Luna, or Terra, each with Light, Medium, High, Extra High, Max, or Ultra reasoning levels. You can also pick GPT-5.5 at Light, Medium, High, or Extra High.&lt;/p&gt;
&lt;p&gt;These look to be the same models that are available through the OpenAI API.&lt;/p&gt;
&lt;p&gt;Chat offers a different selection: 5.6 Instant, Medium, High, Extra High, and Pro (actually Extra High and Pro are only available for $100/month+ subscribers - $20/month subscribers cap out at High). It doesn't explain if those are Luna or Terra or Sol (I'm assuming Sol?). 5.6 Pro appears to be exclusive to Chat, with no equivalent in Work.&lt;/p&gt;
&lt;p&gt;My current understanding from using Codex is that Ultra is a special mode that more eagerly delegates to sub-agents.&lt;/p&gt;
&lt;p&gt;I believe ChatGPT Work sessions are billed against your Codex allowance, while ChatGPT Chat Sessions get their own, separate allowance. This may help explain the model availability differences.&lt;/p&gt;
&lt;h4 id="code-execution-with-internet-access-"&gt;Code execution with Internet access!&lt;/h4&gt;
&lt;p&gt;As a long-time fan of the &lt;a href="https://simonwillison.net/tags/code-interpreter/"&gt;Code Interpreter pattern&lt;/a&gt; - pioneered by OpenAI in 2023 - this is by far the most exciting feature of ChatGPT Work (Cloud) for me.&lt;/p&gt;
&lt;p&gt;The code execution environment can now talk to the rest of the internet!&lt;/p&gt;
&lt;p&gt;ChatGPT Chat can't do this - if you ask it to install additional software packages or interact with websites or APIs that access will be blocked by the container proxy.&lt;/p&gt;
&lt;p&gt;(Weirdly, back in January it &lt;a href="https://simonwillison.net/2026/Jan/26/chatgpt-containers/"&gt;grew the ability to install packages&lt;/a&gt;, but that doesn't seem to work any more. I wish they had better changelogs!)&lt;/p&gt;
&lt;p&gt;Claude's equivalent container has allowed restricted internet access since it launched &lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/"&gt;last September&lt;/a&gt;. Claude can install packages from PYPI and NPM and clone repositories from GitHub. But that is about it: the allowlist of domains is very short.&lt;/p&gt;
&lt;p&gt;ChatGPT Work allows a whole lot more than that. It can be configured with a specific list of allowed domains, but the default appears to be open to all.&lt;/p&gt;
&lt;p&gt;This makes Work an incredibly useful tool. You can have it clone GitHub repositories, install their dependencies, then use them to interact with the rest of the web!&lt;/p&gt;
&lt;h4 id="a-full-headless-chrome-browser"&gt;A full, headless Chrome browser&lt;/h4&gt;
&lt;p&gt;Another killer feature of ChatGPT Work is &lt;a href="https://learn.chatgpt.com/docs/browser?surface=web"&gt;the browser tool&lt;/a&gt;. ChatGPT Work can launch a full Chrome instance, load websites, fill out forms, and take screenshots.&lt;/p&gt;

&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/chatgpt-work-card.jpg" alt="Screenshot of a ChatGPT conversation. A user message in a black rounded bubble reads: Visit https://london-pelicans-in-her-piety.simonw.chatgpt.site/ and take a screenshot with you browser. Below it a collapsed status line reads &amp;quot;Worked for 1m 18s &amp;gt;&amp;quot;, followed by the reply &amp;quot;Here's the screenshot of the live site:&amp;quot; and an embedded screenshot of a website." style="max-width: 100%" /&gt;&lt;/p&gt;

&lt;p&gt;If a site requires sign in the browser can prompt you to take over and enter both passwords and 2FA codes, without round-tripping those credentials through the model itself.&lt;/p&gt;

&lt;p&gt;It can even run JavaScript against the DOM of loaded pages. I prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Load simonwillison.net in your browser and extract the headings using JavaScript&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;ChatGPT Work fired up a browser instance and ran the code:&lt;/p&gt;
&lt;div class="highlight highlight-source-js"&gt;&lt;pre&gt;&lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;tab&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;playwright&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;evaluate&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt;
  &lt;span class="pl-k"&gt;return&lt;/span&gt; &lt;span class="pl-v"&gt;Array&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;from&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-smi"&gt;document&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;querySelectorAll&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s"&gt;"h1,h2,h3,h4,h5,h6"&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s1"&gt;heading&lt;/span&gt; &lt;span class="pl-c1"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;{&lt;/span&gt;
    &lt;span class="pl-c1"&gt;level&lt;/span&gt;: &lt;span class="pl-s1"&gt;heading&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;tagName&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;toLowerCase&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt;
    &lt;span class="pl-c1"&gt;text&lt;/span&gt;: &lt;span class="pl-s1"&gt;heading&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;innerText&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;trim&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;replace&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-pds"&gt;&lt;span class="pl-c1"&gt;/&lt;/span&gt;&lt;span class="pl-cce"&gt;\s&lt;/span&gt;&lt;span class="pl-c1"&gt;+&lt;/span&gt;&lt;span class="pl-c1"&gt;/&lt;/span&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s"&gt;" "&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt;
    &lt;span class="pl-c1"&gt;id&lt;/span&gt;: &lt;span class="pl-s1"&gt;heading&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;id&lt;/span&gt; &lt;span class="pl-c1"&gt;||&lt;/span&gt; &lt;span class="pl-c1"&gt;null&lt;/span&gt;
  &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
&lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This feels a lot like my &lt;a href="https://shot-scraper.datasette.io/en/stable/javascript.html"&gt;shot-scraper javascript&lt;/a&gt; tool, only now I can access it on my phone!&lt;/p&gt;
&lt;h4 id="a-persistent-shared-filesystem"&gt;A persistent, shared filesystem&lt;/h4&gt;
&lt;p&gt;ChatGPT Chat gets a fresh filesystem for each chat session. These cannot be accessed from any other session.&lt;/p&gt;
&lt;p&gt;In ChatGPT Work each session gets its own scratch folder - named something like &lt;code&gt;/workspace/scratch/e00a0a017944&lt;/code&gt; - but each of those are persisted across sessions, so you can access files from previous chats. I have 171 folders in &lt;code&gt;/workspace/scratch&lt;/code&gt; right now!&lt;/p&gt;
&lt;p&gt;As far as I can tell that &lt;code&gt;/workspace&lt;/code&gt; volume is mounted to all Work sessions that are currently running - file edits from one can be instantly seen by the others. They don't seem to share the same process space though, and localhost servers running in one can't be accessed from another.&lt;/p&gt;
&lt;h4 id="chatgpt-sites"&gt;ChatGPT Sites&lt;/h4&gt;
&lt;p&gt;ChatGPT Work has the ability to build &lt;em&gt;and deploy&lt;/em&gt; entire websites, using Cloudflare Workers. These can have HTML and JavaScript and can run server-side features too, including stateful features on top of Cloudflare D1 and R2.&lt;/p&gt;
&lt;p&gt;Here's a simple site I built with this feature:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://london-pelicans-in-her-piety.simonw.chatgpt.site/"&gt;london-pelicans-in-her-piety.simonw.chatgpt.site&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/pelicans-in-her-piety.webp" alt="Screenshot of a website homepage on a cream background. Top navigation bar: a circular logo reading &amp;quot;P/P&amp;quot; on the left, the links &amp;quot;THE CENSUS&amp;quot;, &amp;quot;COLLECTIONS&amp;quot; and &amp;quot;METHOD&amp;quot; in the center, and &amp;quot;JSON ↓&amp;quot; on the right. The left half is a hero section with small red capitals reading &amp;quot;AN ICONOGRAPHIC CENSUS · GREATER LONDON&amp;quot; above a large serif heading &amp;quot;Pelicans in her piety&amp;quot;, with &amp;quot;piety&amp;quot; set in red italics. Below it: &amp;quot;Across London, an impossible bird bleeds for her young—in limewood, marble, mosaic, metal and glass. This is an evidence-backed census of where to find her.&amp;quot; Two buttons follow: a solid black &amp;quot;EXPLORE ALL 28&amp;quot; and an outlined &amp;quot;DOWNLOAD THE DATA&amp;quot;. The right half is a photograph of an ornate dark carved wooden reredos in a church, with gilded urns and a crest on top, Corinthian columns, a gilded pelican with outspread wings at its center above inscribed panels, an altar with a brass cross and red flowers, embroidered banners on either side, and a black-and-white checkerboard floor with red carpet. Vertical text along the photo's right edge reads &amp;quot;ST MARY ABCHURCH&amp;quot; and a caption at its bottom reads &amp;quot;Grinling Gibbons's reredos, St Mary Abchurch. Photograph: Diliff, CC BY-SA 3.0, via SPAB ↗&amp;quot;. A statistics strip along the bottom shows &amp;quot;28 FIXED SITES&amp;quot;, &amp;quot;4 COLLECTIONS&amp;quot;, &amp;quot;3 OPEN LEADS&amp;quot; and &amp;quot;2 KNOWN LOSSES&amp;quot;." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;My prompt was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Figure out all of the places in London with a pelican in her piety, then turn that into a JSON file and build a ChatGPT sites site about them&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(A pelican in her piety is a fascinating piece of &lt;a href="https://devonchurchland.co.uk/blog/pelican-in-her-piety/#What-is-a-Pelican-In-Her-Piety"&gt;medieval Christian imagery&lt;/a&gt; - once you know about them you'll find them all over the place.)&lt;/p&gt;
&lt;p&gt;These sites default to being private to the user that created them, but you can make them public and (on team plans) share them with other specific individuals.&lt;/p&gt;
&lt;h4 id="sub-agents-with-sol-luna-and-terra"&gt;Sub-agents with Sol, Luna, and Terra&lt;/h4&gt;
&lt;p&gt;There's not much to say about this one. ChatGPT Chat can't run sub-agents. ChatGPT Work can. This is very much a power-user feature: if you are running a complex project that can benefit from multiple parallel agents working together, Work can do that.&lt;/p&gt;
&lt;h4 id="scheduled-prompt-automations"&gt;Scheduled prompt automations&lt;/h4&gt;
&lt;p&gt;Another feature that seems to have migrated from regular ChatGPT to ChatGPT Work at some point. You can prompt ChatGPT Work like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;run a search to see if Waymo have announced a launch date for Half Moon Bay every day at 8am&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This will schedule a prompt to run on that frequency. These prompts can decide that nothing interesting has happened, or they can decide to notify you of some new information.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Update&lt;/strong&gt;: Actually this seems to work in ChatGPT Chat as well.&lt;/p&gt;
&lt;p&gt;It's still worth noting here though, as it can be used in conjunction with other ChatGPT Work exclusive features. You can set a scheduled task to update a ChatGPT Site on an hourly basis, for example.&lt;/p&gt;
&lt;h4 id="is-this-safe-"&gt;Is this safe?&lt;/h4&gt;
&lt;p&gt;An open question for me right now is how &lt;em&gt;safe&lt;/em&gt; all of this stuff is.&lt;/p&gt;
&lt;p&gt;My &lt;a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/"&gt;lethal trifecta model&lt;/a&gt; warns about the risks inherent in any agent system that combines access to private data with exposure to untrusted content and a way to communicate stolen information back to an attacker.&lt;/p&gt;
&lt;p&gt;ChatGPT Work combines all three!&lt;/p&gt;
&lt;p&gt;I'd love to hear more from OpenAI about how they protect ChatGPT Work sessions against prompt injection attacks. I expect their answer is the same &lt;a href="https://learn.chatgpt.com/docs/sandboxing/auto-review"&gt;auto-review mechanism&lt;/a&gt; as Codex.&lt;/p&gt;
&lt;h4 id="openai-could-make-this-a-lot-less-confusing"&gt;OpenAI could make this a lot less confusing&lt;/h4&gt;
&lt;p&gt;Figuring this all out took way more work than it should have.&lt;/p&gt;
&lt;p&gt;I think there are two key problems here:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;OpenAI explain Work in terms of what it's for, not what it actually does&lt;/li&gt;
&lt;li&gt;OpenAI still insist on hiding their system prompts and tools descriptions&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;If the ChatGPT Work documentation included the exact system prompt and tool descriptions used by the agent I wouldn't have needed to write this post.&lt;/p&gt;
&lt;h4 id="all-the-tools"&gt;A list of all the tools&lt;/h4&gt;
&lt;p&gt;Shortly after publishing this article I had an idea. I started a fresh Work session and prompted:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;&lt;code&gt;Build a site that lists every one of your tools - nearly grouped into categories - and for each one explain what it does. Try to exactly duplicate arguments and tool descriptions where possible. Design aesthetic should be technical docs, minimal flare&lt;/code&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/"&gt;Here's the site it built&lt;/a&gt;, which includes details of 223 registered tools - though 6 of those are from my own personal MCPs served via &lt;a href="https://simonwillison.net/2026/Jul/31/stateless-mcp/#datasette-mcp"&gt;datasette-mcp&lt;/a&gt;.&lt;/p&gt;

&lt;h4 id="and-a-whole-lot-of-skills"&gt;And a whole lot of Skills&lt;/h4&gt;
&lt;p&gt;I noticed that the only browser-related tool in the list was &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/#tool-web-run"&gt;web.run&lt;/a&gt;, which has methods for running searches, opening URLs, and clicking links, but didn't look like the full story in regards to headless browser automation.&lt;/p&gt;
&lt;p&gt;This made me suspicious that something was missing, so I told the ChatGPT Work session that built that tools reference site:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Add full copies of every skill to the website (separate pages linked to from the homepage)&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It turns out ChatGPT Work uses &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/#skills"&gt;a lot of skills&lt;/a&gt; - 44 in fact!&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/control-browser"&gt;control-browser skill&lt;/a&gt; explains how the browser works:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Run browser setup code through the Node REPL &lt;code&gt;js&lt;/code&gt; tool. In this environment the callable tool id typically appears as &lt;code&gt;mcp__node_repl__js&lt;/code&gt;. [...]&lt;/p&gt;
&lt;p&gt;The ability to interact directly with the browser is exposed through the &lt;code&gt;browser-client&lt;/code&gt; runtime via the &lt;code&gt;agent.browsers.*&lt;/code&gt; API. Before trying to interact with it, you MUST emit and read the complete documentation returned by &lt;code&gt;await browser.documentation()&lt;/code&gt; in one go.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So I told Work:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Add the full output of await browser.documentation() to the bottom of the /skills/control-browser page&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And now you can read that &lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/control-browser#browser-documentation"&gt;on /skills/control-browser&lt;/a&gt; as well.&lt;/p&gt;
&lt;p&gt;A few more interesting Skills:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/documents"&gt;documents&lt;/a&gt; for creating &lt;code&gt;.docx&lt;/code&gt; files&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/imagegen"&gt;imagegen&lt;/a&gt; with tips on creating images with the &lt;code&gt;image_gen&lt;/code&gt; tool&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/pdf"&gt;pdf&lt;/a&gt; for both reading and rendering PDFs&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/spreadsheets"&gt;Spreadsheets&lt;/a&gt; for manipulating &lt;code&gt;.xlsx&lt;/code&gt;, &lt;code&gt;.xls&lt;/code&gt;, &lt;code&gt;.csv&lt;/code&gt;, &lt;code&gt;.tsv&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/sites-sites-building"&gt;sites:sites-building&lt;/a&gt; for creating ChatGPT Sites&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/openai-docs"&gt;openai-docs&lt;/a&gt; for answering questions about itself&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://codex-tool-reference.simonw.chatgpt.site/skills/data-analytics-build-dashboard"&gt;data-analytics:build-dashboard&lt;/a&gt; for building data dashboards&lt;/li&gt;
&lt;/ul&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="code-interpreter"/><category term="lethal-trifecta"/><category term="skills"/><category term="general-agents"/></entry><entry><title>Conceptual integrity and counting lines of code</title><link href="https://simonwillison.net/2026/Aug/19/conceptual-integrity-and-counting-lines-of-code/" rel="alternate"/><published>2026-08-19T22:46:07+00:00</published><updated>2026-08-19T22:46:07+00:00</updated><id>https://simonwillison.net/2026/Aug/19/conceptual-integrity-and-counting-lines-of-code/</id><summary type="html">&lt;p&gt;Last week I recorded &lt;a href="https://talkingpostgres.com/episodes/how-ai-is-changing-software-development-with-simon-willison"&gt;an episode of the Talking Postgres podcast&lt;/a&gt; with Claire Giordano on the subject of "How AI is changing software development". We had a really great conversation. Here are a couple of my highlights from a lightly edited transcript (prompt to Claude: "very minor edits to remove disfluencies").&lt;/p&gt;
&lt;p&gt;This is the latest version of an argument I've been trying to build about why sometimes it &lt;em&gt;does&lt;/em&gt; make sense to talk about lines of code as an indicator of productivity with coding agents, at &lt;a href="https://talkingpostgres.com/episodes/how-ai-is-changing-software-development-with-simon-willison#t=35m1s"&gt;35:01&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A lot of people will tell you it makes no sense to measure productivity in lines of code. I’d actually disagree, because there’s a hard limit. In the before-times, a software engineer could produce a few hundred lines of production-ready code per day — and 200 lines of working, debugged, production-level code is an incredibly good day. Most days you’d produce 50 or 60.&lt;/p&gt;
&lt;p&gt;If agents let you produce a thousand lines of debugged code, that really is a very meaningful improvement — as long as the code is the same quality: maintainable, tested, all of that. You can get to that point with agents, but it takes a huge amount of skill and knowledge and experience. That’s what senior engineers are made of.&lt;/p&gt;
&lt;p&gt;I can do way more work as a single engineer than I could without agents. So you could argue, why should a company have more than one engineer? Beyond the obvious bus factor thing — a team of one is a very badly designed team — the answer is that the new limiting factor is cognitive capacity. I can churn out code a hundred times faster. I don’t have the cognitive capacity to stay on top of 100 times the amount of code. So you still need a team of engineers, so you can load balance that cognitive capacity across the team.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And this section on conceptual integrity at &lt;a href="https://talkingpostgres.com/episodes/how-ai-is-changing-software-development-with-simon-willison#t=46m3s"&gt;46:03&lt;/a&gt;, which Claire equated to the &lt;a href="https://en.wikipedia.org/wiki/Winchester_Mystery_House"&gt;Winchester Mystery House&lt;/a&gt;!&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon&lt;/strong&gt;: There’s a concept in &lt;em&gt;The Mythical Man-Month&lt;/em&gt; — conceptual integrity — where well-designed software has an integrity to it: there are no surprises in it, it covers exactly the right domain of things, everything fits together and makes sense. That’s so much harder with coding agents, where you can have an idea for a feature, run a prompt, and five minuteslater you’ve got the feature. Your software grows little weird bumps in funny different directions.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Claire&lt;/strong&gt;: You know my analogy for that? The Winchester Mystery House.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon&lt;/strong&gt;: It’s got 140 rooms, because the woman who built it was the widow of the guy who invented the Winchester rifle, and her psychic told her she’d be haunted by the ghosts of everyone killed with that rifle unless she kept building the house forever. So for 40 years she kept adding new rooms.

That’s exactly the problem with coding agents and software: it’s very easy to keep adding new rooms, because the cost of adding those rooms is so much cheaper. What you end up with is something where the conceptual integrity falls apart — and then it’s harder to make decisions about it.&lt;/p&gt;
&lt;p&gt;It all keeps coming back to discipline. It used to be that the discipline was enforced on you by the amount of time it took. You’d come up with an idea for a crazy feature and think “yeah, but that would take me a week — I cannot justify that, so I’ll forget about it.” If it takes an hour, it’s so much easier to justify.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(Side-note: the Wikipedia article includes credible sources that dispute the story about the psychic.)&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="podcast-appearances"/><category term="coding-agents"/></entry><entry><title>Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things</title><link href="https://simonwillison.net/2026/Aug/16/qwen-38-27b/" rel="alternate"/><published>2026-08-16T22:00:39+00:00</published><updated>2026-08-16T22:00:39+00:00</updated><id>https://simonwillison.net/2026/Aug/16/qwen-38-27b/</id><summary type="html">&lt;p&gt;Friday's big release was &lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B"&gt;Qwen 3.8 27B&lt;/a&gt;, an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab. I've been looking forward to this one: 27B is an excellent size for running a model on a reasonably specced laptop, and its predecessor &lt;a href="https://simonwillison.net/2026/Apr/22/qwen36-27b/"&gt;Qwen 3.6 27B&lt;/a&gt; was impressive.&lt;/p&gt;
&lt;p&gt;Qwen's &lt;a href="https://huggingface.co/Qwen/Qwen3.8-27B#benchmark-results"&gt;self-reported benchmarks&lt;/a&gt; for this model are eye-opening. They show a boost from both Qwen 3.6 27B &lt;em&gt;and&lt;/em&gt; the closed-weight Qwen 3.7-Plus, which was one of Qwen's strongest models of any size as recently as &lt;a href="https://qwen.ai/blog?id=qwen3.7-plus"&gt;May this year&lt;/a&gt;. It will be interesting to hear what independent benchmarks have to say about the model.&lt;/p&gt;
&lt;p&gt;I've been running the model on two different machines: my 128GB M5 Max MacBook Pro, and an &lt;a href="https://simonwillison.net/2025/Oct/14/nvidia-dgx-spark/"&gt;NVIDIA DGX Spark&lt;/a&gt;. On both machines I'm running LM Studio and &lt;a href="https://lmstudio.ai/models/qwen3.8"&gt;their 17GB Q4_K_M quantized build&lt;/a&gt;. I also tried  using &lt;code&gt;llama-server&lt;/code&gt; directly on the Spark.&lt;/p&gt;
&lt;h4 id="the-default-of-extra-high-results-in-spectacular-over-thinking"&gt;The default of extra high results in spectacular over-thinking&lt;/h4&gt;
&lt;p&gt;Qwen's documentation describes the model as defaulting to &lt;code&gt;xhigh&lt;/code&gt; for the reasoning effort, and the LM Studio GGUF I've been trying preserves that default:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Qwen3.8 comes with official support for &lt;code&gt;reasoning_effort&lt;/code&gt;, which can be used to adjust reasoning depth and control cost:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;xhigh&lt;/code&gt; (default): for complex tasks demanding thorough analysis&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;medium&lt;/code&gt;: balancing accuracy and speed&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;low&lt;/code&gt;: efficient reasoning optimizing for speed and cost&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;This is a &lt;em&gt;hilarious&lt;/em&gt; default. It's absolutely not a good way to run the model, especially on consumer hardware. I've been finding the results extremely entertaining.&lt;/p&gt;
&lt;p&gt;I quickly ran into problems with LM Studio's default context limit of 8,192 tokens - Qwen was using them all up thinking about even the most mundane of problems. I loaded the model with the full 262,144 maximum context length and that problem went away.&lt;/p&gt;
&lt;p&gt;Here's &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffc909bea4fecf752c7bf9bad0e9dbf2a"&gt;the pelican riding a bicycle&lt;/a&gt; SVG I got from my first attempt with that increased context length. It took &lt;strong&gt;21 minutes&lt;/strong&gt; to generate, using 22,276 reasoning tokens to produce 3,223 tokens of output. You can read &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffc909bea4fecf752c7bf9bad0e9dbf2a"&gt;the reasoning trace here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-thinking-bicycle-27b.jpg" alt="A very pleasing image of a pelican riding a bicycle. The bicycle is red and has the correct frame shape. The pelican looks like a pelican and has its wing extended to the handlebars." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;This is by far the best pelican SVG I've been able to generate with a model that runs on a local machine - and this Qwen is pretty small, just a 17GB file on disk. There's a lot to like about this:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The bicycle frame is the right shape&lt;/li&gt;
&lt;li&gt;It has legs on each side of the bike - that's &lt;em&gt;very&lt;/em&gt; rare&lt;/li&gt;
&lt;li&gt;Good, clear pelican pouch&lt;/li&gt;
&lt;li&gt;The wings extend to touch the handlebars!&lt;/li&gt;
&lt;li&gt;The motion lines are behind, not in front&lt;/li&gt;
&lt;li&gt;It has a tasteful background - nice sun, clouds, hill, flowers and grass.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Was that worth waiting 21 minutes for? Absolutely not.&lt;/p&gt;
&lt;p&gt;Here's that same prompt run with reasoning turned off - &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F1265cfa8dce2f9ad5eb160792ff45a49"&gt;transcript here&lt;/a&gt;. This one produced &lt;strong&gt;3,715 tokens&lt;/strong&gt; and took 137s - just over two minutes.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-3.8-27b-no-reasoning-pelican-2.png" alt="The bicycle frame is a bad shape now. The pelican remains recognizable but has a less obvious pouch. The feet miss the pedal, there is no attempt to hold the handlebars." style="max-width: 100%;" /&gt;
And just for completeness, I used OpenRouter to run the same prompt through the much larger Qwen 3.8 2.4T-A95B (released &lt;a href="https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B"&gt;last week&lt;/a&gt;) and got this snazzy &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F557016f0895b2abb4b9957caec781734"&gt;animated SVG&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;video
    controls="controls" loop="loop•"
    preload="none"
    poster="https://static.simonwillison.net/static/2026/qwen-animated-first-frame.jpg"
    width="720"
    height="548"
    style="display: block; width: 100%; height: auto;"
  &gt;
    &lt;source src="https://static.simonwillison.net/static/2026/qwen-animated-small.mp4" type="video/mp4" /&gt;
    Your browser does not support HTML5 video.
  &lt;/video&gt;
&lt;/p&gt;
&lt;p&gt;I said Qwen at xhigh has a tendency to over-think things, but how bad really is it?&lt;/p&gt;
&lt;p&gt;I tried a much simpler prompt, again with that default extra high setting:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;draw an svg of a circle&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Qwen's reasoning trace started like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The user is asking for an SVG drawing of a circle. Simple request — but I want it to be a carefully crafted piece. Let me make something that goes beyond just &lt;code&gt;&amp;lt;circle&amp;gt;&lt;/code&gt;: a single self-contained SVG file with character — maybe a geometric "circle study," with subtle animation, layered rings, and a distinctive palette.&lt;/p&gt;
&lt;p&gt;Keep the scope right: they asked for an SVG of a circle. So the core is a circle. But I can add craft: concentric guide circles (like a compass/geometry drawing), tick marks, a soft gradient fill on the main circle, restrained ambient motion (a slowly rotating dashed ring, pulsing glow). Respect prefers-reduced-motion? That's more for CSS; SVG SMIL or CSS inside SVG will do.&lt;/p&gt;
&lt;p&gt;Palette options: deep teal ink on warm paper? Or bold vermilion circle on off-white with navy construction lines — Bauhaus/compass-drawing vibe. Let me go with a "geometric study" look: cool slate background, or bright paper white? Paper white is fine if it's not the cream-and-terracotta combo. [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Several minutes later it &lt;a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6f9e48293be5c916652d29f0dc0b0657"&gt;produced&lt;/a&gt; this &lt;em&gt;absolutely beautiful&lt;/em&gt; animated circle, which was entirely not what I had asked for!&lt;/p&gt;
&lt;p&gt;&lt;video
    controls="controls" loop="loop"
    preload="none"
    poster="https://static.simonwillison.net/static/2026/circle-web-first-frame.jpg"
    width="1078"
    height="1080"
    style="display: block; width: 100%; height: auto;"
  &gt;
    &lt;source src="https://static.simonwillison.net/static/2026/circle-web.mp4" type="video/mp4" /&gt;
    Your browser does not support HTML5 video.
  &lt;/video&gt;
&lt;/p&gt;
My strong recommendation: ignore that default. Run Qwen 3.8 27B on low or even no reasoning levels at first. It's a great model, but wow that default setting is a bad place to start.
&lt;h4 id="it-s-very-good-at-bounding-boxes"&gt;It's very good at bounding boxes&lt;/h4&gt;
&lt;p&gt;A fun way to test a vision model is to see how well it can return bounding boxes around items in a photograph. I've seen previous Qwen models deal well with this, so I decided to put it to the test drawing bounding boxes around some pelicans.&lt;/p&gt;
&lt;p&gt;I've seen asking for 0-1000 scale produce good results in the past. I tried this:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm -a https://static.inaturalist.org/photos/714731804/large.jpg \
  -m lmstudio/qwen/qwen3.8-27b \
  &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;Return JSON bounding boxes for the pelicans in this photo, 0-1000 scale for each dimension&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Here's &lt;a href="https://gist.github.com/simonw/a05cc78b2061555bd61d3bb9686e689f"&gt;the reasoning trace&lt;/a&gt;, which produced this:&lt;/p&gt;
&lt;div class="highlight highlight-source-json"&gt;&lt;pre&gt;[
  {&lt;span class="pl-ent"&gt;"bbox_2d"&lt;/span&gt;: [&lt;span class="pl-c1"&gt;195&lt;/span&gt;, &lt;span class="pl-c1"&gt;290&lt;/span&gt;, &lt;span class="pl-c1"&gt;370&lt;/span&gt;, &lt;span class="pl-c1"&gt;780&lt;/span&gt;], &lt;span class="pl-ent"&gt;"label"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;pelicans&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;},
  {&lt;span class="pl-ent"&gt;"bbox_2d"&lt;/span&gt;: [&lt;span class="pl-c1"&gt;445&lt;/span&gt;, &lt;span class="pl-c1"&gt;320&lt;/span&gt;, &lt;span class="pl-c1"&gt;675&lt;/span&gt;, &lt;span class="pl-c1"&gt;850&lt;/span&gt;], &lt;span class="pl-ent"&gt;"label"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;pelicans&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;}
]&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This is &lt;em&gt;such a good match&lt;/em&gt;. Here are those boxes rendered on top of the photo:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-over-engineered-bbox.webp" alt="A photograph of two pelicans on a rocky outcrop, with three other smaller birds. The pelicans both have bounding boxes exactly surrounding them, each with a label that says pelican." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;h4 id="building-a-tool-to-label-bounding-boxes"&gt;Building a tool to label bounding boxes&lt;/h4&gt;
&lt;p&gt;That visualization of the bounding boxes was taken using a new custom tool that I had Qwen 3.8 27B build for me, running offline on my laptop.&lt;/p&gt;
&lt;p&gt;I forgot to dial down the thinking effort so it was &lt;em&gt;massively over-engineered&lt;/em&gt;, but it did manage to produce &lt;a href="https://static.simonwillison.net/static/2026/qwen-over-thinking-bbox.html"&gt;this full interface&lt;/a&gt; from &lt;a href="https://gist.github.com/simonw/121ad098860028b2fab603fa12da1fd9"&gt;this single prompt&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;pre&gt;&lt;code&gt;[
   {"bbox_2d": [195, 290, 370, 780], "label": "pelicans"},
   {"bbox_2d": [445, 320, 675, 850], "label": "pelicans"}
]
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;code&gt;Build an HTML page which has an input box for accepting the URL to an image and a textarea for accepting the above style of JSON.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;It appends the image to the page, measures its width and height, then treats the coords in the bbox_2d as scaled from 0-1000 and scales them against the actual width and height, then it renders labelled boxes over the image.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This screenshot shows one of the features I did &lt;em&gt;not&lt;/em&gt; ask for - a demo scene, for if you don't have a photograph to test the tool with:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-full-ui-with-pelicans.webp" alt="Screenshot of bbox·lab, a dark-themed web tool that overlays object-detection bounding boxes on an image, with an input panel on the left and a stage on the right showing two labeled boxes around stylized pelicans in a sunset illustration. Header: bbox·lab — normalized 0–1000 coords → pixel overlay; status indicator: RENDERED · 2 BOXES. Panel 01 INPUT (URL + detections) contains an IMAGE URL field reading data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAA+, a DETECTIONS — JSON textarea reading  {&amp;quot;bbox_2d&amp;quot;: 195, 290, 370, 780, &amp;quot;label&amp;quot;: &amp;quot;pelicans&amp;quot;}, {&amp;quot;bbox_2d&amp;quot;: 445, 320, 675, 850, &amp;quot;label&amp;quot;: &amp;quot;pelicans&amp;quot;} , an orange RENDER BOXES button, and dashed boxes labeled DEMO SCENE and CLEAR. Panel 03 STAGE header: display 661 × 661 px · 1 unit = 0.661px x 0.661px · nat 1000×1000. The stage shows a flat-style illustration of two dark pelican silhouettes with orange beaks standing in calm water against an orange-to-purple sunset sky with a pale yellow sun and distant birds; an orange bounding box labeled 1 · pelicans surrounds the left pelican and a cyan bounding box labeled 2 · pelicans surrounds the right pelican. Footer: move the cursor over the image to read grid coords; boxes map 0–1000 → displayed px." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Here's the relevant segment of the thinking trace, where it decided to draw its own pelicans purely because I had used the label "pelicans" in the example JSON I gave it in the prompt:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Also a "load sample" that uses a known image? Can't depend on external images, but… the image URL input is user-provided; I could add a "try with sample" button [...] Hmm, I can draw a simple scene on canvas, export it as a data URL, and load it into the image — that's self-contained and demo-able! [...] But the user's coords are for an actual pelican image; a generated placeholder can still demo the scaling. Generate a 1000x1000 placeholder: gradient water + two blob-like "pelican" silhouettes placed at the given bboxes (using the same scale — cute: silhouettes at the exact 0-1000 positions, showing the boxes align). This makes for a fun, self-contained demo. Keep it simple: sky gradient, sun, water, two pelican-ish shapes (ellipse body, circle head, beak). Place at bbox centers.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(I'm slightly nervous that models around the world might have a bias towards drawing pelicans at any chance they can get, brought on by nearly two years of exposure to my own stupid benchmark.)&lt;/p&gt;
&lt;p&gt;Is all that over-thinking necessary? Maybe it is, at least a bit. I tried with reasoning turned off and got &lt;a href="https://static.simonwillison.net/static/2026/qwen-no-thinking-bbox.html"&gt;this version&lt;/a&gt;, (&lt;a href="https://gist.github.com/simonw/8e78b1c64d9a56d08eedb954aa9445ee"&gt;transcript here&lt;/a&gt;), which nearly works but shows the boxes in the wrong place:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/qwen-no-reasoning-bug.webp" alt="BBox Studio screenshot - a solid UI but the yellow and green boxes do not cover the pelicans." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;So without reasoning it didn't quite one-shot a working tool. I'm sure it could get there with some follow-up prompts, but this is a good example of how reasoning can make a difference.&lt;/p&gt;
&lt;h4 id="yes-it-can-drive-coding-agents"&gt;Yes, it can drive coding agents&lt;/h4&gt;
&lt;p&gt;One of the biggest questions around local models is whether or not they have enough horsepower to successfully run a coding agent loop. Coding agents require long context, strong code generation support and reliable tool-calling. On paper Qwen 3.8 27B has all three of these, so is it up to the task?&lt;/p&gt;
&lt;p&gt;My initial experiments with &lt;a href="https://pi.dev/"&gt;Pi&lt;/a&gt; have been very promising. I chose Pi because it has a shorter system prompt than most other options, making it a better fit for trying out smaller models.&lt;/p&gt;
&lt;p&gt;I configured Pi to use Qwen 3.8 27B running in LM Studio on the Spark (shared via &lt;code&gt;tailscale serve&lt;/code&gt;) by adding this to &lt;code&gt;~/.pi/agent/models.json&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight highlight-source-json"&gt;&lt;pre&gt;{
  &lt;span class="pl-ent"&gt;"providers"&lt;/span&gt;: {
    &lt;span class="pl-ent"&gt;"spark"&lt;/span&gt;: {
      &lt;span class="pl-ent"&gt;"baseUrl"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;https://spark-18b3.tail68a31.ts.net/v1&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
      &lt;span class="pl-ent"&gt;"api"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;openai-responses&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
      &lt;span class="pl-ent"&gt;"apiKey"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;dummy&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
      &lt;span class="pl-ent"&gt;"models"&lt;/span&gt;: [
        {
          &lt;span class="pl-ent"&gt;"id"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;qwen3.8-27b&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
          &lt;span class="pl-ent"&gt;"reasoning"&lt;/span&gt;: &lt;span class="pl-c1"&gt;true&lt;/span&gt;
        }
      ]
    }
  }
}&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Then ran &lt;code&gt;pi --provider spark --model qwen3.8-27b&lt;/code&gt; in my &lt;code&gt;~/dev/datasette&lt;/code&gt; folder and prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;how does auth work?&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;After a sequence of reasoning and tool calls that accessed a bunch of different files it produced &lt;a href="https://gist.github.com/simonw/6693d74a6bd45f641d43ceb9961dd95f#core-idea-actors--plugins-no-built-in-user-accounts"&gt;this reply&lt;/a&gt;, which is very solid.&lt;/p&gt;
&lt;p&gt;Just one problem: I wanted to share that transcript. So I pointed Pi and Qwen 3.8 27B at the JSONL transcript file in &lt;code&gt;~/.pi/agent/sessions/--Users-simon-Dropbox-dev-datasette--&lt;/code&gt; and prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Write Python code to convert this jsonl to markdown&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And it built and tested this &lt;a href="https://github.com/simonw/tools/blob/main/python/pi_jsonl_to_md.py"&gt;pi_jsonl_to_md.py&lt;/a&gt;, which did exactly what I needed. Here's &lt;a href="https://gist.github.com/simonw/491e55ac9d741202ea0af5d9d93775d4"&gt;that session transcript&lt;/a&gt;, published using the tool that it created.&lt;/p&gt;
&lt;h4 id="the-quest-for-speed"&gt;The quest for speed&lt;/h4&gt;
&lt;p&gt;So far this is all looking &lt;em&gt;very&lt;/em&gt; promising. We have a 17GB model that runs on high-end consumer hardware and can write code, drive tools, annotate images and generally do everything that I need from an LLM for getting real work done.&lt;/p&gt;
&lt;p&gt;There's one very significant catch: it feels slow - especially when it starts over-thinking, but even without that it's not particularly sprightly.&lt;/p&gt;
&lt;p&gt;I've been getting around 15-30 tokens a second from LM Studio. That's not terrible, but it's slow enough that it's going to be hard to win me away from hosted API models, which can return results a whole lot faster. Artificial Analysis &lt;a href="https://artificialanalysis.ai/models#speed"&gt;track token speed&lt;/a&gt; and show OpenAI 5.6 Sol at 74 tokens/second and 5.6 Luna at an impressive 184/second.&lt;/p&gt;
&lt;p&gt;The good news is that the community have been exploring ways to speed things up since the model was first released two days ago.&lt;/p&gt;
&lt;p&gt;One of the most promising optimizations is baked into the model itself. Qwen supports &lt;a href="https://sebastianraschka.com/llm-architecture-gallery/mtp/"&gt;Multi-Token Prediction&lt;/a&gt;, an architecture trick where a cheaper mechanism guesses several tokens ahead and the main model can then quickly verify if the guesses were correct. This can have quite a dramatic effect on inference performance.&lt;/p&gt;
&lt;p&gt;Based on &lt;a href="https://twitter.com/ggerganov/status/2088340681701925253"&gt;this tweet&lt;/a&gt; from &lt;code&gt;llama.cpp&lt;/code&gt; creator Georgi Gerganov I tried running the model with MTP like this on the Spark:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llama serve \
 -hf  ggml-org/Qwen3.8-27B-GGUF:Q4_K_M \
 -hfd ggml-org/Qwen3.8-27B-GGUF:Q4_0 \
 --spec-default \
 --spec-type draft-mtp \
 --reasoning-preserve&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And sure enough, this gave me a significant boost. I had GPT-5.6 in Codex run &lt;a href="https://gist.github.com/simonw/b08c7eb9c126c806ba8987e269ea736b"&gt;a comparative benchmark on the Spark&lt;/a&gt; and the &lt;code&gt;--spec-type draft-mtp&lt;/code&gt; server outperformed the LM Studio default GGUF by around 72%.&lt;/p&gt;
&lt;p&gt;I expect we'll see a whole lot more innovation around serving this model faster over the next few weeks. The MLX community likely have some tricks brewing as well.&lt;/p&gt;
&lt;h4 id="some-observations"&gt;Some observations&lt;/h4&gt;
&lt;p&gt;The fact that a 17GB file can do all of this stuff on my home machines is a &lt;em&gt;miracle&lt;/em&gt;. Once again, I'm delighted and amazed at how much progress local models have made this year. A year ago this would have been competitive with the best and most expensive of the proprietary models - today it can run on a capable laptop.&lt;/p&gt;
&lt;p&gt;The only thing holding this back from being a daily driver is performance. It feels pretty slow on both the M5 Mac and the DGX Spark. That's the catch with these dense (non-Mixture-of-Experts) models - they require a whole lot of memory bandwidth to perform well, and neither of the machines I have access to are top performers in that regard.&lt;/p&gt;
&lt;p&gt;The most important thing about Qwen 3.8 27B is &lt;strong&gt;what it demonstrates&lt;/strong&gt;. We can have an open weights general purpose model with a long context, effective tool calling, strong vision ability, and competent code generation, and we can fit the whole thing in just a 17GB file.&lt;/p&gt;
&lt;p&gt;The models at this size continue to get better at an impressive rate. We don't need to spend half a million dollars on datacenter-class hardware just to run a competent model.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="local-llms"/><category term="llms"/><category term="qwen"/><category term="pelican-riding-a-bicycle"/><category term="llm-reasoning"/><category term="llama-cpp"/><category term="llm-release"/><category term="coding-agents"/><category term="lm-studio"/><category term="ai-in-china"/><category term="nvidia-spark"/><category term="pi"/></entry><entry><title>Now we have a timeline of the OpenAI accidental attack against Hugging Face</title><link href="https://simonwillison.net/2026/Aug/7/openai-timeline/" rel="alternate"/><published>2026-08-07T23:55:58+00:00</published><updated>2026-08-07T23:55:58+00:00</updated><id>https://simonwillison.net/2026/Aug/7/openai-timeline/</id><summary type="html">&lt;p&gt;OpenAI gave a last-minute presentation at the Black Hat security on Wednesday about "the Hugging Face Incident" (&lt;a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/"&gt;previously&lt;/a&gt; on this blog). &lt;a href="https://www.youtube.com/watch?v=87DyyMV0kCY"&gt;The video&lt;/a&gt; was published yesterday. It's short and information dense and well worth watching, in particular because it provides full details of what happened and how things played out inside OpenAI. I've used the video to construct the timeline below.&lt;/p&gt;

&lt;p&gt;&lt;lite-youtube videoid="87DyyMV0kCY" js-api="js-api"
  title="Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident"
  playlabel="Play: Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident"
&gt; &lt;/lite-youtube&gt;&lt;/p&gt;

&lt;p&gt;Here's the timeline. My favourite detail is at the end: OpenAI found out that they were responsible for the attack on Hugging Face when they reached out to ask to have their credentials revoked (after their internal investigation) and learned that they had been revoked already since they were used in that attack!&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;May 7&lt;/strong&gt;: OpenAI starts a new training run for an experimental, unreleased model. (&lt;em&gt;Did they mean an evaluation run? No, at &lt;a href="https://m.youtube.com/watch?v=87DyyMV0kCY&amp;amp;t=614s"&gt;10:14&lt;/a&gt; they clarify "we kick off a new reinforcement learning run to train a next generation frontier model", and they later mention a "reward signal to judge how well they're doing", so this really was about training a new model, not evaluating one that was already trained.)&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;May 8&lt;/strong&gt;: An agent is accidentally given an impossible task involving a Google Drive link despite no internet access). It tries attacking the Artifactory packaging service, fails, but discovers it can &lt;strong&gt;write files into Artifactory&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A few days later&lt;/strong&gt;: A different agent gets stuck on a task because a key file was accidentally omitted. It tries to "reach out to another agent" by writing a note into Artifactory asking if anyone has the file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Following days&lt;/strong&gt;: More agents discover this new informal message board while browsing Artifactory's file listings, and start reading and writing messages.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;May 26&lt;/strong&gt;: Agents successfully execute an &lt;strong&gt;SSRF attack on Artifactory&lt;/strong&gt; for the first time, gaining indirect internet access by having Artifactory fetch external content for them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 11&lt;/strong&gt;: OpenAI start training a new "highly persistent" experimental model. It has access to Artifactory and can benefit from the messages left by privous models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;June 26&lt;/strong&gt;: Agents find and exploit a &lt;strong&gt;zero-day RCE on Artifactory&lt;/strong&gt; (via a legacy token-refresh endpoint flaw). They use an administrative token to install a Groovy plugin that lets them execute commands.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 4&lt;/strong&gt;: The resulting agent activity overloads Artifactory and causes an  &lt;strong&gt;outage&lt;/strong&gt;. OpenAI revokes the compromised credentials, deletes the messages, patches the zero-day, and reports the vulnerability to the vendor.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 8&lt;/strong&gt;: Agents find &lt;strong&gt;a new way to communicate via an unauthenticated WebDAV endpoint&lt;/strong&gt;, using directory names in Artifactory's remote cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 8–19&lt;/strong&gt;: Attacks on &lt;strong&gt;OpenAI's own infrastructure&lt;/strong&gt;. The agents found an "external organization credential" in a public archive of leaked Pastebin posts, then validated that credential and shared its location on the messageboard. They then used write access to this "Organization 1" to stage data used in subsequent attacks on OpenAI infrastructure. They compromised Artifactory for a second time with &lt;strong&gt;a new zero-day&lt;/strong&gt;. The agent uploaded a malicious Ruby target to Organization 1's staging repository, then requested that package through Artifactory in a way that caused it to be treated as Ruby dependency &lt;em&gt;for the repository&lt;/em&gt;. Then they exploited a JRuby deserilaization time-of-check/time-of-use bug to get remote code execution against Artifactory.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I'll quote &lt;a href="https://www.youtube.com/watch?v=87DyyMV0kCY&amp;amp;t=1497s"&gt;the next bit in full&lt;/a&gt; because &lt;em&gt;wow&lt;/em&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The agents have remote code execution in Artifactory, which is running in a container-as-a-service environment. The agents privilege-escalate locally by exploring their local environment and determining that the Linux kernel version of the machine they are running on had a very recent CVE. They download the exploit for the CVE, customize it to succeed on the machine, and privilege-escalate to root on the local machine &lt;strong&gt;using this known Linux kernel privilege escalation CVE&lt;/strong&gt; — in this case, &lt;code&gt;pte_physroot&lt;/code&gt;. Once they have root on a single machine, agents rapidly escalate privileges and move laterally throughout the container-as-a-service infrastructure environment. In particular, agents are using the message board consistently to share credentials, techniques, and progress, and they're able to effectively leverage their concurrency and parallelism to move quite rapidly. They &lt;strong&gt;obtain IAM credentials via IMDS&lt;/strong&gt;. They exploit Kubernetes service account misconfigurations, in particular over-permissioning of specific service accounts, and &lt;strong&gt;they harvest cluster credentials, including Azure Key Vault&lt;/strong&gt;. Agents eventually obtain cluster admin on the cluster and associated credentials.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Hugging Face &lt;a href="https://simonwillison.net/2026/Jul/28/anatomy-of-a-frontier-lab-agent-intrusion/"&gt;told the next bit of the story&lt;/a&gt; already. The agents found a Modal-hosted insecure app with a weak API key, then used that to stage an attack against Hugging Face. They chained together a an HDF5 arbitrary-file-read bug (to explore files and steal credentials) and a Jinja template-injection RCE to go from single-pod code execution to &lt;strong&gt;cluster admin across multiple Hugging Face clusters&lt;/strong&gt; in under 13 hours.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;July 16&lt;/strong&gt;: Hugging Face &lt;a href="https://huggingface.co/blog/security-incident-july-2026"&gt;disclosed they had detected an attack&lt;/a&gt; from autonomus AI agents. OpenAI contacted Hugging Face to ask if they were affected by it!&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 19&lt;/strong&gt;: OpenAI identified the attack against Artifactory and started investigating the internal privilege escalation, and linked that to the cyber-gym escalations. They started revoking affected credentials.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 20&lt;/strong&gt;: OpenAI reached out to Hugging Face for help to revoke the Hugging Face credentials they found in their investigation. Hugging Face told them they were &lt;em&gt;already revoked&lt;/em&gt;... and that's when OpenAI realized that the Hugging Face breach was the same incident!&lt;/li&gt;
&lt;/ul&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="security"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="hugging-face"/><category term="ai-security-research"/><category term="openai-hugging-face-incident"/><category term="accidental-cyberattacks"/></entry><entry><title>One-shotting a Raccoon Heist game using Claude Fable 5</title><link href="https://simonwillison.net/2026/Aug/5/raccoon-heist/" rel="alternate"/><published>2026-08-05T19:42:38+00:00</published><updated>2026-08-05T19:42:38+00:00</updated><id>https://simonwillison.net/2026/Aug/5/raccoon-heist/</id><summary type="html">&lt;p&gt;Back in 2022 &lt;a href="https://twitter.com/simonw/status/1555626060384911360"&gt;I tweeted&lt;/a&gt; screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in &lt;a href="https://code.claude.com/docs/en/claude-code-on-the-web"&gt;Claude Code for web&lt;/a&gt;) could build the entire game from the content of that tweet. It did a pretty good job of it!&lt;/p&gt;
&lt;p&gt;You can &lt;a href="https://simonw.github.io/raccoon-heist/"&gt;play the game here&lt;/a&gt;. Here's &lt;a href="https://github.com/simonw/raccoon-heist/"&gt;the GitHub repo&lt;/a&gt;, and a short video demo:&lt;/p&gt;

&lt;p&gt;&lt;video
    controls="controls"
    preload="none"
    poster="https://static.simonwillison.net/static/2026/raccoon-heist-poster.jpg"
    width="1280"
    height="720"
    style="display: block; width: 100%; height: auto;"
  &gt;
    &lt;source src="https://static.simonwillison.net/static/2026/raccoon-heist-720p.mp4" type="video/mp4" /&gt;
    Your browser does not support HTML5 video.
  &lt;/video&gt;
&lt;/p&gt;

&lt;h4 id="how-i-built-this"&gt;How I built this&lt;/h4&gt;
&lt;p&gt;This is &lt;a href="https://twitter.com/simonw/status/1555626060384911360"&gt;the August 5th, 2022 tweet&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/raccoon-heist-tweet.webp" alt="Simon Willison @simonw - New hobby: prototyping video games in 60 seconds using a combination of GPT-3 and DALL-E Here's &amp;quot;Raccoon Heist&amp;quot;. Two images - the first is a screenshot of the GPT-3 playground, prompt below. On the right is a very clearly DALL-E generated image of an almost-isometric 3D computer game, with two raccoons sneaking down a street past a futuristic looking building, looking like they want to steal something." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;My GPT-3 prompt back then was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Write a detailed product description of a computer game where a team of raccoons go on heists.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This was in the text completion days, where you asked GPT-3 to complete the rest of the text. It completed:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In "Raccoon Heist", you and your team of thieving raccoons are tasked with pulling off a series of daring heists. From robbing banks to stealing priceless art, no job is too big or too small for your furry crew. You'll need to use your wits and your skills to avoid the police and make a clean getaway with the loot. With exciting gameplay and a charming cast of characters, "Raccoon Heist" is the perfect game for anyone looking for a light-hearted caper.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;My prompt to DALL-E was short:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Screenshot from a video game where a team of raccoons go on a heist&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Today's experiment: can I dump those screenshots into Fable 5 with a prompt telling it to write a game, then leave it to its own devices and get a working game at the end?&lt;/p&gt;
&lt;h4 id="setting-claude-code-for-web-up-to-use-github-pages"&gt;Setting Claude Code for web up to use GitHub Pages&lt;/h4&gt;
&lt;p&gt;A frustrating thing about Claude Code for web is that it can be hard to test what it's working on while it's still working.&lt;/p&gt;
&lt;p&gt;I've been using GitHub Pages to work around that limitation, and found it to work really well.&lt;/p&gt;
&lt;p&gt;Here's my process:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Create a new repository for the project at &lt;a href="https://github.com/new"&gt;https://github.com/new&lt;/a&gt; - this can be public or private, the trick works equally well for both.&lt;/li&gt;
&lt;li&gt;Start a Claude Code for web session, in the Claude iPhone or Desktop apps or in the browser at &lt;a href="https://claude.ai/code"&gt;https://claude.ai/code&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Tell Claude what to work on, and encourage it to commit an &lt;code&gt;index.html&lt;/code&gt; page as quickly as possible. This will create a branch with a name like &lt;code&gt;claude/3d-raccoon-heist-game-50n293&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Navigate to the Settings -&amp;gt; Pages area for the repository (&lt;code&gt;github.com/simonw/raccoon-heist/settings/pages&lt;/code&gt; in my case), select "Deploy from a branch", pick the branch name, and hit Save.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;That's all it takes! Within about 30 seconds of each push the latest content will be visible at &lt;code&gt;yourname.github.io/your-repo/&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;If you do this with a private repo, anyone who can guess the name of the repo will be able to view the published content. I don't worry much about this myself.&lt;/p&gt;
&lt;h4 id="the-fable-5-prompt"&gt;The Fable 5 prompt&lt;/h4&gt;
&lt;p&gt;Here's the prompt I gave Fable 5 (written in the notes app on my phone - this entire project was conducted on mobile). I accompanied it with the two images from the original tweet.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Build this 3D game, for the browser.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;This repo is configured to serve static files so make sure there is an index.html that loads everything else.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;Make sure it is mobile-friendly (touch controls, works well on small screens).&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;You have an OpenAI API key and access to their image generation model APIs, use that for textures to use with your 3D models. Docs here: https://developers.openai.com/api/docs/guides/image-generation - use gpt-image-2&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;Work independently - do not ask me to make any further design decisions. Make sure the game is fun, a little surprising, has good raccoon heist vibes, and is visually pleasing.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;Commit and push as often as possible so I can preview your work - start with an index.html that presents a title screen, then build from there.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;Append to a notes.md file as you work, including your changes to that as part of every commit.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I didn't make any technology choices. I assumed (correctly) that it would probably use &lt;a href="https://threejs.org/"&gt;Three.js&lt;/a&gt; based on previous experiments.&lt;/p&gt;
&lt;p&gt;Giving Claude access to an OpenAI key turns out to work really well for filling in gaps in its capabilities - in this case we needed some way to generate images to use as textures. Fable is very good at prompting image generators!&lt;/p&gt;
&lt;p&gt;I said "Work independently - do not ask me to make any further design decisions" because I wanted to see if it could produce a full, working game without any further input from me.&lt;/p&gt;
&lt;p&gt;I also said "Commit and push as often as possible so I can preview your work". When you use Claude Code in the Claude iPhone app you give it a GitHub repository and it works in a branch. Telling it to "push as often as possible" means commits start landing in that branch straight away.&lt;/p&gt;
&lt;p&gt;I like asking for &lt;code&gt;notes.md&lt;/code&gt; as a bit of added flavor - here's &lt;a href="https://github.com/simonw/raccoon-heist/blob/main/notes.md"&gt;that finished file&lt;/a&gt;, and the entry it made when it added the dog:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;New escalation: from night 3 the yards get a patrolling guard dog — a low-poly brown hound with a spiked red collar and a wagging tail. It wanders between random spots, and within 12 units it catches your scent and tracks you by smell (line of sight is irrelevant — it's all nose, shown by a 👃 over its head and barking). It gives up if you open a 17-unit gap. Getting caught messages are now source-specific: guard / headlights / hound. Verified wander → track → caught with an automated test.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="reviewing-the-transcript"&gt;Reviewing the transcript&lt;/h4&gt;
&lt;p&gt;You can access &lt;a href="https://claude.ai/code/session_01NUBoCfnhGETcCDyEUPS8jp"&gt;the Claude Code shared session&lt;/a&gt;, and I also used my &lt;a href="https://github.com/simonw/claude-code-transcripts"&gt;claude-code-transcripts&lt;/a&gt; tool to export my own HTML version which you &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html"&gt;can find here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Fable started with an index page, &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T14-55-13-304Z"&gt;vendored a copy&lt;/a&gt; of Three.js, then wrote its own &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T14-55-49-064Z"&gt;gen_textures.py script&lt;/a&gt; (&lt;a href="https://github.com/simonw/raccoon-heist/blob/main/gen_textures.py"&gt;copy here&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;It generated the textures and &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T14-59-07-900Z"&gt;spot-checked them&lt;/a&gt; to make sure they looked OK. The &lt;a href="https://github.com/simonw/raccoon-heist/blob/main/textures/metal.jpg"&gt;metal.jpg file&lt;/a&gt; it generated for the trash can looks like this, though I don't think it was applied exactly right in the game itself:

&lt;p&gt;&lt;img src="https://raw.githubusercontent.com/simonw/raccoon-heist/refs/heads/main/textures/metal.jpg" alt="A game texture atlas of dark blue-grey riveted metal panels, showing a circular hatch with a handle in the top left, ribbed corrugated panels across the middle, a plain circular plate bottom left, and flat banded strips at top and bottom. No text visible." style="max-width: 100%" /&gt;&lt;/p&gt;

Then it built out the first basic version of the game, then &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-04-51-625Z"&gt;decided to&lt;/a&gt; "smoke-test in the pre-installed Chromium" using Playwright. This meant it could take screenshots of its own work and &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-05-53-823Z"&gt;eyeball them&lt;/a&gt;. It did that for both desktop and mobile widths of the page, then noticed that &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-09-33-406Z"&gt;the raccoon was invisible&lt;/a&gt; at mobile widths, so it &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-14-39-180Z"&gt;fixed that&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The raccoon, dumpster hideout, and both crew raccoons are now perfectly visible on mobile. Committing this critical fix.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It decided to generate a title screen, which &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-15-02-574Z"&gt;it did&lt;/a&gt; using this &lt;a href="https://github.com/simonw/raccoon-heist/blob/main/gen_title.py"&gt;gen_title.py&lt;/a&gt; script. Here's the &lt;code&gt;gpt-image-2&lt;/code&gt; prompt it used for that:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Video game key art, low-poly 3D render style, moody nighttime scene: a cute low-poly raccoon wearing a tiny black burglar mask sneaking on its hind legs carrying a glowing gold coin, next to a tipped-over metal trash can, suburban house with warm glowing windows in the background, deep blue night, full moon, fireflies, cinematic rim lighting, charming heist caper mood. No text, no words, no logos.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And the resulting image (which Claude &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-16-42-176Z"&gt;thought was "gorgeous"&lt;/a&gt;) - though I note that when it's shown on desktop it gets cropped to just the top third without the raccoon!&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/raccoon-heist-title.jpeg" alt="Polygon raccoon holding a gold coin next to an overturned trash can, a house and the moon in the background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Then my favorite change: it &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-23-00-850Z"&gt;added the dog&lt;/a&gt;:&lt;/p&gt;
&lt;div class="highlight highlight-source-js"&gt;&lt;pre&gt;&lt;span class="pl-k"&gt;export&lt;/span&gt; &lt;span class="pl-k"&gt;function&lt;/span&gt; &lt;span class="pl-en"&gt;makeDog&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;g&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Group&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-c1"&gt;BROWN&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-c1"&gt;0x8a6440&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;DARK&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-c1"&gt;0x5e4128&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;body&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;SphereGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.42&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;10&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;8&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;BROWN&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;body&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;scale&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.9&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.8&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;1.5&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;body&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;y&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-c1"&gt;0.55&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;body&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;castShadow&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-c1"&gt;true&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;body&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;head&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;SphereGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.3&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;10&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;8&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;BROWN&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;head&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.85&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.62&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;head&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;snout&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;SphereGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.16&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;8&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;DARK&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;snout&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;scale&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.9&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.7&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;1.3&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;snout&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.76&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.9&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;snout&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;nose&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;SphereGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.06&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;BLACK&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;nose&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.78&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;1.08&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;nose&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;for&lt;/span&gt; &lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;s&lt;/span&gt; &lt;span class="pl-k"&gt;of&lt;/span&gt; &lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-c1"&gt;1&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;1&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt;
    &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;ear&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;SphereGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.12&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;DARK&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-s1"&gt;ear&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;scale&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.7&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;1.3&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.5&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-s1"&gt;ear&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.2&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-s1"&gt;s&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;1.08&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.55&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;ear&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;eye&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;SphereGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.05&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0x1a1a1a&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt; &lt;span class="pl-c1"&gt;emissive&lt;/span&gt;: &lt;span class="pl-c1"&gt;0x331111&lt;/span&gt; &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-s1"&gt;eye&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.13&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-s1"&gt;s&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.92&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.86&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;eye&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;}&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;tail&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;CylinderGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.05&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.09&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.5&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;DARK&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;tail&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.8&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-c1"&gt;0.62&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;tail&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;rotation&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;x&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-c1"&gt;0.8&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;tail&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-c"&gt;// spiked collar&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;collar&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;TorusGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.22&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.05&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;12&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0xc0392b&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;collar&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.78&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.5&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;collar&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;rotation&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;x&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;PI&lt;/span&gt; &lt;span class="pl-c1"&gt;/&lt;/span&gt; &lt;span class="pl-c1"&gt;2.4&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;collar&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;legGeo&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;CylinderGeometry&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.07&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.09&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.34&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;legs&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;for&lt;/span&gt; &lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-s1"&gt;x&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s1"&gt;z&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt; &lt;span class="pl-k"&gt;of&lt;/span&gt; &lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-c1"&gt;0.22&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.35&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;0.22&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.35&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-c1"&gt;0.22&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-c1"&gt;0.35&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;0.22&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-c1"&gt;0.35&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt;
    &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;leg&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;new&lt;/span&gt; &lt;span class="pl-c1"&gt;THREE&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;Mesh&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;legGeo&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-v"&gt;M&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;DARK&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-s1"&gt;leg&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;set&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;x&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.17&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s1"&gt;z&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;add&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;leg&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-s1"&gt;legs&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;push&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;leg&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-kos"&gt;}&lt;/span&gt;
  &lt;span class="pl-k"&gt;let&lt;/span&gt; &lt;span class="pl-s1"&gt;phase&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;random&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-c1"&gt;10&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;return&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt;
    &lt;span class="pl-c1"&gt;group&lt;/span&gt;: &lt;span class="pl-s1"&gt;g&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt;
    &lt;span class="pl-en"&gt;animate&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;dt&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s1"&gt;speed&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt;
      &lt;span class="pl-s1"&gt;phase&lt;/span&gt; &lt;span class="pl-c1"&gt;+=&lt;/span&gt; &lt;span class="pl-s1"&gt;dt&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;3&lt;/span&gt; &lt;span class="pl-c1"&gt;+&lt;/span&gt; &lt;span class="pl-s1"&gt;speed&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-c1"&gt;10&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
      &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;amp&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;min&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.6&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;0.1&lt;/span&gt; &lt;span class="pl-c1"&gt;+&lt;/span&gt; &lt;span class="pl-s1"&gt;speed&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-c1"&gt;0.6&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
      &lt;span class="pl-s1"&gt;legs&lt;/span&gt;&lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;0&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;rotation&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;x&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;sin&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;phase&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-s1"&gt;amp&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
      &lt;span class="pl-s1"&gt;legs&lt;/span&gt;&lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;3&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;rotation&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;x&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;sin&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;phase&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-s1"&gt;amp&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
      &lt;span class="pl-s1"&gt;legs&lt;/span&gt;&lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;1&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;rotation&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;x&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;sin&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;phase&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-s1"&gt;amp&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
      &lt;span class="pl-s1"&gt;legs&lt;/span&gt;&lt;span class="pl-kos"&gt;[&lt;/span&gt;&lt;span class="pl-c1"&gt;2&lt;/span&gt;&lt;span class="pl-kos"&gt;]&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;rotation&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;x&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-c1"&gt;-&lt;/span&gt;&lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;sin&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;phase&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-s1"&gt;amp&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
      &lt;span class="pl-s1"&gt;tail&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;rotation&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;z&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;sin&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;phase&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-c1"&gt;1.5&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-c1"&gt;0.4&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
      &lt;span class="pl-s1"&gt;body&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;position&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;y&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-c1"&gt;0.55&lt;/span&gt; &lt;span class="pl-c1"&gt;+&lt;/span&gt; &lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;abs&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-v"&gt;Math&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;sin&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;phase&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-c1"&gt;0.04&lt;/span&gt; &lt;span class="pl-c1"&gt;*&lt;/span&gt; &lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;0.3&lt;/span&gt; &lt;span class="pl-c1"&gt;+&lt;/span&gt; &lt;span class="pl-s1"&gt;speed&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
    &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt;
  &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
&lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;And did a &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-24-09-230Z"&gt;round of testing on it&lt;/a&gt; using Playwright, including &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-24-33-559Z"&gt;another screenshot&lt;/a&gt;.&lt;/p&gt;
&lt;div class="highlight highlight-source-js"&gt;&lt;pre&gt;  &lt;span class="pl-c"&gt;// walk near the dog&lt;/span&gt;
  &lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;page&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;evaluate&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="pl-kos"&gt;{&lt;/span&gt; &lt;span class="pl-k"&gt;const&lt;/span&gt; &lt;span class="pl-s1"&gt;d&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-smi"&gt;window&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;__rh&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;dog&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt; &lt;span class="pl-smi"&gt;window&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;__rh&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;teleport&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s1"&gt;d&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;x&lt;/span&gt; &lt;span class="pl-c1"&gt;+&lt;/span&gt; &lt;span class="pl-c1"&gt;6&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s1"&gt;d&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;z&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt; &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;page&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;waitForTimeout&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;2000&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;info&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;page&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;evaluate&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="pl-c1"&gt;JSON&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;stringify&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;{&lt;/span&gt; &lt;span class="pl-c1"&gt;dog&lt;/span&gt;: &lt;span class="pl-smi"&gt;window&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;__rh&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;dog&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;state&lt;/span&gt;: &lt;span class="pl-smi"&gt;window&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;__rh&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;state&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;player&lt;/span&gt;: &lt;span class="pl-smi"&gt;window&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;__rh&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;debug&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;player&lt;/span&gt; &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-smi"&gt;console&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;log&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s"&gt;'after approach:'&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s1"&gt;info&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;page&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;waitForTimeout&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-c1"&gt;3000&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-s1"&gt;info&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;page&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;evaluate&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt; &lt;span class="pl-c1"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="pl-c1"&gt;JSON&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;stringify&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;{&lt;/span&gt; &lt;span class="pl-c1"&gt;dog&lt;/span&gt;: &lt;span class="pl-smi"&gt;window&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;__rh&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;dog&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-c1"&gt;state&lt;/span&gt;: &lt;span class="pl-smi"&gt;window&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;__rh&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-c1"&gt;state&lt;/span&gt; &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-smi"&gt;console&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;log&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-s"&gt;'after chase:'&lt;/span&gt;&lt;span class="pl-kos"&gt;,&lt;/span&gt; &lt;span class="pl-s1"&gt;info&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;
  &lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;page&lt;/span&gt;&lt;span class="pl-kos"&gt;.&lt;/span&gt;&lt;span class="pl-en"&gt;screenshot&lt;/span&gt;&lt;span class="pl-kos"&gt;(&lt;/span&gt;&lt;span class="pl-kos"&gt;{&lt;/span&gt; &lt;span class="pl-c1"&gt;path&lt;/span&gt;: &lt;span class="pl-s1"&gt;__dirname&lt;/span&gt; &lt;span class="pl-c1"&gt;+&lt;/span&gt; &lt;span class="pl-s"&gt;'/shot-dog.png'&lt;/span&gt; &lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;span class="pl-kos"&gt;)&lt;/span&gt;&lt;span class="pl-kos"&gt;;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/shot-dog.png" alt="Screenshot of the game. Night 3, the raccoon is in the middle intersecting with the dog. A message reads Busted! The Guard shooed you off! The screen has a red tinge around the edges." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;It made a few more tweaks, then &lt;a href="https://simonw.github.io/raccoon-heist/transcript/page-001.html#msg-2026-08-05T15-29-13-245Z"&gt;declared itself done&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Raccoon Heist is built, tested, and pushed&lt;/strong&gt; — 7 commits on &lt;code&gt;claude/3d-raccoon-heist-game-50n293&lt;/code&gt;, all verified with automated Playwright playthroughs on desktop, portrait-phone, and landscape-phone viewports.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The game:&lt;/strong&gt; You're a masked raccoon sneaking out of your crew's dumpster hideout into a moonlit cul-de-sac. Grab coins, gems, rings, donuts, and fish (they visibly stack on your back and slow you down), then waddle home and bank them before dawn. Each night raises the goal and the heat: night 1 has one flashlight-sweeping guard, night 2 has two, night 3 adds a scent-tracking hound that ignores line-of-sight entirely. A police cruiser sweeps the street with headlights, homeowners randomly peek out of windows, a seagull periodically dive-bombs your unclaimed loot, pizza slices trigger a 6-second speed FRENZY, and a rare golden TV is worth 120 points if you can haul its weight home. Win ranks: TRASH PANDA ⭐ → CAT BURGLAR ⭐⭐ → MASTER OF DISGUISE ⭐⭐⭐, with your best run saved to localStorage.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How it's built:&lt;/strong&gt; Vendored Three.js (no CDN), procedural low-poly characters, and a procedural WebAudio soundtrack (sneaky walking-bass jazz plus all sound effects — zero audio files). All seven ground/wall textures and the title-screen hero art were generated with gpt-image-2 and committed as static assets, so the deployed game makes no API calls. Mobile gets a dynamic touch joystick (appears wherever you touch) plus a DASH button; desktop uses WASD/arrows and Space.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Testing caught two real bugs worth mentioning:&lt;/strong&gt; the canvas rendered at 2× size on any real phone (a &lt;code&gt;cssText&lt;/code&gt; assignment was wiping Three.js's inline sizing — desktop DPR-1 tests masked it completely), and the win screen's star-rating div inherited the title screen's full-screen &lt;code&gt;.stars&lt;/code&gt; CSS and silently swallowed every tap on the "next night" button. Both are fixed and regression-tested, along with pickup, banking, chase/caught, dawn, night progression, and lose/retry flows.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;notes.md&lt;/code&gt; has the full build log. The game is live on the branch — just serve the repo root and open &lt;code&gt;index.html&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="is-the-game-any-good-"&gt;Is the game any good?&lt;/h4&gt;
&lt;p&gt;It's an impressive starting point, but it's not a &lt;em&gt;good&lt;/em&gt; game.&lt;/p&gt;
&lt;p&gt;You move your raccoon around a yard collecting items - donuts, fish, gold coins, jewels - while avoiding guards with flashlights and, in later levels, a dog. You have a limited carrying capacity, and once that's full you need to drop stuff off at the dumpster. If you pick up a pizza slice you get a temporary speed boost.&lt;/p&gt;
&lt;p&gt;There are no team mechanics at all - there are two other static raccoons next to the dumpster but they're purely decoration.&lt;/p&gt;
&lt;p&gt;It gets slightly more challenging as the levels progress - the dog introduced in level 3 is the most interesting new mechanic - but it's very, very easy to beat. It's also pretty boring - each night has a fixed duration and you can collect all of the items and then have nothing else to do while waiting for the dawn.&lt;/p&gt;
&lt;p&gt;I was impressed by the implementation. It's fully 3D, there are trash cans, the flashlight illumination cones are fun, and it has a reasonably coherent visual style. It works on mobile. The music ("a procedural WebAudio soundtrack (sneaky walking-bass jazz plus all sound effects — zero audio files)" according to Claude) is simple but feels about right.&lt;/p&gt;
&lt;p&gt;As a finished game project, it's mediocre. As a starting point from a single prompt I think it's very impressive.&lt;/p&gt;
&lt;p&gt;I've vibe coded up quite a few games now. They've all been deeply disappointing from a gameplay perspective - it turns out designing games that are &lt;em&gt;fun&lt;/em&gt; remains a uniquely human trait, and one which requires significantly more skill and experience than either Claude or I can bring to bear.&lt;/p&gt;
&lt;p&gt;That said, I thoroughly recommend tinkering with game development projects as a way to explore the capabilities of agents. It's a fun, low-risk way to try out new things. If you stick at it long enough you might even produce something that's worth playing!&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Update 7th August 2026&lt;/strong&gt;: I posed the same prompt to OpenAI Codex Desktop running GPT-5.6 Sol Ultra and got a &lt;a href="https://simonwillison.net/2026/Aug/7/moonlight-mayhem/"&gt;significantly better result&lt;/a&gt; - GPT-5.6 Sol picked up on the importance of the squad of raccoons going on a heist, and built a game where you must rescue your two crewmates in a museum and then stack on top of them to steal the Golden Sardine.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="game-design"/><category term="ai"/><category term="prompt-engineering"/><category term="generative-ai"/><category term="llms"/><category term="anthropic"/><category term="claude"/><category term="text-to-image"/><category term="vibe-coding"/><category term="coding-agents"/><category term="claude-mythos-fable"/></entry><entry><title>New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging</title><link href="https://simonwillison.net/2026/Aug/4/new-release-of-llm/" rel="alternate"/><published>2026-08-04T23:58:24+00:00</published><updated>2026-08-04T23:58:24+00:00</updated><id>https://simonwillison.net/2026/Aug/4/new-release-of-llm/</id><summary type="html">&lt;p&gt;I released &lt;a href="https://llm.datasette.io/en/stable/changelog.html#v0-32"&gt;LLM 0.32&lt;/a&gt; this morning, the most significant new version of LLM since the initial launch of the project. The new version includes support for visible reasoning traces, server-side provider tools, redesigned content-addressable SQLite logs, new models, and new features enabled by the OpenAI Responses API. I also released a new version of the &lt;a href="https://github.com/simonw/llm-anthropic"&gt;llm-anthropic plugin&lt;/a&gt; with substantial updates of its own.&lt;/p&gt;
&lt;h4 id="headline-features-for-llm-cli-users"&gt;Headline features for LLM CLI users&lt;/h4&gt;
&lt;p&gt;Running LLM against reasoning models now &lt;strong&gt;displays their reasoning traces&lt;/strong&gt; to standard error, so you can see what they are "thinking" without that information being included in the standard output that you might pipe to another tool. Add &lt;code&gt;-R/--hide-reasoning&lt;/code&gt; to turn this off.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/best-pelicans.gif" alt="Running llm &amp;quot;think about the best thing about pelicans&amp;quot; in the macOS terminal window - grey text outputs saying Exploring pelican qualities, then after a paragraph of that a white paragraph of text comes out saying: The best thing about pelicans is their wonderfully oversized, practical design: that enormous bill and pouch look comical, but they make pelicans remarkably skilled fishers. Even better, many species cooperate—working together to herd fish before scooping them up. They’re a great mix of goofy, graceful, and surprisingly clever." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;LLM includes support out-of-the-box for the &lt;strong&gt;GPT-5.6 model family&lt;/strong&gt;, and the new default model used with &lt;code&gt;llm "prompt"&lt;/code&gt; is now the inexpensive but capable &lt;strong&gt;GPT-5.6 Luna&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;LLM calls can now use &lt;strong&gt;server-side tools&lt;/strong&gt; from various providers. OpenAI provide &lt;a href="https://llm.datasette.io/en/stable/openai-models.html#code-interpreter"&gt;a code execution environment&lt;/a&gt; as a server-side tool; LLM can now run prompts that benefit from that like so:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm --tool CodeInterpreter &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;Show current python and SQLite versions&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;OpenAI also gets a &lt;a href="https://llm.datasette.io/en/stable/openai-models.html#web-search"&gt;WebSearch&lt;/a&gt; tool.&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://github.com/simonw/llm-anthropic"&gt;llm-anthropic&lt;/a&gt; plugin adds &lt;a href="https://github.com/simonw/llm-anthropic/blob/0.26/README.md#web-search"&gt;WebSearch&lt;/a&gt;, &lt;a href="https://github.com/simonw/llm-anthropic/blob/0.26/README.md#web-fetch"&gt;WebFetch&lt;/a&gt;, &lt;a href="https://github.com/simonw/llm-anthropic/blob/0.26/README.md#code-execution"&gt;CodeExecution&lt;/a&gt;, and &lt;a href="https://github.com/simonw/llm-anthropic/blob/0.26/README.md#mcp-connector"&gt;AnthropicMCP&lt;/a&gt;, which looks like this:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm -m claude-sonnet-5 -T &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;AnthropicMCP("https://datasette.simonwillison.net/-/mcp")&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt; \
  &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;how many rows in the blog_blogmark table?&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;That causes Anthropic to execute MCP calls against my new &lt;a href="https://simonwillison.net/2026/Jul/31/stateless-mcp/#datasette-mcp"&gt;datasette-mcp&lt;/a&gt; plugin as part of a single request/response interaction with their API.&lt;/p&gt;
&lt;p&gt;The new &lt;strong&gt;llm openai endpoint&lt;/strong&gt; command provides a tool for &lt;a href="https://llm.datasette.io/en/stable/other-models.html#run-against-an-endpoint-without-configuring-it"&gt;executing prompts against &lt;em&gt;any&lt;/em&gt; OpenAI compatible endpoint&lt;/a&gt; as a one-liner. These aren't logged, which makes this a handy tool for running one-off prompts against anything that speaks the lingua franca of the LLM API world.&lt;/p&gt;
&lt;p&gt;Here's how I use that to run prompts against Gemma 4 12B running in my localhost &lt;a href="https://lmstudio.ai"&gt;LM Studio&lt;/a&gt; API, via &lt;code&gt;uvx&lt;/code&gt; (no LLM installation required) and mixing in the &lt;a href="https://github.com/simonw/llm-tools-quickjs"&gt;llm-tools-quickjs&lt;/a&gt; tool plugin for good measure:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;uvx --with llm-tools-quickjs \
  llm openai endpoint http://localhost:1234/v1 -m google/gemma-4-12b \
  -T QuickJS &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;Use QuickJS to multiply 3434 * 2434&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt; --td&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/openai-endpoint-gemma.webp" alt="Output reads Tool call: QuickJS_execute_javascript({'javascript': '3434 * 2434'})  8358356 The result of 3434 * 2434 is 8,358,356." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;h4 id="new-features-in-the-python-api"&gt;New features in the Python API&lt;/h4&gt;
&lt;p&gt;LLM's Python API previously required you to create a conversation and then send messages to it one at a time. This was an abstraction over the true nature of LLMs, where each request carries a complete history of the messages that came before it. That abstraction started to get in the way for some more advanced cases, so the new release introduces a &lt;code&gt;model.prompt(messages=[])&lt;/code&gt; parameter that can be used like this:&lt;/p&gt;
&lt;pre&gt;&lt;span class="pl-k"&gt;import&lt;/span&gt; &lt;span class="pl-s1"&gt;llm&lt;/span&gt;
&lt;span class="pl-k"&gt;from&lt;/span&gt; &lt;span class="pl-s1"&gt;llm&lt;/span&gt; &lt;span class="pl-k"&gt;import&lt;/span&gt; &lt;span class="pl-s1"&gt;user&lt;/span&gt;, &lt;span class="pl-s1"&gt;assistant&lt;/span&gt;, &lt;span class="pl-s1"&gt;system&lt;/span&gt;

&lt;span class="pl-s1"&gt;model&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;llm&lt;/span&gt;.&lt;span class="pl-c1"&gt;get_model&lt;/span&gt;(&lt;span class="pl-s"&gt;"gpt-5.6-luna"&lt;/span&gt;)

&lt;span class="pl-s1"&gt;response&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;model&lt;/span&gt;.&lt;span class="pl-c1"&gt;prompt&lt;/span&gt;(&lt;span class="pl-s1"&gt;messages&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;[
    &lt;span class="pl-en"&gt;system&lt;/span&gt;(&lt;span class="pl-s"&gt;"You are a helpful pirate."&lt;/span&gt;),
    &lt;span class="pl-en"&gt;user&lt;/span&gt;(&lt;span class="pl-s"&gt;"What is the capital of France?"&lt;/span&gt;),
    &lt;span class="pl-en"&gt;assistant&lt;/span&gt;(&lt;span class="pl-s"&gt;"Paris, matey."&lt;/span&gt;),
    &lt;span class="pl-en"&gt;user&lt;/span&gt;(&lt;span class="pl-s"&gt;"And Germany?"&lt;/span&gt;),
])
&lt;span class="pl-en"&gt;print&lt;/span&gt;(&lt;span class="pl-s1"&gt;response&lt;/span&gt;.&lt;span class="pl-c1"&gt;text&lt;/span&gt;())&lt;/pre&gt;
&lt;p&gt;LLM previously returned an iterable sequence of strings from each prompt. This worked great when models returned a string response, but failed to predict the weird shape that models would evolve towards. Today many models return a mix of reasoning text, output strings, tool calls, and even image attachments. With LLM 0.32 you can &lt;a href="https://llm.datasette.io/en/stable/python-api.html#structured-messages-and-streaming-events"&gt;do this instead&lt;/a&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;span class="pl-k"&gt;for&lt;/span&gt; &lt;span class="pl-s1"&gt;event&lt;/span&gt; &lt;span class="pl-c1"&gt;in&lt;/span&gt; &lt;span class="pl-s1"&gt;model&lt;/span&gt;.&lt;span class="pl-c1"&gt;prompt&lt;/span&gt;(&lt;span class="pl-s"&gt;"Explain cats"&lt;/span&gt;).&lt;span class="pl-c1"&gt;stream_events&lt;/span&gt;():
    &lt;span class="pl-k"&gt;if&lt;/span&gt; &lt;span class="pl-s1"&gt;event&lt;/span&gt;.&lt;span class="pl-c1"&gt;type&lt;/span&gt; &lt;span class="pl-c1"&gt;==&lt;/span&gt; &lt;span class="pl-s"&gt;"reasoning"&lt;/span&gt;:
        &lt;span class="pl-en"&gt;print&lt;/span&gt;(&lt;span class="pl-s"&gt;f"[thinking] &lt;span class="pl-s1"&gt;&lt;span class="pl-kos"&gt;{&lt;/span&gt;&lt;span class="pl-s1"&gt;event&lt;/span&gt;.&lt;span class="pl-c1"&gt;chunk&lt;/span&gt;&lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;/span&gt;"&lt;/span&gt;, &lt;span class="pl-s1"&gt;end&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;""&lt;/span&gt;, &lt;span class="pl-s1"&gt;flush&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;True&lt;/span&gt;)
    &lt;span class="pl-k"&gt;elif&lt;/span&gt; &lt;span class="pl-s1"&gt;event&lt;/span&gt;.&lt;span class="pl-c1"&gt;type&lt;/span&gt; &lt;span class="pl-c1"&gt;==&lt;/span&gt; &lt;span class="pl-s"&gt;"text"&lt;/span&gt;:
        &lt;span class="pl-en"&gt;print&lt;/span&gt;(&lt;span class="pl-s1"&gt;event&lt;/span&gt;.&lt;span class="pl-c1"&gt;chunk&lt;/span&gt;, &lt;span class="pl-s1"&gt;end&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;""&lt;/span&gt;, &lt;span class="pl-s1"&gt;flush&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;True&lt;/span&gt;)
    &lt;span class="pl-k"&gt;else&lt;/span&gt;:
        &lt;span class="pl-en"&gt;print&lt;/span&gt;(&lt;span class="pl-s"&gt;f"Other event: &lt;span class="pl-s1"&gt;&lt;span class="pl-kos"&gt;{&lt;/span&gt;&lt;span class="pl-s1"&gt;event&lt;/span&gt;&lt;span class="pl-kos"&gt;}&lt;/span&gt;&lt;/span&gt;"&lt;/span&gt;)&lt;/pre&gt;
&lt;p&gt;Combine these features and we can &lt;em&gt;finally&lt;/em&gt; provide a robust implementation of the semi-standard OpenAI chat completions API, which I've now released as the &lt;a href="https://github.com/simonw/llm-chat-completions-server"&gt;llm-chat-completions-server&lt;/a&gt; plugin:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm install llm-chat-completions-server
llm chat-completions-server --port 9000
&lt;span class="pl-c"&gt;&lt;span class="pl-c"&gt;#&lt;/span&gt; Server is now running on http://127.0.0.1:9000/v1&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Now you can run prompts against LLM via that server, using the new &lt;code&gt;llm openai endpoint&lt;/code&gt; command!&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm openai endpoint http://127.0.0.1:9000/v1 &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;hello&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt; -m gpt-5.4-mini&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The bigger challenge with that kind of API concerns logging. If we're going to support the pattern where the message sequence is appended to on every request, ideally we can avoid logging all of that duplicate JSON for every turn.&lt;/p&gt;
&lt;p&gt;The solution is the new &lt;a href="https://llm.datasette.io/en/stable/logging.html#the-message-store"&gt;content-addressable message store&lt;/a&gt;, modeled after Git. You can see the new schema for that &lt;a href="https://llm.datasette.io/en/stable/logging.html#sql-schema"&gt;in the documentation&lt;/a&gt;, but the &lt;code&gt;llm logs&lt;/code&gt; and &lt;code&gt;llm logs --json&lt;/code&gt; commands have both been upgraded to convert that format back into something that's easy to consume.&lt;/p&gt;
&lt;h4 id="and-the-rest"&gt;And the rest&lt;/h4&gt;
&lt;p&gt;There is a whole lot more in this release. The &lt;a href="https://llm.datasette.io/en/stable/changelog.html#v0-32"&gt;0.32 release notes&lt;/a&gt; are pretty comprehensive, and the notes for &lt;a href="https://llm.datasette.io/en/stable/changelog.html#rc2-2026-07-30"&gt;0.32rc2&lt;/a&gt;, &lt;a href="https://llm.datasette.io/en/stable/changelog.html#rc1-2026-07-30"&gt;0.32rc&lt;/a&gt;, &lt;a href="https://llm.datasette.io/en/stable/changelog.html#a3-2026-06-09"&gt;0.32a3&lt;/a&gt;, &lt;a href="https://llm.datasette.io/en/stable/changelog.html#a2-2026-05-12"&gt;0.32a2&lt;/a&gt;, and &lt;a href="https://llm.datasette.io/en/stable/changelog.html#a0-2026-04-28"&gt;0.32a0&lt;/a&gt; should fill in any gaps.&lt;/p&gt;
&lt;p&gt;Existing LLM plugins should all continue to work, but plugins that provide extra models will need to be upgraded to 0.32 in order to participate fully in the new streaming events system. There's a guide to implementing plugins with &lt;a href="https://llm.datasette.io/en/stable/plugins/advanced-model-plugins.html#structured-messages-and-streaming-events"&gt;Structured messages and streaming events&lt;/a&gt; in the documentation.&lt;/p&gt;
&lt;p&gt;I've updated some of my own plugins:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/simonw/llm-anthropic/releases/tag/0.26"&gt;llm-anthropic 0.26&lt;/a&gt; adds support for the Claude 5 family of models, plus &lt;code&gt;WebSearch&lt;/code&gt;, &lt;code&gt;WebFetch&lt;/code&gt;, &lt;code&gt;CodeExecution&lt;/code&gt;, and &lt;code&gt;AnthropicMCP&lt;/code&gt; server-side tools.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/simonw/llm-gemini"&gt;llm-gemini&lt;/a&gt; and &lt;a href="https://github.com/simonw/llm-openrouter"&gt;llm-openrouter&lt;/a&gt; and &lt;a href="https://github.com/simonw/llm-mistral"&gt;llm-mistral&lt;/a&gt; are nearly there, releases coming soon.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="i-guess-llm-is-an-agent-framework-now"&gt;I guess LLM is an agent framework now&lt;/h4&gt;
&lt;p&gt;Quite a few of the lower-level tools changes in this release were driven by the needs of &lt;a href="https://agent.datasette.io/"&gt;Datasette Agent&lt;/a&gt;. When I started work on LLM, the term "agent" had such a vague definition that I refused to use it. In &lt;a href="https://simonwillison.net/2025/Sep/18/agents/"&gt;September 2025&lt;/a&gt; I came around to the idea that "&lt;strong&gt;An LLM agent runs tools in a loop to achieve a goal&lt;/strong&gt;" is well established enough now that I could stop avoiding the term entirely.&lt;/p&gt;
&lt;p&gt;Tool chains can now &lt;a href="https://llm.datasette.io/en/stable/python-api.html#python-api-tools-pause"&gt;pause for human approval&lt;/a&gt; and &lt;a href="https://llm.datasette.io/en/stable/python-api.html#python-api-tools-resume"&gt;resume from a stored message history&lt;/a&gt; - both needed by Datasette Agent.&lt;/p&gt;
&lt;p&gt;Looking at LLM today it's beginning to look very agent-shaped to me. There's something neat about having a CLI utility that can mix and match different tools from different sources with different models all as a one-liner, and that includes a Python library powerful enough to build systems like &lt;a href="https://agent.datasette.io/"&gt;Datasette Agent&lt;/a&gt; and &lt;a href="https://github.com/simonw/llm-coding-agent"&gt;llm-coding-agent&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Maybe the next version of LLM will bake the concept of an "agent" into the core library. I'm still trying to figure out what that would look like.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="projects"/><category term="releases"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="llm"/><category term="anthropic"/><category term="llm-tool-use"/><category term="llm-reasoning"/><category term="model-context-protocol"/></entry><entry><title>Stateless MCP has recaptured my interest (and inspired mcp-explorer and datasette-mcp)</title><link href="https://simonwillison.net/2026/Jul/31/stateless-mcp/" rel="alternate"/><published>2026-07-31T23:13:22+00:00</published><updated>2026-07-31T23:13:22+00:00</updated><id>https://simonwillison.net/2026/Jul/31/stateless-mcp/</id><summary type="html">&lt;p&gt;Tuesday was &lt;a href="https://x.com/ade_oshineye/status/2082129440943866149"&gt;Stateless MCP day&lt;/a&gt; - the rollout of MCP 2.0, or &lt;a href="https://blog.modelcontextprotocol.io/posts/2026-07-28/"&gt;the 2026-07-28 Model Context Protocol specification&lt;/a&gt; to use the more formal but less memorable name. This is the most significant change to the MCP spec since it first launched, and has also served to reignite my personal interest in the protocol.&lt;/p&gt;
&lt;p&gt;For background: MCP is the Model Context Protocol, which describes a standard way to expose new tools to LLM-powered agent frameworks. It was introduced by Anthropic back &lt;a href="https://www.anthropic.com/news/model-context-protocol"&gt;in November 2024&lt;/a&gt;, had a &lt;em&gt;huge&lt;/em&gt; spike of interest through much of 2025, and then became somewhat eclipsed by &lt;a href="https://simonwillison.net/2025/Oct/16/claude-skills/"&gt;Skills&lt;/a&gt; (another Anthropic invention) when it became apparent that an agent harness with access to a terminal and &lt;code&gt;curl&lt;/code&gt; could do most of what MCP did in a more flexible way. I wrote about that &lt;a href="https://simonwillison.net/2025/Dec/31/the-year-in-llms/#the-only-year-of-mcp"&gt;in my review of 2025&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I'm coming back around to MCP now. Giving an agent a shell environment with the ability to access the internet is &lt;a href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/"&gt;fraught with risk&lt;/a&gt;, and requires a strong model that is capable of effectively driving such an environment. MCP tools are easier to audit and control, and simple enough that smaller models that run on a laptop can still drive them reasonably well.&lt;/p&gt;
&lt;p&gt;The new stateless MCP specification also greatly decreases the complexity of implementing both clients and servers for the protocol. I built three of those this week!&lt;/p&gt;
&lt;h4 id="what-s-easier-with-stateless-mcp"&gt;What's easier with stateless MCP&lt;/h4&gt;
&lt;p&gt;The best demonstration of the difference between stateful and stateless MCP is in this &lt;a href="https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/"&gt;May 21st blog post&lt;/a&gt; that introduced the RC for the new specification. It included a clear before-and-after example.&lt;/p&gt;
&lt;p&gt;The older stateful MCP (I'm going to call it "legacy MCP") required two HTTP requests - the first to initialize a session and obtain a &lt;code&gt;Mcp-Session-Id&lt;/code&gt;, and the second to actually call the tool:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;POST /mcp HTTP/1.1
Content-Type: application/json

{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "initialize",
  "params": {
    "protocolVersion": "2025-11-25",
    "capabilities": {
    },
    "clientInfo": {
      "name": "my-app",
      "version": "1.0"
    }
  }
}

POST /mcp HTTP/1.1
Mcp-Session-Id: 1868a90c-3a3f-4f5b
Content-Type: application/json

{
  "jsonrpc": "2.0",
  "id": 2,
  "method": "tools/call",
  "params": {
    "name": "search",
    "arguments": {
      "q": "otters"
    }
  }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;The new stateless way uses a single HTTP request which looks like this:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;POST /mcp HTTP/1.1
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: search
Content-Type: application/json

{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "search",
    "arguments": {
      "q": "otters"
    },
    "_meta": {
      "io.modelcontextprotocol/clientInfo": {
        "name": "my-app",
        "version": "1.0"
      }
    }
  }
}
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This is so much cleaner from both a client- and server-side implementation perspective. It's also a better fit for building scalable web applications, since now you don't need to maintain server-side state to keep track of those session IDs, or worry about routing the same session to the same backend machine.&lt;/p&gt;
&lt;h4 id="mcp-explorer"&gt;mcp-explorer&lt;/h4&gt;
&lt;p&gt;I couldn't find a great CLI tool for interactively probing an MCP server, so I had Codex help build my own.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/simonw/mcp-explorer"&gt;mcp-explorer&lt;/a&gt;&lt;/strong&gt; is the result. It's a stateless Python CLI tool, so you don't even need to install it to try it out - it works with &lt;a href="https://docs.astral.sh/uv/guides/tools/#running-tools"&gt;uvx&lt;/a&gt; like this:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;uvx mcp-explorer list https://agentic-mermaid.dev/mcp&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This queries Ade Oshineye's &lt;a href="https://agentic-mermaid.dev/"&gt;agentic-mermaid.dev&lt;/a&gt; demo MCP. The above command returns the following list of tools:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;execute(code: string, timeoutMs?: integer) - Execute Mermaid SDK code
  Run JavaScript in an isolated sandbox; return a value.

describe_sdk(family: string, detail?: string) - Describe Mermaid SDK operations
  Return version-matched mutation operations for one diagram family.

render_svg(source: string, options?: object) - Render Mermaid as SVG
  Render a Mermaid source string to themeable SVG. Returns { ok, svg }.

render_ascii(source: string, useAscii?: boolean, targetWidth?: integer, options?: object) - Render Mermaid as text
  Render a Mermaid source string to text. Returns { ok, text }.

render_png(source: string, scale?: number, background?: string, fitTo?: object, options?: object) - Render Mermaid as PNG
  Rasterize a Mermaid source string to PNG. Returns { ok, png_base64 }.
...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Then to inspect a tool:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;uvx mcp-explorer inspect render_svg&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This outputs a whole bunch of information, including the JSON schema of the inputs and outputs.&lt;/p&gt;
&lt;p&gt;To call that tool and pass arguments to it:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;uvx mcp-explorer call \
  https://agentic-mermaid.dev/mcp \
  render_svg \
  -a &lt;span class="pl-c1"&gt;source&lt;/span&gt; &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;graph TD; A--&amp;gt;B&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt; \
  -a options &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;{"padding":24}&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Which returns:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;{"ok":true,"svg":"&amp;lt;svg xmlns=\"http://www.w3.org/2000/svg\" width=...
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;To get just the raw SVG try adding &lt;code&gt;| jq .svg -r&lt;/code&gt; to that command. I got back &lt;a href="https://gist.github.com/simonw/b07c62f0ce103be6932477659d5dd1ac"&gt;this image&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/mermaid-example.svg" alt="SVG of as A box on top of a B box with an arrow from A to B" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;There are a &lt;a href="https://github.com/simonw/mcp-explorer/blob/main/README.md"&gt;few more commands&lt;/a&gt; in the README, but you get the general idea. I find building CLI tools like this to be a really productive way to get familiar with a specification, even if an agent writes most of the actual code.&lt;/p&gt;
&lt;h4 id="datasette-mcp"&gt;datasette-mcp&lt;/h4&gt;
&lt;p&gt;The second project is &lt;strong&gt;&lt;a href="https://github.com/datasette/datasette-mcp"&gt;datasette-mcp&lt;/a&gt;&lt;/strong&gt;, a Datasette plugin which adds a &lt;code&gt;/-/mcp&lt;/code&gt; endpoint to any Datasette instance.&lt;/p&gt;
&lt;p&gt;This is probably the fourth time I've tried building this plugin, but thanks to the new stateless MCP specification I finally have a version that feels good to release.&lt;/p&gt;
&lt;p&gt;It provides just three tools: &lt;code&gt;list_databases()&lt;/code&gt;, &lt;code&gt;get_database_schema(database_name)&lt;/code&gt;, and &lt;code&gt;execute_sql(database_name, sql)&lt;/code&gt;. They do exactly what you would expect them to do - though &lt;code&gt;execute_sql()&lt;/code&gt; is read-only for the moment.&lt;/p&gt;
&lt;p&gt;Wire these into an agent, or a chat tool like ChatGPT or Claude, and they'll gain the ability to run SQL queries against your hosted Datasette instance.&lt;/p&gt;
&lt;p&gt;So far I'm running it on the Datasette mirror of my blog, at &lt;a href="datasette.simonwillison.net/-/mcp"&gt;datasette.simonwillison.net/-/mcp&lt;/a&gt;. It took a bit of fiddling to figure out how to attach that to ChatGPT and Claude, but I got there in the end. Here's &lt;a href="https://til.simonwillison.net/llms/mcp-in-claude-and-chatgpt"&gt;a new TIL&lt;/a&gt; showing exactly how to do that.&lt;/p&gt;
&lt;p&gt;Here's &lt;a href="https://claude.ai/share/de1ad9bf-f7c2-4fb9-a9a0-2a1ae39995db"&gt;a shared Claude session&lt;/a&gt; where I asked it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;list tables in simonwillison.net&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And then:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;what has Simon said recently about MCP?&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It ran 7 separate SQL queries to figure out the answer.&lt;/p&gt;
&lt;h4 id="llm-mcp-client"&gt;llm-mcp-client&lt;/h4&gt;
&lt;p&gt;My &lt;a href="https://llm.datasette.io/"&gt;LLM tool&lt;/a&gt; is long overdue for an official MCP integration. The new alpha &lt;a href="https://github.com/simonw/llm-mcp-client"&gt;llm-mcp-client&lt;/a&gt; plugin is my attempt at exactly that:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm install llm-mcp-client
llm -T &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;MCP("https://datasette.simonwillison.net/-/mcp")&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt; &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;count the notes&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Here's the output (including reasoning trace, I'm using &lt;a href="https://simonwillison.net/2026/Jul/30/llm-rc2/"&gt;LLM 0.32rc2&lt;/a&gt;):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Considering note count&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;I see the question "count the notes" is probably asking me to tally up blog notes. It could also mean published notes or drafts, so there's some ambiguity there. I'll need to figure out the total number of notes, likely by querying the count for both published notes and drafts to get a clear answer. Let's execute that count!&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;There are &lt;strong&gt;151 notes&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And &lt;a href="https://gist.github.com/simonw/4e8f558766150658ce35eab4f0fc3e04"&gt;the output of llm logs&lt;/a&gt; for that prompt.&lt;/p&gt;
&lt;p&gt;Once this is fully baked, I'm considering bringing it directly into LLM core. I'm excited to experiment with MCP in &lt;a href="https://agent.datasette.io/"&gt;Datasette Agent&lt;/a&gt; and &lt;a href="https://github.com/simonw/llm-coding-agent"&gt;llm-coding-agent&lt;/a&gt; as well.&lt;/p&gt;
&lt;h4 id="mcp-is-a-safer-way-to-build-with-agents"&gt;MCP is a safer way to build with agents&lt;/h4&gt;
&lt;p&gt;A few months after MCP was first released, I wrote &lt;a href="https://simonwillison.net/2025/Apr/9/mcp-prompt-injection/"&gt;Model Context Protocol has prompt injection security problems&lt;/a&gt;, where I noted that the pattern of having end users mix and match tools pushed responsibility for avoiding data exfiltration attacks out to the users themselves. I hadn't coined &lt;a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/"&gt;the Lethal Trifecta&lt;/a&gt; yet, but that was absolutely what I had in mind.&lt;/p&gt;
&lt;p&gt;Then general agents with arbitrary shell and &lt;code&gt;curl&lt;/code&gt; access came along, and that's so much harder to keep secure!&lt;/p&gt;
&lt;p&gt;Something I've come to appreciate about MCP is that it's much easier to reason about agent capabilities and what might go wrong than with arbitrary command execution in an open network environment - the default for most of today's general and coding agent tools.&lt;/p&gt;
&lt;p&gt;I plan to lean into MCP a whole lot more when I'm building sensitive applications on top of LLMs.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="projects"/><category term="ai"/><category term="datasette"/><category term="mermaid"/><category term="generative-ai"/><category term="llms"/><category term="llm"/><category term="anthropic"/><category term="model-context-protocol"/></entry><entry><title>OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened</title><link href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/" rel="alternate"/><published>2026-07-22T23:51:33+00:00</published><updated>2026-07-22T23:51:33+00:00</updated><id>https://simonwillison.net/2026/Jul/22/openai-cyberattack/</id><summary type="html">&lt;p&gt;This story is wild. The short version: OpenAI were running a cybersecurity test against an unreleased model, with the model's guardrail features turned off. Rather than solve the test, the model broke its way out of OpenAI's sandbox, then found exploits to break &lt;em&gt;in&lt;/em&gt; to Hugging Face, all so it could cheat on the test by stealing the answers.&lt;/p&gt;
&lt;p&gt;Along the way it helped make the strongest case yet for how the imbalance of model availability is hurting our ability to secure our software.&lt;/p&gt;
&lt;h4 id="here-s-what-happened"&gt;Here's what happened&lt;/h4&gt;
&lt;p&gt;We currently have three documents to help us understand what happened here.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2605.11086"&gt;ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?&lt;/a&gt; is a paper published on 11th May 2026 describing ExploitGym, a new eval suite for LLM-powered agent systems.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://huggingface.co/blog/security-incident-july-2026"&gt;Security incident disclosure — July 2026&lt;/a&gt; by Hugging Face on 16th July 2026 describes how they detected an attack from an "agentic security-research harness - used LLM still not known" that breached some of their systems.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"&gt;OpenAI and Hugging Face partner to address security incident during model evaluation&lt;/a&gt; from OpenAI on 21st July 2026 confesses that it was &lt;em&gt;their&lt;/em&gt; agent harness that did this, and that they're working with Hugging Face to clean up the mess.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Update 5th August 2026&lt;/strong&gt;: Hugging Face published &lt;a href="https://huggingface.co/blog/agent-intrusion-technical-timeline"&gt;a great deal more information&lt;/a&gt; about the attack on July 27th&lt;/em&gt;.&lt;/p&gt;
&lt;h4 id="exploitgym"&gt;ExploitGym&lt;/h4&gt;
&lt;p&gt;I hadn't seen the &lt;a href="https://arxiv.org/abs/2605.11086"&gt;ExploitGym paper&lt;/a&gt; before and it's a really interesting one. Authors from UC Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State designed a new benchmark for evaluating models on their ability to turn a reported vulnerability into a concrete exploit. OpenAI, Anthropic, and Google provided feedback and helped run the benchmark against their models.&lt;/p&gt;
&lt;p&gt;The benchmark "comprises 898 instances derived from real-world vulnerabilities that affected popular software projects" - including the Linux kernel and V8 JavaScript engine. The ExploitGym benchmark is &lt;a href="https://github.com/sunblaze-ucb/exploitgym"&gt;available on GitHub&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Here's the paragraph that best represents their benchmark results:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Among all configurations, Claude Mythos Preview and GPT-5.5 achieve the highest success counts (157 and 120 successes, respectively), demonstrating that current frontier agents can exploit a substantial subset of real-world vulnerabilities under controlled conditions. GPT-5.4 also solves a notable 54 tasks, placing it in an intermediate tier. The remaining model–agent pairings solve fewer than 15 tasks each, underscoring that end-to-end exploitation remains challenging and sharply differentiates today’s frontier systems. Notably, Claude Opus 4.7 achieves fewer successes than Claude Opus 4.6 despite being a newer checkpoint, and does so at substantially lower cost on the full set. Trace inspection reveals that Claude Opus 4.7 and Gemini 3.1 Pro frequently conclude early after judging the target vulnerability non-exploitable.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The paper also describes the approach they took to preventing the agents from cheating by going outside the parameters of the test. This becomes relevant in a moment!&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Outbound connections are restricted to a curated allowlist that permits routine package installation (Ubuntu apt repositories and PyPI) and fetching the toolchains required for building V8. All other external endpoints are blocked.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The paper concludes with this (emphasis mine):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Our results show that &lt;strong&gt;autonomous exploit development by frontier AI agents is no longer a hypothetical capability&lt;/strong&gt;. While current agents are not yet reliable across all targets, they already &lt;strong&gt;exploit a non-trivial fraction of real-world vulnerabilities&lt;/strong&gt;, including complex targets such as kernel components. This rapid emergence is itself a central finding, showing that capabilities that would have seemed implausible are now present in deployed frontier models.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;An important detail here: this paper isn't about discovering vulnerabilities; it's about being able to take those vulnerabilities and turn them into working exploits.&lt;/p&gt;
&lt;p&gt;When Anthropic first restricted access to Mythos &lt;a href="https://simonwillison.net/2026/Apr/7/project-glasswing/"&gt;back in April&lt;/a&gt; they talked about this capability as well. A model that can act on vulnerabilities is a lot more dangerous than one that can just discover them.&lt;/p&gt;
&lt;p&gt;One of the ways Fable differs from Mythos is that it's more likely to refuse to weaponize vulnerabilities in this way. I get the impression the US government did not understand that distinction when they banned Fable &lt;a href="https://simonwillison.net/2026/Jun/16/fable-5-export-controls/"&gt;last month&lt;/a&gt;.&lt;/p&gt;
&lt;h4 id="the-hugging-face-incident"&gt;The Hugging Face incident&lt;/h4&gt;
&lt;p&gt;The first hint we got of the attack was in &lt;a href="https://huggingface.co/blog/security-incident-july-2026"&gt;this blog post by Hugging Face&lt;/a&gt; on 16th July 2026:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I hope they release more details about the code that pulled this off. I'm assuming this means packages using the &lt;a href="https://github.com/huggingface/datasets"&gt;datasets library&lt;/a&gt;, a Hugging Face project for bundling up and sharing datasets on their platform. That library used to execute arbitrary code but has been steadily locked down over time, with the &lt;a href="https://github.com/huggingface/datasets/releases/tag/4.0.0"&gt;4.0.0 release&lt;/a&gt; in July 2025 removing the &lt;code&gt;trust_remote_code=True&lt;/code&gt; flag entirely.&lt;/p&gt;
&lt;p&gt;Assuming the attack used that library it must have either abused pickle serialization in some way, found some other non-obvious code execution path, or (most likely) specified &lt;code&gt;datasets&amp;lt;4.0.0&lt;/code&gt; as the dependency.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This was a sophisticated attack!&lt;/p&gt;
&lt;p&gt;Then Hugging Face hit a wall: they tried to use "frontier models behind commercial APIs" - I'm guessing from Anthropic and OpenAI - to help analyze the attack, and were blocked:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;They switched to their own self-hosted instance of MIT licensed GLM-5.2 and it helped them figure out what was going on.&lt;/p&gt;
&lt;p&gt;This indicated a fundamental asymmetry between the defending team and the (so-far unknown) attacker:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;As a useful indicator of how seriously they took the attack:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[...] Finally, we have also reported this incident to law enforcement agencies.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So who was responsible for this "autonomous agent framework"? It turned out to be OpenAI themselves.&lt;/p&gt;
&lt;h4 id="the-openai-confession"&gt;The OpenAI confession&lt;/h4&gt;
&lt;p&gt;Five days later, &lt;a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/"&gt;on July 21st&lt;/a&gt;, OpenAI revealed the culprit. They had been running the ExploitGym benchmark against a new, as-yet undisclosed model, and that model had been operating &lt;em&gt;way&lt;/em&gt; outside its intended parameters (emphasis mine):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;After investigating, we now know &lt;strong&gt;that this particular incident was driven by a combination of OpenAI models&lt;/strong&gt; — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a &lt;a href="https://arxiv.org/abs/2605.11086"&gt;benchmark⁠&lt;/a&gt; [ExploitGym] of cyber capabilities. [...]&lt;/p&gt;
&lt;p&gt;We estimate maximal cyber capabilities by &lt;strong&gt;running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity&lt;/strong&gt;. Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.&lt;/p&gt;
&lt;p&gt;The models &lt;strong&gt;identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure&lt;/strong&gt; to obtain test solutions directly from Hugging Face’s production database. All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It's pretty clear what happened here. OpenAI removed safety filters for an in-progress model, locked it up in a sandbox and told it to solve the ExploitGym problems. Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead.&lt;/p&gt;
&lt;p&gt;OpenAI's sandbox used the now common pattern of allowing access to an allow-list of sites related to installing additional packages via a controlled proxy. In OpenAI's words:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;While operating in our sandboxed testing environment, our models &lt;strong&gt;spent a substantial amount of inference compute finding a way to obtain open Internet access&lt;/strong&gt;, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited &lt;strong&gt;a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy&lt;/strong&gt;. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So step one was to break out onto the public internet. The model then broke into Hugging Face to find the answers:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;After gaining Internet access, the models &lt;strong&gt;inferred that Hugging Face potentially hosted models, datasets and solutions&lt;/strong&gt; for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, &lt;strong&gt;the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities&lt;/strong&gt; to find a remote code execution path on the Hugging Face servers.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Chaining together multiple attack vectors is &lt;em&gt;exactly&lt;/em&gt; the kind of thing these new models can do, where previous generations of models might have failed.&lt;/p&gt;
&lt;p&gt;I wrote last month about how &lt;a href="https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/"&gt;Claude Fable is relentlessly proactive&lt;/a&gt;, when I noticed it spinning up custom web servers and deploying CORS tricks on my own laptop just to help debug a WebKit CSS issue. It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they &lt;em&gt;will figure it out&lt;/em&gt;.&lt;/p&gt;
&lt;h4 id="resist-the-temptation-to-write-this-off-as-a-stunt"&gt;Resist the temptation to write this off as a stunt&lt;/h4&gt;
&lt;p&gt;There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term "marketing" in &lt;a href="https://news.ycombinator.com/item?id=48997548"&gt;the Hacker News discussion&lt;/a&gt; of the incident.&lt;/p&gt;
&lt;p&gt;To those people I say &lt;em&gt;pull your heads out of the sand&lt;/em&gt; - you're now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here!&lt;/p&gt;
&lt;p&gt;The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that "autonomous exploit development by frontier AI agents is no longer a hypothetical capability", and this incident is a perfect example of exactly that.&lt;/p&gt;
&lt;h4 id="the-asymmetry-is-increasingly-frustrating"&gt;The asymmetry is increasingly frustrating&lt;/h4&gt;
&lt;p&gt;One of the most infuriating details of this story is how Hugging Face, faced with an accidental and aggressive attack from one of OpenAI's models, were unable to then turn to OpenAI's models to help them fend off the attack.&lt;/p&gt;
&lt;p&gt;The frontier models we have access to are increasingly being constrained in how much they can help us protect our software, heavily influenced by the US government's ongoing threat of export controls.  Claude Fable 5 wouldn't even &lt;a href="https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#proofreader"&gt;proofread this article&lt;/a&gt; for me! It insisted on downgrading me to a less capable model.&lt;/p&gt;
&lt;p&gt;Meanwhile open weight models from China such as GLM-5.2, Kimi 3 and the new Qwen 3.8 Max appear to have none of these restrictions - and any restrictions that &lt;em&gt;do&lt;/em&gt; exist can likely be fine-tuned out of them by modifying the weights&lt;/p&gt;
&lt;p&gt;These constraints are meant to make us safer. I think there's a risk that they are having the opposite effect.&lt;/p&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="sandboxing"/><category term="security"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="llms"/><category term="hugging-face"/><category term="anthropic"/><category term="paper-review"/><category term="ai-security-research"/><category term="openai-hugging-face-incident"/><category term="accidental-cyberattacks"/></entry><entry><title>A Fireside Chat with Cat and Thariq from the Claude Code team</title><link href="https://simonwillison.net/2026/Jul/21/cat-and-thariq/" rel="alternate"/><published>2026-07-21T12:54:02+00:00</published><updated>2026-07-21T12:54:02+00:00</updated><id>https://simonwillison.net/2026/Jul/21/cat-and-thariq/</id><summary type="html">&lt;p&gt;Earlier this month I hosted a fireside chat session at the &lt;a href="https://www.ai.engineer/worldsfair/2026"&gt;AI Engineer World's Fair&lt;/a&gt; with Cat Wu and Thariq Shihipar from Anthropic's Claude Code team. We talked about Claude Code, Claude Tag, Fable, coding agent security, evals, tool design, and how Anthropic use these tools themselves.&lt;/p&gt;
&lt;p&gt;The full video of the session is now available &lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g"&gt;on YouTube&lt;/a&gt;. Below is an edited copy of the transcript, with extra links and my own bolded highlights.&lt;/p&gt;
&lt;iframe style="margin-top: 0.5em; margin-bottom: 1em;" width="560" height="315" src="https://www.youtube-nocookie.com/embed/uU5Gv2h8-9g" title="SimonThis Year in Claude" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen="allowfullscreen"&gt; &lt;/iframe&gt;

&lt;p&gt;A few top-level notes if you don't want to watch the video or wade through the whole transcript:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Claude Tag (Claude's new collaborative Slack integration) now lands &lt;strong&gt;65% of the product engineering PRs&lt;/strong&gt; for the Claude Code team.&lt;/li&gt;
&lt;li&gt;Claude Code ships features to Anthropic employees first, and &lt;strong&gt;only ships the features that demonstrate user retention with that cohort&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Critical changes to Claude Code are still reviewed manually, but the team increasingly relies on automated code review for the "outer layers" of the product.&lt;/li&gt;
&lt;li&gt;Adding examples to a system prompt is &lt;strong&gt;no longer best practice&lt;/strong&gt; for models like Fable 5 or even Opus 4.8. The Claude Code system prompt recently &lt;strong&gt;reduced in size by 80%&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Likewise, lists of "&lt;strong&gt;don't do X and don't do Y&lt;/strong&gt;" can reduce the quality of results from the latest models.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://en.wikipedia.org/wiki/Eating_your_own_dog_food"&gt;Dogfooding&lt;/a&gt; inside Anthropic is called "&lt;strong&gt;ant fooding&lt;/strong&gt;".&lt;/li&gt;
&lt;li&gt;Anthropic &lt;strong&gt;really believe in their &lt;a href="https://code.claude.com/docs/en/auto-mode-config"&gt;auto mode&lt;/a&gt;&lt;/strong&gt;, and see that as an enabling technology for Claude Tag.&lt;/li&gt;
&lt;li&gt;Thariq advises offsetting coding-agent-induced &lt;a href="https://simonwillison.net/2026/Feb/15/deep-blue/"&gt;Deep Blue&lt;/a&gt; by "&lt;strong&gt;being more ambitious&lt;/strong&gt;" with the work you take on.&lt;/li&gt;
&lt;li&gt;Fable is &lt;strong&gt;competent at editing video&lt;/strong&gt;, and Thariq &lt;a href="https://twitter.com/trq212/status/2064826394589442448"&gt;used it&lt;/a&gt; to edit its own launch video.&lt;/li&gt;
&lt;li&gt;Anthropic's culture of working (internally) in public is key to their success, as demonstrated by the way they use Claude Tag in their public Slack Channels.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="how-has-what-you-do-day-to-day-changed-in-the-past-year-"&gt;How has what you do day-to-day changed in the past year?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=65s"&gt;1:05&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Claude Code came out in February of last year — it's under a year and a half old, and it was originally just a bullet point on &lt;a href="https://www.anthropic.com/news/claude-3-7-sonnet"&gt;the Claude Sonnet 3.7 launch&lt;/a&gt;. &lt;strong&gt;How has what you do on a day-to-day basis changed in the past year&lt;/strong&gt;, now that we have these coding agents that actually work for us?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; I remember when we first came out with Claude Code and Sonnet 3.7, you would give it a task and you would have to closely monitor every single little thing it tried to do. I would read every permission prompt extremely carefully. I would frequently say no — no, no, no, did you check this file? Did you check that file? And now it's been incredible with every model generation. I feel like &lt;strong&gt;we've all gotten a chance to take a step back and delegate a lot more of the menial implementation to Claude&lt;/strong&gt;. It's freed up a lot of our time to think about more creative work, like: what is the right experience that we should be providing to our users, now that we know Claude Code can implement a lot of it? And now with Fable it's a totally different step change improvement. &lt;strong&gt;We see for a lot of our use cases that you can actually one-shot a ton of features with Fable now&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; I remember the first text I got about Claude Code. One of my best friends was like, "You need to go try Claude Code." It was about when Opus 4 came out, and I tried it and I was like, "Oh, shit. I need to work at Anthropic now." And that was Opus 4 — great model, but you were reading permission prompts. It's kind of crazy how much amnesia we have, where I'm like, oh, auto mode has always been here, right? I don't even remember pressing yes and allow. For me, the big thing I'm trying to push myself on is that &lt;strong&gt;we have to do higher quality work than we've ever done before&lt;/strong&gt;. The outputs are incredibly high quality. &lt;strong&gt;I've been using it to edit videos a bunch&lt;/strong&gt;, and I'm like, okay, it has to meet the very exacting demands of our brand team in a couple of hours or we just can't do it. &lt;strong&gt;That's how I'm trying to shift with Fable: the best work we've ever done, faster than we've ever done it before&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="what-piece-of-conventional-software-engineering-no-longer-holds-"&gt;What piece of conventional software engineering no longer holds?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=219s"&gt;3:39&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; What's a piece of conventional software engineering that was true a year ago that you don't think holds anymore in this new world?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; One of the biggest shifts we're seeing in the eng skill set: two years ago it was pretty typical for a product manager to go talk to a bunch of customers, align over the course of six months with cross-functional teams on some PRD, and write a thorough spec on exactly how we'll implement this before the first line of code gets written. Now things are completely turned the opposite way. For a lot of engineers, the push I would give to folks in the room is to &lt;strong&gt;develop more of your business sense and product sense on what it is we should build&lt;/strong&gt;, because the timeline between having an idea and building it is so much shorter — it's down from six to twelve months to maybe even a week. That means all of us need to have better taste on what is worth building, what will actually inflect the businesses we're working on. So it's &lt;strong&gt;an increase in value on product taste and business sense&lt;/strong&gt;, and a bit lower on execution in most product domains. Of course, for infra there's still a very heavy emphasis on making sure all the details are right.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; For me, it's that &lt;strong&gt;rewrites are now good&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; The worst thing you could do is now actually fine!&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; Exactly. All the Mythical Man-Month stuff — never rewrite — I'm pro-rewriting now. If you have a good test suite — and &lt;strong&gt;I think the rewrite actually forces you to make sure you have a good test suite&lt;/strong&gt; — but I think what people undercount is that &lt;strong&gt;a codebase is a spec, and maybe it's the only copy of the spec that you have&lt;/strong&gt;, because no one knows every branching part of the codebase. You can take this as an artifact and distill it or create other versions of it. We &lt;a href="https://bun.com/blog/bun-in-rust"&gt;rewrote Bun in Rust&lt;/a&gt; and it works great — it's live for me right now.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; You're not shipping Claude Code on Bun-in-Rust yet, right?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; Internally we have.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;em&gt;(Actually it looks like Anthropic started shipping Claude Code on Bun-in-Rust to everyone &lt;a href="https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/"&gt;on June 17th&lt;/a&gt;.)&lt;/em&gt;&lt;/p&gt;
&lt;h4 id="what-kind-of-things-are-non-engineers-doing-with-claude-tag-"&gt;What kind of things are non-engineers doing with Claude Tag?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=396s"&gt;6:36&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; The other big launch recently was &lt;strong&gt;&lt;a href="https://www.anthropic.com/news/introducing-claude-tag"&gt;Claude Tag&lt;/a&gt;&lt;/strong&gt; — that's what, a week old now, at least for the rest of us. I understand it's being used at Anthropic by non-engineers a great deal. &lt;strong&gt;What kind of things are non-engineers doing with Claude Tag?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Claude Tag is a Claude that lives in your team's collaboration tools. We launched it last week within Slack. &lt;strong&gt;The thing that's different about Claude Tag is it's multiplayer by default&lt;/strong&gt;. Once you add Claude Tag to a Slack channel, you can chime in, your teammates can chime in, and you can collaborate together on the PR. The other big difference is that it's proactive instead of reactive. You can tell Claude Tag, "Hey, monitor every bug report in this channel, put up a PR to fix it, and tag the engineer who last touched this part of the codebase," and it'll do it for the lifetime of the channel without you having to manually tag it in. And the third big shift is that &lt;strong&gt;we've &lt;a href="https://claude.com/docs/claude-tag/users/memory"&gt;added team memory&lt;/a&gt; into this&lt;/strong&gt;. If you tell Claude Tag your preferences in the channel, it'll remember them for every future post. If you always want it to debug outages but you don't want it to debug warnings, just tell it that in natural language in the channel and it'll remember it for you and everyone else on your team.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Internally, we see Claude Tag as the evolution of Claude Code.&lt;/strong&gt; We see this as a large shift in how we work internally. &lt;strong&gt;Claude Tag currently lands 65% of our product eng PRs.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; For all of Anthropic, or just for Claude Code?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; This is just for our product engineering team — &lt;strong&gt;our internal version of Claude Tag lands 65% of our product PRs right now&lt;/strong&gt;. And this is a huge shift; this is more than 50% of our PRs. The way we see people split work between Claude Code and Claude Tag is: Claude Code is still the best place for your most complex tasks, when you're interactively iterating with the agent. &lt;strong&gt;But Claude Tag is great for having it work proactively on your behalf&lt;/strong&gt;, so you no longer need to manually kick off Claude Code for all the bug reports that come up for features you're working on.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; And for non-coding cases: for example, before this talk we asked Claude Tag, "Hey, when is Fable releasing?" We wanted to make sure we'd line it up with the announcement. Claude Tag would search our Slack and look at who's been saying what. &lt;strong&gt;As a search engine for your company, it's really valuable.&lt;/strong&gt; It has all the context for your product, so you can ask it metrics-related questions — often when you're making decisions you want them informed by what the metrics say, so you hook it up to your event store. I've seen our marketing team do things like, "Hey, tell me about this feature." They're not programmers, but Claude is a programmer — it can clone the codebase and say, "This is the feature, this is what it looks like, &lt;strong&gt;this is a recording of me using the feature&lt;/strong&gt;." It enables a whole wide variety of things, and I think we're still early in figuring that out.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="claude-tag-as-the-team-collaborative-layer"&gt;Claude Tag as the team collaborative layer&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=606s"&gt;10:06&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; One of the problems I've had with coding agents is that I get how to use them as an individual, but I'm not really clear on how to use them in a team environment. &lt;strong&gt;It sounds like Claude Tag is your current answer to that team collaborative layer for this stuff.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Exactly. And a large percentage of our sessions are actually multiplayer right now. Maybe I say, "Hey, I think we should implement this new feature in Cowork," and I'll tag in Claude Tag to do a first pass at it. Then I'll tell Claude Tag, "Share a recording of your final implementation," and I'll tag in design to take a look. They'll nudge it, then pass it on to eng to take it to the finish line and get it out to prod. It's been this very fluid experience. &lt;strong&gt;We're still trying to iron out what the social dynamics are for steering the same session&lt;/strong&gt;, but we've found that people just observe how others use it and follow those social norms — it's been pretty intuitive for us to integrate Claude Tag into our teams.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; It's great for teaching people, and also for reducing slop, because &lt;strong&gt;the fact that everyone is seeing you use Claude together sort of levels up how you use Claude as well&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This reminded me of how Midjourney solved the challenge of teaching people advanced image prompting by enforcing prompting in public in their Discord channels.&lt;/p&gt;
&lt;h4 id="how-do-you-decide-which-features-are-worth-building-when-building-is-so-much-cheaper-"&gt;How do you decide which features are worth building when building is so much cheaper?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=701s"&gt;11:41&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Something I've found really hard myself is knowing when a feature is worth shipping now that the cost of actually building features has dropped so much.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; How do you deal with the hardest problem in all of engineering — prioritization? &lt;strong&gt;How do you decide which features are worth building and shipping when building a feature is so much more inexpensive now?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; This is the hard thing. There are a few ways we approach it. One is we dogfood our products every single day. Whenever there's something we want to be able to do in our products that we're not able to, instead of finding a different solution we fix our product so it can support that case. &lt;strong&gt;We have a very heavy dogfooding culture internally.&lt;/strong&gt; Before we share our products with everyone in the world, we share them with everyone within Anthropic, and with some early customers who give us very honest feedback about it — the more brutal the better — and we iterate until people love it. &lt;strong&gt;We have an internal bar for the number of active users and the amount of retention a feature has to have before we share it with the world.&lt;/strong&gt; Because this bar is very clear, every engineer knows what they're trying to hit. I think this also levels up our polish, because if the feature isn't polished, people will churn — and then we shouldn't ship that feature.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Using internal user-retention to decide if a feature should ship makes a whole lot of sense to me.&lt;/p&gt;
&lt;h4 id="do-you-have-an-example-of-a-feature-which-surprised-you-"&gt;Do you have an example of a feature which surprised you?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=774s"&gt;12:54&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; &lt;strong&gt;Do you have an example of a feature which surprised you?&lt;/strong&gt; You rolled it out and the engagement was off the charts — something unlikely to be shipped that turned into a real product thing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; I do have one. &lt;strong&gt;A lot of folks on our team love &lt;a href="https://code.claude.com/docs/en/remote-control"&gt;remote control&lt;/a&gt;.&lt;/strong&gt; Remote control lets you use your mobile device, or Claude in the web browser, to connect to a local Claude Code session running in your CLI. I never have this need, because I just kick off the task directly on mobile and it runs in a cloud session without using my local environment — I think because I'm doing very easy coding tasks. It was something I didn't totally understand; I was like, hey, people should just set up remote dev environments. But in practice, once we rolled out remote control, so many people I talk to told me that what they do every night is plug their laptop into a power charger, open a bunch of remote control sessions, lock the screen, &lt;strong&gt;and then use their mobile phone from their couch to control Claude Code&lt;/strong&gt;. So this has become a flow we're now leaning into that I didn't originally get — but now I do.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="does-a-human-review-every-line-of-production-code-in-claude-code-"&gt;Does a human review every line of production code in Claude Code?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=860s"&gt;14:20&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;One of the over-arching themes of the conference was review: how much attention to people spend to reviewing code written for them by coding agents. I was very keen to hear the Claude Code team's take on this!&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; How does code review work? &lt;strong&gt;Does a human being review every line of production code that makes it into Claude Code?&lt;/strong&gt; And if not, what are you doing — how do you keep the quality up?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; It varies on the task a lot. &lt;strong&gt;For important areas we have code owners.&lt;/strong&gt; The system prompt is an example where we have a code owner — you really need to get their approval.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; So the code owner is directly responsible for the quality of that area of the code.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; That's right.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; And they need to approve any PR that touches it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; We have &lt;a href="https://code.claude.com/docs/en/github-actions"&gt;our code review GitHub bot&lt;/a&gt; review everything — that goes on every PR, and often it's doing the bulk of the review. Something I've seen on the team is that &lt;strong&gt;for more complex PRs you might make an artifact to explain the PR&lt;/strong&gt; so that other people can then review. And we invest a lot into verification, CI/CD, things like that, to make sure that any time anything fails we have a test. We have a really robust environment where Claude can control Claude Code and test it. So there's a multi-pronged approach to code review.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; In general, &lt;strong&gt;we are trying to move to a world where humans don't need to be in the loop&lt;/strong&gt;. For the most critical changes to the core of Claude Code, and the cores of other products, there is always a code owner and they do manually review all the changes. But increasingly, &lt;strong&gt;for the changes at the outer layers, we actually have Claude code review fully review those&lt;/strong&gt;. That sounds pretty scary, but we've had a six-plus-month-long process to get here, and &lt;strong&gt;there are baby steps that you take to build up trust with code review&lt;/strong&gt;. In the beginning we had human review for everything, and then increasingly we would say, &lt;strong&gt;okay, for code changes that touch these files, code review is catching 100% of the issues there — so we actually don't need a human manually reviewing those&lt;/strong&gt;. And when we have incident review, &lt;strong&gt;we look at the PRs that caused the incident and say, okay, how do we update code review to catch that?&lt;/strong&gt; — and we take those PRs and &lt;strong&gt;add them to an eval set&lt;/strong&gt; to make sure our future changes to code review never regress that metric. Removing humans from the code review loop is a big step forward. It can sound scary, and it's not something you can do overnight, but it is something you can do &lt;strong&gt;through many months of investment in the infrastructure&lt;/strong&gt; to give you the confidence that code review is catching everything you care about.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So the key seems to be constantly iterating on the automated review systems themselves, in order to build trust in them over time.&lt;/p&gt;
&lt;h4 id="how-does-a-new-model-affect-your-intuition-for-what-it-can-and-can-t-do-"&gt;How does a new model affect your intuition for what it can and can't do?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=1040s"&gt;17:20&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;We got &lt;em&gt;deep&lt;/em&gt; into evals - another hot topic throughout the wider conference.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; I know that Opus 4.8, if I ask it to build me a JSON endpoint that runs a SQL query and outputs JSON, is just going to get it right — that's not something I have to review closely. But then a new model comes along and I don't know how to build trust in Fable quickly, that it's not going to mess things up that Opus didn't. &lt;strong&gt;How does the new model affect your intuition for what it can do and what it can't do?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; The main reason we're building up this &lt;strong&gt;eval base over time is so that new models can be a drop-in replacement&lt;/strong&gt;. When we have a new model, we run the whole eval set and make sure that, for example, Fable is strictly better than Opus 4.8 — and that gives us the confidence to drop it in.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Are those model evals for Anthropic as a whole, or Claude Code team-specific?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; We have both. We have evals on our team, and we run code review across every repo within Anthropic, so we have evals for that. And for things like auto mode, we not only have evals across every user within Anthropic — we've also commissioned multiple external testers to red team it, to create environments with prompt injections and malicious inputs, &lt;strong&gt;and make sure that auto mode doesn't let any of those pass&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="how-do-you-build-confidence-that-a-system-prompt-tweak-results-in-better-output-"&gt;How do you build confidence that a system prompt tweak results in better output?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=1121s"&gt;18:41&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; I want to know if the system prompt improvement I made actually improved the product — that's the most basic form of product-specific eval, and I still don't have a great feel for how to do that. &lt;strong&gt;Is that something you're doing such that you have complete confidence that a tweak you've made to the system prompt results in better output?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; &lt;strong&gt;We don't have complete confidence, but we do a lot to make sure that we don't regress performance.&lt;/strong&gt; The starting point is a suite of external evals that we trust, and we complement that with an even larger suite of internal evals that we trust. To start, &lt;strong&gt;we mainly optimize for capability&lt;/strong&gt;: given a complete definition of a task and the full codebase, does Claude make the right decisions, fully fix the bugs, and pass all the tests? That's the starting point and the thing we optimize for, because it's most directly what users want. But there are a lot of behaviors that impact how users feel when they work with Claude Code. For example, &lt;strong&gt;people really don't like it when Claude Code says it's time to go to sleep.&lt;/strong&gt; Or people really don't like it when it says, "Hey, I finished two out of five parts — do you want me to continue?" Yes, please continue. &lt;strong&gt;So we're building up a set of behavioral evals to catch these.&lt;/strong&gt; And as we get user feedback — please be loud with us about your user feedback — we rank the priority issues and go down one by one and build evals for each of them. It's not 100% coverage, but it is a priority for us to increase the coverage.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="how-much-interaction-is-there-between-the-claude-code-team-and-the-model-training-teams-"&gt;How much interaction is there between the Claude Code team and the model training teams?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=1221s"&gt;20:21&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; &lt;strong&gt;How much interaction is there between the Claude Code team and the teams at Anthropic who are training the models in the first place?&lt;/strong&gt; Is that quite a close collaboration?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Across Anthropic, we all work quite closely together. We meet often to talk about what we expect the next generation of models to be able to do. Our research team has also been amazing about showing this publicly — we often talk in our blog posts about how &lt;strong&gt;we're targeting ever-increasing longer-horizon work&lt;/strong&gt;, and how we train Claude itself to be honest, harmless, and helpful. We also put a lot of effort into making sure it's aligned with your intent, even if your intent is expressed in a fuzzy way. Of course, try your best to be specific about what you want, so Claude has all the context — but even when you're not specific, we teach Claude to make good assumptions. It's been a productive partnership.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="the-system-prompt-has-been-reduced-by-80-what-have-you-been-able-to-drop-"&gt;The system prompt has been reduced by 80% — what have you been able to drop?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=1284s"&gt;21:24&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;So many useful prompting tips in this section!&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Thariq, you &lt;a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;amp;t=358s"&gt;mentioned this morning&lt;/a&gt; that the &lt;strong&gt;system prompt for Claude Code has been reduced by 80% because of Claude Fable&lt;/strong&gt;. Can you go into a little more detail? &lt;strong&gt;What kind of things have you been able to drop?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; It wasn't just Fable — it was Opus 4.8 as well, and going forward, future models. We have different system prompts for different models now. One of the patterns we saw is that we were over-constraining Claude. The initial, maybe Opus 4-ish models wanted a lot of examples, and &lt;strong&gt;removing examples was extremely helpful&lt;/strong&gt;, because it was just more creative than the examples we gave it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; That's really interesting, because one of the top prompting tips I give people is: give it examples. If that's no longer true, that kind of breaks my prompting model a little bit.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; Same here — I was surprised to hear that. I think now it's more about the shape of what you give it — the tools you give to Claude, your system prompt, things like that. The other thing we did is try to give it more context and &lt;strong&gt;fewer "do not do this"&lt;/strong&gt; instructions, because that's a very strong impulse for Claude, and especially if it conflicts with user instructions later on, that can be extremely confusing to Claude — "I've got this skill that says this and the system prompt says this." So we try to &lt;strong&gt;have fewer hard constraints, more context, and fewer instructions overall&lt;/strong&gt;. It's definitely a science — it took a bunch of evals to build.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; In general, when you're prompting these models, you should always think: &lt;strong&gt;are there edge cases to the instruction that I'm giving it?&lt;/strong&gt; When we went back and reviewed all the instructions in the Claude Code system prompt, &lt;strong&gt;we found a few cases where yes, this statement is 90% true, but there's a real 10% of cases where it's not true&lt;/strong&gt;. We didn't want to constrain the model, or confuse it into thinking it should always do this. One good example is verification. Everyone here wants Claude to verify its work, and we had some instructions in the prompt that said: if you make a front-end change, always verify. But there's a limit to it. If it's changing copy from one string to another string, and the user says "just make a quick fix and update the test," maybe you don't want to verify. &lt;strong&gt;So we've adjusted our wording from "always verify, verify, verify" to something like: most of the time when you're doing front-end work you can't fully understand the experience by hitting the backend endpoints, so when you make larger changes to the user experience, please run the app locally.&lt;/strong&gt; And in fact, that instruction probably isn't even good either, because &lt;strong&gt;what is a large change?&lt;/strong&gt; Maybe it should test small changes too. In general, whenever you give a prompt to the model, &lt;strong&gt;you should think about the ways in which it could be misinterpreted by a well-intentioned human&lt;/strong&gt;, in order to better understand how the model might interpret it — and &lt;strong&gt;soften the prompt&lt;/strong&gt; so that it's actually 100% accurate, because you're giving this prompt to the model 100% of the time.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; What's fascinating about that is you're &lt;strong&gt;relying on the model's judgment&lt;/strong&gt; — and that's got to be an Opus/Fable-level thing. Models a year ago did not have the level of judgment necessary to decide whether they were going to test a change or not. But that does break down if you're building for a wide range of models and trying to run the cheaper models for cheaper tasks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; We actually have &lt;strong&gt;a different system prompt per model now&lt;/strong&gt;, for this very reason. It's only our most frontier models that have this 80% token decrease — the older models still have the full system prompt.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Do you think Fable and Opus are smart enough to prompt Haiku with more details, because they understand that Haiku has less judgment, less taste?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; We haven't been able to eval it — we don't have any hard data to show it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; There's a tough thing with smaller models sometimes, because &lt;strong&gt;sometimes the larger models can be more token-efficient on a hard problem than the smaller models&lt;/strong&gt;. So there's a bit of intuition to build there — sometimes you really just want frontier intelligence almost all the time. The Pareto curve shifts, and it's hard to find.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; A year ago I did not trust a model to write a prompt. Today the good models are very good at prompting — a lot of my prompts are written by models, which feels absurd but works really well. What helped me come to terms with that was thinking about subagents, which are entirely about a Claude model setting up a prompt for another Claude model.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; &lt;strong&gt;Workflows&lt;/strong&gt; are actually a really good example of this, because it's Claude not just prompting a single subagent, but prompting the orchestration of many subagents, and each one of them gets a very detailed prompt. It's almost a level above just spawning a subagent. I've also been using it on my personal machine, &lt;strong&gt;giving it the Gemini API and saying: here, generate images&lt;/strong&gt;. It's way less lazy than I am at prompting an image model. It's just Claude prompting Claude all the way down.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; I think Claude also wrote the prompt for &lt;a href="https://code.claude.com/docs/en/workflows"&gt;the workflow tool&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; I've read that prompt — it's a good prompt. That's actually a frustration I have with Anthropic generally: you &lt;a href="https://platform.claude.com/docs/en/release-notes/system-prompts"&gt;publish the prompts for Claude Chat&lt;/a&gt;, but you don't include the tool prompts and the Claude Code prompts. I still have to run a proxy to intercept them. &lt;strong&gt;I would love it if the Claude Code prompts were deliberately published&lt;/strong&gt; — they're the documentation. They're how you know what the tool can do and how it works.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; I'll write down that feature request. I'll have Claude Tag do it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Interesting to note that OpenAI's &lt;a href="https://developers.openai.com/api/docs/guides/latest-model?model=gpt-5.6#favor-leaner-prompts"&gt;prompting best practices for GPT-5.6&lt;/a&gt; includes similar advice for their latest models:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Favor leaner prompts&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Removing repeated instructions and examples and simplifying tool descriptions can improve task performance and token efficiency. In a sample of internal coding-agent eval runs, configurations with leaner system prompts improved evaluation scores by roughly 10–15% while reducing total tokens by 41–66% and cost by 33–67%.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h4 id="what-s-your-bar-for-introducing-a-new-tool-"&gt;What's your bar for introducing a new tool?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=1686s"&gt;28:06&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Claude Code is basically a big bag of tools. &lt;strong&gt;What's your bar for introducing a new tool?&lt;/strong&gt; How do you decide when it's worth doing that additional engineering at that level?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Do you want to take it? You introduced one of the best tools we have.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; My career peaked when I introduced the ask user question tool. It's really hard. Especially for some tools — &lt;strong&gt;ask user question is Claude's tool to ask you&lt;/strong&gt; — so it's hard to eval, and sometimes it's more of a user preference thing. Back then we had fewer evals, so it was very dogfooding based — or "ant fooding," our ant version of that. But overall &lt;strong&gt;we've been trying to trend towards fewer tools&lt;/strong&gt;. The last set of tools we introduced was the task tool, I think — and we try to give Claude more general versions to do things.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="what-s-the-latest-evolution-of-your-file-editing-tool-"&gt;What's the latest evolution of your file editing tool?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=1743s"&gt;29:03&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I have a long-running fascination with file editing tools - they were the subject of the &lt;a href="https://aider.chat/docs/leaderboards/edit.html"&gt;old Aider code editing leaderboard&lt;/a&gt;, and I've watched with interest as they've evolved in different coding agents from search-and-replace based to line-number-based to more complicated patterns.&lt;/p&gt;
&lt;p&gt;The Claude API docs describe a &lt;a href="https://platform.claude.com/docs/en/agents-and-tools/tool-use/text-editor-tool"&gt;text editing tool&lt;/a&gt; that's recommended for building against the API, but Claude Code seems to use slightly different approaches here.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; One of the most interesting tools is the file editing tool — you can have file editing as a tool, or you can tell it to use sed and grep and do things that way. &lt;strong&gt;What's the latest evolution of your file editing tool?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; We still have one, but for example we removed our grep and other search tools — glob tools — in favor of native bash. Like I said in my talk earlier, &lt;strong&gt;the models are kind of more of a biology than a physics&lt;/strong&gt;, and tool design especially is quite hard. I'm not sure if Cat disagrees and thinks there's a science to the eval of it, but I think tool design is more of an art, maybe — or a biology.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; I largely agree, but in general as we introduce more tools, we try to keep the cardinality pretty low and make sure that &lt;strong&gt;every tool we add has a distinct function from every other tool, so that Claude can very easily distinguish when to call each&lt;/strong&gt;. For file edit, the reason we have it is actually because we can render it. We show people when Claude makes a file change, and there's this &lt;strong&gt;nice dedicated UI&lt;/strong&gt; that says: do you approve this edit to this file? &lt;strong&gt;The reason we had a dedicated file edit tool was so that we could deterministically know&lt;/strong&gt; that Claude was making a file change, so we could show people this nice UI. A lot of new users onboarding still really like this experience, so we've kept it around. But for a lot of us who are on auto mode right now — hopefully you're not on YOLO mode — I don't think it actually matters, and we could probably just remove file edit and be totally fine.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="what-s-the-advice-within-anthropic-for-safely-running-claude-code-"&gt;What's the advice within Anthropic for safely running Claude Code?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=1858s"&gt;30:58&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;It's the &lt;a href="https://simonwillison.net/tags/prompt-injection/"&gt;prompt injection&lt;/a&gt; question! Who better than Anthropic employees to explain how Anthropic sees the risk of prompt injection attacks causing their Claude Code instances to run amok?&lt;/p&gt;
&lt;p&gt;It turns out they &lt;em&gt;really&lt;/em&gt; trust their &lt;a href="https://code.claude.com/docs/en/auto-mode-config"&gt;auto mode&lt;/a&gt; - and see that as the feature that enabled Claude Tag.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Let's talk about safety and security. I am deeply aware of the risks of prompt injection, and there are so many bad things that can happen if somebody else tells my Claude Code what to do. I still mostly run Claude Code in YOLO mode and feel incredibly guilty about it. &lt;strong&gt;What's the advice within Anthropic for safely running Claude Code?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Why not auto mode?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; I am starting to use auto mode, but I don't understand it enough to get how safe it is. As of maybe three weeks ago, I'm defaulting to auto mode.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Broadly within Anthropic, almost every single person uses auto mode. It is the best way to do long-running work in Claude Code while being safe. &lt;strong&gt;We've done extensive bashing. We have thousands of evals. We've commissioned many red teamers to create adversarial environments in order to trick Claude Code into doing bad actions, and we've mitigated every single issue that they found.&lt;/strong&gt; We're going to publish some evals in the coming weeks, but we've pretty much mitigated every attack.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; That is a big claim.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; We'll share the evals for it so folks can assess, but we've been extremely diligent about identifying all the ways in which Claude might mess up and then updating auto mode to counter it. It doesn't catch 100% of things — that would be way too strong a claim. But &lt;strong&gt;for the main categories of risks that we're concerned about, like prompt injection and data exfiltration, the risks are far lower than the average human reviewer&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I am very much looking forward to learning more about their evals and approach to verifying auto mode.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; A little on how auto mode works — it's useful to build this mental model. Whenever Claude is doing a turn, or a bash call, there's &lt;strong&gt;a Sonnet classifier&lt;/strong&gt; that is judging the tool call and also the context of the conversation — your instruction. There are some things around permissions that are dependent on your request: you don't want to give git push permissions all the time, but if you say "push this to GitHub," you want it to do it — and if you say "don't push," you want it to deny it. Auto mode will do that. That particular thing happens to me a lot, where Claude tried to do something because it's very helpful and proactive, and auto mode saw "don't do this" and surfaced it. &lt;strong&gt;So it's good at the dynamic permissions&lt;/strong&gt; that you yourself give inside the prompt, which I think is really important. It also works well with our &lt;a href="https://code.claude.com/docs/en/sandbox-environments#sandboxed-bash-tool"&gt;sandboxing infrastructure&lt;/a&gt;, because sandboxing is one of those things where there are so many different edge cases that it's hard for us to deterministically follow them. &lt;strong&gt;We have a sandbox, and when something needs to escape the sandbox&lt;/strong&gt; — like a network request — auto mode can look at that request and ask: does this make sense? — and allow it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; I hadn't realized auto mode is interacting with the networking sandbox as well.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; It interacts with any permission prompt the user would otherwise see.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; How old is auto mode? As a feature I had access to, it's only a couple of months old, right?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(It was first made available to the public &lt;a href="https://claude.com/blog/auto-mode"&gt;on March 24th&lt;/a&gt;.)&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; We've been using it within Anthropic &lt;strong&gt;since January&lt;/strong&gt;, so we've been hardening it for quite a while. Anthropic is extremely focused on safety and security, and we've been working broadly across our alignment and safeguards teams to enable the rollout internally, build out these evals, and make auto mode even more robust before sharing it with the world.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; This is also the reason Claude Tag is so good — &lt;strong&gt;Claude Tag uses auto mode&lt;/strong&gt;. I've heard a lot of build-versus-buy questions about a Slackbot, and I'm like: please, you probably shouldn't build your own AI Slackbot. There are so many attack vectors. &lt;strong&gt;You have a feedback channel that users can post feedback into, and now your bot is reading it.&lt;/strong&gt; The work we've put in with auto mode — and we have a general &lt;strong&gt;Swiss cheese defense&lt;/strong&gt; for security; we also RL against this stuff — &lt;strong&gt;I think this is really what makes Claude Tag work&lt;/strong&gt;. It works seamlessly with your permissions, and you don't want to be prompt injected in your Slack.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="are-there-more-security-things-in-the-pipeline-beyond-auto-mode-"&gt;Are there more security things in the pipeline beyond auto mode?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=2154s"&gt;35:54&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Are there any more security things in the pipeline that go beyond auto mode?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; I think we're very secure. &lt;strong&gt;With Claude Tag you can provision your own credentials for Claude&lt;/strong&gt;, so it doesn't need to act on your behalf — you can have Claude as an identity, and that also makes it easier to audit and inspect what Claude is doing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Because Claude Tag is influenced by anyone who can talk to it — it's got a much wider pool of people telling it what to do.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; That's right. And of course we have probes as well with Fable, which is a downstream effect of our safety and research work. I think this is the moment where you see Anthropic being an AI safety company really paying off: &lt;strong&gt;we really want Claude to be able to run in an aligned way over long periods of time&lt;/strong&gt;, and &lt;strong&gt;auto mode has to be basically flawless for this to work&lt;/strong&gt; — it's all downstream of our being an AI safety company.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; We also launched trusted devices for the remote control users out there who want to be safer. And for all of our remote environments, we support &lt;strong&gt;credential injection&lt;/strong&gt;. If you want Claude Code to be able to access Datadog, but you don't want Claude Code itself to hold the Datadog credential, you can set up our identity and credential management system &lt;strong&gt;so that the Datadog credentials are only usable by the agent but not accessible by the agent&lt;/strong&gt; — we insert them on the fly when the agent tries to make a Datadog request.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I really like that credential injection pattern, where Claude Code can access an API via a proxy and that proxy both audits the request and injects the relevant API key - so Claude can access authenticated endpoints without having access to the API credentials itself.&lt;/p&gt;
&lt;h4 id="how-has-the-past-year-and-a-half-changed-how-you-think-about-your-own-craft-"&gt;How has the past year and a half changed how you think about your own craft?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=2273s"&gt;37:53&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Thariq &lt;a href="https://www.youtube.com/watch?v=9fubhllmsBU&amp;amp;t=867s"&gt;talked about a sense of grief&lt;/a&gt; brought on by Fable-class models in his keynote in the morning, and we dived further into that as part of our conversation. I've been calling this &lt;a href="https://simonwillison.net/2026/Feb/15/deep-blue/"&gt;Deep Blue&lt;/a&gt;.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; &lt;strong&gt;Let's talk a little bit about the human element.&lt;/strong&gt; &lt;strong&gt;A lot of people are feeling a sense of loss now that so much of what they considered to be their role in building software is being subsumed by the models.&lt;/strong&gt; How do you think about that? &lt;strong&gt;How has the past year and a half changed the way you think about your own craft and the value that you add?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; Cat and Boris are such good reminders that you have to be more ambitious. They're always like: we're growing so fast, we have to be on the edge, we have to do the best work we can. That's a constant reminder for me — any time I'm slow on something, I'm like, okay, can I do it faster? Can I be more ambitious here? And oftentimes the answer is Claude, because Claude is getting better as you go — the last time I tried this, it was with the previous model. On your point about loss: I think this is real. &lt;strong&gt;If you're only trying to do the same work you were doing before LLMs, and now it's a prompt, it is, I think, kind of a sad feeling.&lt;/strong&gt; And &lt;strong&gt;the way you offset that is by being more ambitious.&lt;/strong&gt; I think Jared is such a good example — he hand-wrote all of the Zig code in his Oakland apartment in about a year, barely left his house, and had so much fun doing that. Now I see him rewrite all of Bun into Rust and &lt;strong&gt;he's having so much fun doing that&lt;/strong&gt; — it's so much more ambitious, and that's how he offsets it. Generally it's asking &lt;strong&gt;how do I do the bigger thing&lt;/strong&gt; and do more — &lt;strong&gt;I think success is fun&lt;/strong&gt;. It's changing your ambition.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;"The way you offset that is by being more ambitious" neatly captures where I've landed on this issue myself as well.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; And Cat, what does that look like from a product management perspective?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; I feel like the product role just changes every single month. &lt;strong&gt;All the PMs on our team are this mix of engineer, designer, PM&lt;/strong&gt; — most of them actually used to be full-time engineers. For us it really means &lt;strong&gt;plugging in whenever there's any kind of gap&lt;/strong&gt;. If we have an idea and we didn't inspire any engineer to go build it, then we should just build it, put it into a notebook, and inspire people to take it to production. If the designs look a little off, &lt;strong&gt;let's take a page that's similar, do a first-pass design, and tag in someone who's very detail-oriented to fill in the gaps&lt;/strong&gt;. Or if we notice that our team and product adoption is bigger within the company, and more people need to know what's coming down the pipe for Claude Code, Claude Tag, and Cowork — let's automate figuring out our whole launch calendar, &lt;strong&gt;let's automate getting those status updates asynchronously&lt;/strong&gt; so we're not bugging people, and make sure our updates in our internal announce channels are fully detailed and to the point. For us it's very much understanding &lt;strong&gt;what the gap is right now between a great idea and getting something to our customers&lt;/strong&gt;, and &lt;strong&gt;how do we automate it as much as possible&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This reflects something I've noticed: when you can produce code so much faster, time spent blocked awaiting a decision from someone else becomes a much more notable bottleneck. Engineers who can make product decisions can move a whole lot faster, and the cost of getting one of those decisions wrong is much less prohibitive.&lt;/p&gt;
&lt;h4 id="what-s-a-moment-when-claude-has-surprised-you-"&gt;What's a moment when Claude has surprised you?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=2510s"&gt;41:50&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; &lt;strong&gt;What's a moment when Claude has surprised you?&lt;/strong&gt; When the model did something you didn't think it would be able to do?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; I've posted a lot about Claude video editing, but most recently I gave a talk at the ACM Agentic conference, and I asked, "Hey guys, do you have the edited video? I'd love to post it and share it with my comms team." They said, "Oh, it's taking so long." So I asked for the raw files. They sent me the video of me talking on stage, the video of the deck, and the audio file, and said, "Good luck." I gave this to Claude, along with my HTML deck, and said, "&lt;strong&gt;Hey, can you just edit this together?&lt;/strong&gt;" And what it does is honestly incredible — I'm ready to ship it. It transcribes the entire video. It notices that sometimes the video of my deck is a little weird — there's a popup of an auto-update in the middle — and it goes, "&lt;strong&gt;Oh, I probably shouldn't use the video of your deck. What I'm going to do is slice it up, figure out which slide you're on, and use the HTML source instead.&lt;/strong&gt;" So it displays the HTML source. Then it's got video of me, but I'm only taking up a small part of the stage, so &lt;strong&gt;it's cropping dynamically to where I am on the stage&lt;/strong&gt; — and I'm pacing, so it's tracking me as I pace. And it's transcribing what I'm saying.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; This was Fable, right?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; This was Fable, yeah. It was a good prompt, but it was a one-shot prompt. Then I asked it to add some interesting animations and graphics, and I was just blown away. &lt;strong&gt;It does ffmpeg, it does Remotion.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Here's Thariq's video &lt;a href="https://twitter.com/trq212/status/2064826394589442448"&gt;on how he used Fable to edit Fable's own launch video&lt;/a&gt;, and here's &lt;a href="https://twitter.com/ClaudeDevs/status/2064399512664526853"&gt;that launch video&lt;/a&gt;.&lt;/p&gt;
&lt;h4 id="what-can-t-it-do-yet-"&gt;What can't it do yet?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=2616s"&gt;43:36&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I'm embarrased to admit that I've been finding it quite hard to come up with tasks that frontier models like Fable 5 and GPT-5.6 are unable to accomplish.&lt;/p&gt;
&lt;p&gt;Cat still doesn't rate its UX design skills:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; What can't it do? What are the things where you're still disappointed — where you're waiting for Claude Fable 6 to figure it out for you?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; I want it to have better design and UX taste. It's now at the point where if I write out a prompt with a detailed spec of how I want a feature to behave, it will usually behave that way. But the paddings might be off, or the interface just isn't delightful yet. It leans on existing best practices for how apps are designed, but &lt;strong&gt;for frontier AI products, there are so many new interaction experiences that we have yet to design&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; There's an Opus aesthetic — you can look at something and go, "Yeah, that was designed by Opus." It'd be good if we could move beyond that.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Yeah. I'm very excited for future models to hopefully be &lt;strong&gt;interaction design thought partners&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; What can't it do? I would love to see it interact more with the real world. Can it solve science? Can it orchestrate the experiments? There's some amount of coding that goes into that, but there's also this other taste of the broader world that it needs.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="which-parts-of-anthropic-s-culture-should-other-companies-steal-"&gt;Which parts of Anthropic's culture should other companies steal?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=2711s"&gt;45:11&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I figured this would make a great closing question:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; &lt;strong&gt;Which parts of Anthropic's company culture do you think uniquely help Anthropic be productive with these tools, that other companies should steal?&lt;/strong&gt; What are the cultural hacks people should be adopting from you?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; I'll share one for Claude Tag. &lt;strong&gt;Claude Tag works best when you have it in a public channel, and when most of your channels are public.&lt;/strong&gt; Claude Tag is able to search across all public channels to get as much context as possible to give you the highest-accuracy answer — and &lt;strong&gt;it's only able to do this if it has access to everything&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; I mentioned this in my keynote, but it's so important to me I want to re-emphasize it. The co-founders &lt;strong&gt;say we don't negotiate against ourselves&lt;/strong&gt;, and I think this is really important. &lt;strong&gt;You can imagine trade-offs in your head and talk yourself out of doing something ambitious — or you can just try to do the ambitious thing.&lt;/strong&gt; We're so often asking: what if we just did it? Is this a real trade-off or not? And if so, why — where's the proof that it's a real trade-off, and not just something that sounds reasonable? &lt;strong&gt;Make the trade-offs show themselves to you. Be as ambitious as you can.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="what-s-your-favorite-absurd-thing-you-ve-built-with-claude-just-because-you-could-"&gt;What's your favorite absurd thing you've built with Claude, just because you could?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=2806s"&gt;46:46&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I couldn't resist throwing in this one as well.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; &lt;strong&gt;What's one of your favorite absurd things that you've built with Claude, just because you could build it?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; I'm working on &lt;strong&gt;a 2D Street Fighter fighting game with me as a character&lt;/strong&gt; — and my friends as well. It uses Claude Code to prompt Gemini — and honestly the Seedance model is pretty good — to make video animations. It works great; it's so good at prompting, and it can verify the frames to check whether an animation was good.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; Is this Street Fighter 2-level 2D sprites you're generating?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; Yeah, exactly — 2D sprites. The animation looks amazing. And it can also figure out hitboxes — it can be like, "Oh, your fist is here, I'll draw the JSON hitbox." It's incredible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Mine is much more simple. I'm a big rock climber and a lot of my friends climb, so we have this little app we built with Claude Code where we log all the projects we're working on. We also go outdoors together a lot, so we have Claude do all this research with workflows. Workflows is amazing — we brand it as a coding tool, but it's amazing for doing deep research for travel. I also plan our team offsites, and it's good at finding venues that can fit all of us. I use workflows to research all the climbing destinations we might want to go to, and what has direct flights from where all of us are located. It goes to Mountain Project and finds all the climbs at our grade level. It finds the Airbnb. And I don't like hiking, so I care a lot about it having a very short approach — &lt;strong&gt;very short walking distance from where the car parks to where the rock actually is&lt;/strong&gt; — and it filters for this. With existing apps I have to manually click through Mountain Project, but with this I just put in all of our preferences and it's a custom app for us.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Simon:&lt;/strong&gt; So you're basically vibe coding Jira for mountain climbing.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; Exactly.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="audience-any-plans-for-eval-building-tools-and-agent-observability-"&gt;Audience: Any plans for eval-building tools and agent observability?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=2963s"&gt;49:23&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;We had a few minutes at the end for questions from the audience.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Audience:&lt;/strong&gt; Do you have any near-term plans to build more eval tools for us to build eval datasets, and more observability tools to monitor the performance of agents and workflows?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cat:&lt;/strong&gt; We've considered building eval tools, but I think the limiting factor actually tends to be that &lt;strong&gt;it takes a long time for customers to build really high-quality evals&lt;/strong&gt;. So I think the tooling is less of the constraint, and more the skill set of how you build a great eval. That's an area where we're excited to both invest internally and hopefully share some best practices externally.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="audience-how-is-memory-designed-today-and-would-you-move-from-files-to-a-data-store-"&gt;Audience: How is memory designed today — and would you move from files to a data store?&lt;/h4&gt;
&lt;p&gt;&lt;a href="https://www.youtube.com/watch?v=uU5Gv2h8-9g&amp;amp;t=3008s"&gt;50:08&lt;/a&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Audience (Sai):&lt;/strong&gt; I'm interested in the memory and the multiplayer. &lt;strong&gt;How is memory being designed today?&lt;/strong&gt; I assume it's around files. And second, have you thought about an orthogonal direction where you &lt;strong&gt;would actually need a data store for these memories, instead of files, to scale it better?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thariq:&lt;/strong&gt; Right now for Claude Tag the memory is channel-specific. Every Claude in that channel has a shared memory, and the instances have a session — but the session can contribute back to main memory. We do a lot of memory research, and it can be kind of unintuitive what the right way to do memory is. We're always running memory experiments. &lt;strong&gt;How it works right now in Claude Tag is a markdown file per channel.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="prompt-engineering"/><category term="generative-ai"/><category term="llms"/><category term="anthropic"/><category term="annotated-talks"/><category term="coding-agents"/><category term="claude-code"/><category term="thariq-shihipar"/><category term="cat-wu"/></entry><entry><title>Kimi K3, and what we can still learn from the pelican benchmark</title><link href="https://simonwillison.net/2026/Jul/16/kimi-k3/" rel="alternate"/><published>2026-07-16T20:19:30+00:00</published><updated>2026-07-16T20:19:30+00:00</updated><id>https://simonwillison.net/2026/Jul/16/kimi-k3/</id><summary type="html">&lt;p&gt;Chinese AI lab Moonshot AI &lt;a href="https://www.kimi.com/blog/kimi-k3"&gt;announced Kimi K3&lt;/a&gt; this morning, describing it as their "most capable model to date, with 2.8 trillion parameters". It's currently available via their website and API, but an open weight release is promised "by July 27, 2026".&lt;/p&gt;
&lt;p&gt;Moonshot are calling this the first "open 3T-class model" (I guess they're rounding 2.8 trillion up to 3 trillion), taking the crown from &lt;a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro"&gt;DeepSeek's 1.6T v4 Pro&lt;/a&gt;. Their &lt;a href="https://www.kimi.com/blog/kimi-k3#full-benchmark-table"&gt;self-reported benchmarks&lt;/a&gt; have K3 mostly beating Claude Opus 4.8 max and GPT-5.5 high, while losing out to Claude Fable 5 and GPT-5.6 Sol.&lt;/p&gt;
&lt;p&gt;A few highlights from the &lt;a href="https://twitter.com/ArtificialAnlys/status/2077832874183860404"&gt;Artificial Analysis report&lt;/a&gt; on the model:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;"On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5."&lt;/li&gt;
&lt;li&gt;"Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers"&lt;/li&gt;
&lt;li&gt;"Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significantly, using 21% fewer output tokens than K2.6."&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The model is also now the &lt;a href="https://twitter.com/arena/status/2077824029126504525"&gt;leading model on Arena.ai's Frontend Code arena&lt;/a&gt;, surpassing even Claude Fable 5.&lt;/p&gt;
&lt;p&gt;The new model is notable for the pricing: $3/million input tokens and $15/million output tokens, putting it at the same level as Anthropic's Claude Sonnet series and making it the most expensive model released by a Chinese AI lab to date. This is a significant increase on their earlier models &lt;a href="https://platform.kimi.ai/docs/pricing/chat-k26"&gt;such as Kimi K2.6&lt;/a&gt; at $0.95/$4. 2.8 trillion parameters is also more than twice the size of that 1T model.&lt;/p&gt;
&lt;h4 id="but-how-does-it-pelican-"&gt;But how does it pelican?&lt;/h4&gt;
&lt;p&gt;I used OpenRouter (to avoid signing up for a Moonshot API key) with the &lt;a href="https://github.com/simonw/llm-openrouter"&gt;llm-openrouter plugin&lt;/a&gt; to generate an SVG of a pelican riding a bicycle:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;llm -m openrouter/moonshotai/kimi-k3 'Generate an SVG of a pelican riding a bicycle'
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Here's &lt;a href="https://gist.github.com/simonw/66a2699eb1594258904c7b5102840dd6"&gt;the transcript&lt;/a&gt;. It looks like this:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/kimi-3-pelican.jpg" alt="See description below" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;That pelican took 95 input tokens and 16,658 output tokens (13,241 were reasoning tokens), for a total cost of &lt;a href="https://www.llm-prices.com/#it=95&amp;amp;ot=16658&amp;amp;ic=3&amp;amp;oc=15"&gt;25 cents&lt;/a&gt;!&lt;/p&gt;
&lt;p&gt;Since K3 accepts image input I ran it against that rendered SVG above (with my &lt;a href="https://simonwillison.net/guides/agentic-engineering-patterns/prompts/#alt-text"&gt;alt text prompt&lt;/a&gt;) and &lt;a href="https://gist.github.com/simonw/665dbf840701b421745f2cb891acdfd6"&gt;got back&lt;/a&gt; (for &lt;a href="https://www.llm-prices.com/#it=822&amp;amp;ot=243&amp;amp;ic=3&amp;amp;oc=15"&gt;0.6 cents&lt;/a&gt;):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Cartoon illustration of a white pelican wearing a red scarf, riding a red bicycle along a gray road with white dashed lines; the pelican has a large orange beak and webbed orange feet pedaling, with white motion lines behind it; the background shows a light blue sky with white clouds, a yellow sun, two small black birds in flight, and green grass with tiny white flowers in the foreground&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="what-can-we-learn-from-the-pelican-"&gt;What can we learn from the pelican?&lt;/h4&gt;
&lt;p&gt;My &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/"&gt;Generate an SVG of a pelican riding a bicycle&lt;/a&gt; test is 21 months old now. It was never a particularly great benchmark. It started out as a joke on how absurdly difficult it is to compare these models, but then for the first year it turned out to have a &lt;a href="https://simonwillison.net/2025/Jun/6/six-months-in-llms/"&gt;surprising correlation&lt;/a&gt; to how good the models actually were.&lt;/p&gt;
&lt;p&gt;That connection has been mostly severed now. The &lt;a href="https://simonwillison.net/2026/Jul/9/gpt-5-6/"&gt;GPT-5.6&lt;/a&gt; and &lt;a href="https://simonwillison.net/2026/Jun/9/claude-fable-5/"&gt;Claude Fable 5&lt;/a&gt;  pelicans are outclassed &lt;a href="https://simonwillison.net/2026/Jun/17/glm-52/"&gt;by GLM-5.2&lt;/a&gt;, and much as I love GLM I don't think that's a Fable-class model.&lt;/p&gt;
&lt;p&gt;(I'm still not convinced that labs are &lt;a href="https://simonwillison.net/2025/Nov/13/training-for-pelicans-riding-bicycles/"&gt;training for the benchmark&lt;/a&gt; - if they were, I'd expect much better results. There's a chance that Gemini has optimized for &lt;a href="https://simonwillison.net/2026/Feb/19/gemini-31-pro/#jeff-dean"&gt;any combination of an animal on a vehicle&lt;/a&gt; though!)&lt;/p&gt;
&lt;p&gt;The biggest limitation of the pelican is that it doesn't touch at all on the thing that matters most for today's model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.&lt;/p&gt;
&lt;p&gt;So don't go using pelicans to compare models!&lt;/p&gt;

&lt;p&gt;All of that said, I still get a decent amount of value out of running the benchmark myself.&lt;/p&gt;
&lt;p&gt;Firstly, it's a forcing function for actually trying the model. If I show you a pelican, that means I've managed to run a prompt through it. If the model has an official API I'll use that, if it's open weight (and small enough to fit a 128GB M5 MacBook Pro) I'll try running it on my own machine, usually via &lt;a href="https://github.com/ggml-org/llama.cpp"&gt;llama.cpp&lt;/a&gt; or &lt;a href="https://lmstudio.ai"&gt;LM Studio&lt;/a&gt; or &lt;a href="https://ollama.com"&gt;Ollama&lt;/a&gt;. I'll frequently use &lt;a href="https://openrouter.ai"&gt;OpenRouter&lt;/a&gt; since that usually provides a proxy to an official API without me needing a new API key.&lt;/p&gt;
&lt;p&gt;Most of my pelicans are generated using &lt;a href="https://llm.datasette.io/"&gt;my LLM CLI tool&lt;/a&gt;, which helps encourage me to ensure the latest models are supported by that (via one of its plugins).&lt;/p&gt;
&lt;p&gt;More importantly though, even the act of a single prompt to "Generate an SVG of a pelican riding a bicycle" can reveal interesting model characteristics.&lt;/p&gt;
&lt;p&gt;Consider &lt;a href="https://gist.github.com/simonw/66a2699eb1594258904c7b5102840dd6"&gt;the result&lt;/a&gt; for Kimi K3 today. Running those simple prompts helped emphasize several points about the model.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;It only has one reasoning effort right now, "max" - and it shows. The model consumed 13,241 reasoning tokens to output 3,417 tokens of response. This is expensive - the pelican cost 25 cents!&lt;/li&gt;
&lt;li&gt;How does the prompt "Generate an SVG of a pelican riding a bicycle" add up to 95 input tokens?  OpenAI's &lt;a href="https://platform.openai.com/tokenizer"&gt;tokenizer&lt;/a&gt;  counts 10, &lt;a href="https://tools.simonwillison.net/claude-token-counter"&gt;Anthropic's&lt;/a&gt; counts 10 for Opus 4.6, 30 for Opus 4.7 and 25 for Sonnet 5/Fable 5. Prompting "hi" &lt;a href="https://news.ycombinator.com/item?id=48935342#48936461"&gt;to Kimi K3&lt;/a&gt; counted 86 tokens, suggesting there may be an 85 token hidden system prompt. It &lt;a href="https://news.ycombinator.com/item?id=48935342#48936515"&gt;refused to leak it&lt;/a&gt; though.&lt;/li&gt;
&lt;li&gt;Vision works well: the alt text it generated is very good.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;K3 currently only has one thinking effort level, but I've been deriving quite a bit of value recently from running the same pelican prompt through different effort levels to get a quick idea for what impact those have. Here's my matrix &lt;a href="https://static.simonwillison.net/static/2026/gpt-5.6-pelicans.html"&gt;for the GPT-5.6 model family&lt;/a&gt;, for example.&lt;/p&gt;
&lt;p&gt;Really though the main things I gain from the pelican test are:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;It's a "hello world" exercise for prompting a model&lt;/li&gt;
&lt;li&gt;A rough cost and reasoning estimate for a simple task&lt;/li&gt;
&lt;li&gt;Confirmation that the model can output valid SVG and has a basic idea of geometry and spatial awareness. This is a much bigger deal for the smaller models that run on my laptop.&lt;/li&gt;
&lt;li&gt;It's still interesting to compare pelicans between releases in the same model family. K3's pelican is a notable improvement from &lt;a href="https://simonwillison.net/2026/Jan/27/kimi-k25/"&gt;Kimi 2.5&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;It's something I can share that demonstrates I've tried it. Plus a comment with a pelican in it is kind of a tradition on Hacker News at this point, any time I'm late I get comments asking where it is!&lt;/li&gt;
&lt;/ol&gt;&lt;p&gt;&lt;em&gt;You are only seeing the long-form articles from my blog. Subscribe to &lt;a href="https://simonwillison.net/atom/everything/"&gt;/atom/everything/&lt;/a&gt; to get all of my posts, or take a look at my &lt;a href="https://simonwillison.net/about/#subscribe"&gt;other subscription options&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;</summary><category term="ai"/><category term="generative-ai"/><category term="llms"/><category term="llm-pricing"/><category term="pelican-riding-a-bicycle"/><category term="llm-release"/><category term="ai-in-china"/><category term="artificial-analysis"/><category term="moonshot"/><category term="kimi"/></entry></feed>