<?xml version="1.0" encoding="utf-8"?>
<feed xml:lang="en-us" xmlns="http://www.w3.org/2005/Atom"><title>Simon Willison's Weblog: chatgpt</title><link href="http://simonwillison.net/" rel="alternate"/><link href="http://simonwillison.net/tags/chatgpt.atom" rel="self"/><id>http://simonwillison.net/</id><updated>2026-07-29T00:13:18+00:00</updated><author><name>Simon Willison</name></author><entry><title>Adding a custom MCP server to Claude and ChatGPT</title><link href="https://simonwillison.net/2026/Jul/29/mcp-in-claude-and-chatgpt/#atom-tag" rel="alternate"/><published>2026-07-29T00:13:18+00:00</published><updated>2026-07-29T00:13:18+00:00</updated><id>https://simonwillison.net/2026/Jul/29/mcp-in-claude-and-chatgpt/#atom-tag</id><summary type="html">
    
        &lt;p&gt;&lt;strong&gt;TIL:&lt;/strong&gt; &lt;a href="https://til.simonwillison.net/llms/mcp-in-claude-and-chatgpt"&gt;Adding a custom MCP server to Claude and ChatGPT&lt;/a&gt;&lt;/p&gt;
        &lt;p&gt;Connecting a custom MCP server to Claude and ChatGPT's standard chat interfaces is possible, but can take quite a few steps.&lt;/p&gt;
    
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/claude"&gt;claude&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/model-context-protocol"&gt;model-context-protocol&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="ai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="claude"/><category term="model-context-protocol"/></entry><entry><title>Quoting OpenAI</title><link href="https://simonwillison.net/2026/Jul/10/openai/#atom-tag" rel="alternate"/><published>2026-07-10T01:05:57+00:00</published><updated>2026-07-10T01:05:57+00:00</updated><id>https://simonwillison.net/2026/Jul/10/openai/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://help.openai.com/en/articles/20001275-chatgpt-work-and-codex"&gt;&lt;p&gt;[...] Work on web and mobile runs in the cloud. Work in the desktop app can also use local files and desktop apps with your permission. At launch, cloud Work conversations do not appear in desktop Work; desktop Work threads and local files remain on that computer.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://help.openai.com/en/articles/20001275-chatgpt-work-and-codex"&gt;OpenAI&lt;/a&gt;, trying (unsuccessfully) to clarify ChatGPT Work&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="chatgpt"/></entry><entry><title>WHY ARE YOU LIKE THIS</title><link href="https://simonwillison.net/2026/Apr/25/why-are-you-like-this/#atom-tag" rel="alternate"/><published>2026-04-25T16:44:01+00:00</published><updated>2026-04-25T16:44:01+00:00</updated><id>https://simonwillison.net/2026/Apr/25/why-are-you-like-this/#atom-tag</id><summary type="html">
    &lt;p&gt;@scottjla &lt;a href="https://twitter.com/scottjla/status/2047535371664457863"&gt;on Twitter&lt;/a&gt; in reply to my &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/"&gt;pelican riding a bicycle&lt;/a&gt; benchmark:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I feel like we need to stack these tests now&lt;/p&gt;
&lt;p&gt;&lt;img alt="AI generated image. A pelican is riding a bicycle along a dirt track, chased by a police car. The pelican looks panicked, likely because there is an astronaut (with prehensile toes for some reason) riding the pelican clinging on to where its ears should be. The astronaut is being ridden by a horse, with an equally wild expression. A slice of pizza and a can and a cowboy hat are falling next to them. A road sign in the background reads WHY ARE YOU LIKE THIS." src="https://static.simonwillison.net/static/2026/why-are-you-like-this.jpeg" /&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I checked to confirm that the model (ChatGPT Images 2.0) added the "WHY ARE YOU LIKE THIS" sign of its own accord and &lt;a href="https://chatgpt.com/share/69ebff27-2220-839f-b065-8c3516ea9b6d"&gt;it did&lt;/a&gt; - the prompt Scott used was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Create an image of a horse riding an astronaut, where the astronaut is riding a pelican that is riding a bicycle. It looks very chaotic but they all just manage to balance on top of each other&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/slop"&gt;slop&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/text-to-image"&gt;text-to-image&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle"&gt;pelican-riding-a-bicycle&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="generative-ai"/><category term="chatgpt"/><category term="slop"/><category term="text-to-image"/><category term="pelican-riding-a-bicycle"/></entry><entry><title>A pelican for GPT-5.5 via the semi-official Codex backdoor API</title><link href="https://simonwillison.net/2026/Apr/23/gpt-5-5/#atom-tag" rel="alternate"/><published>2026-04-23T19:59:47+00:00</published><updated>2026-04-23T19:59:47+00:00</updated><id>https://simonwillison.net/2026/Apr/23/gpt-5-5/#atom-tag</id><summary type="html">
    &lt;p&gt;&lt;a href="https://openai.com/index/introducing-gpt-5-5/"&gt;GPT-5.5 is out&lt;/a&gt;. It's available in OpenAI Codex and is rolling out to paid ChatGPT subscribers. I've had some preview access and found it to be a fast, effective and highly capable model. As is usually the case these days, it's hard to put into words what's good about it - I ask it to build things and it builds exactly what I ask for!&lt;/p&gt;
&lt;p&gt;There's one notable omission from today's release - the API:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;API deployments require different safeguards and we are working closely with partners and customers on the safety and security requirements for serving it at scale. We'll bring GPT‑5.5 and GPT‑5.5 Pro to the API very soon.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;When I run my &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/"&gt;pelican benchmark&lt;/a&gt; I always prefer to use an API, to avoid hidden system prompts in ChatGPT or other agent harnesses from impacting the results.&lt;/p&gt;
&lt;h4 id="the-openclaw-backdoor"&gt;The OpenClaw backdoor&lt;/h4&gt;
&lt;p&gt;One of the ongoing tension points in the AI world over the past few months has concerned how agent harnesses like OpenClaw and Pi interact with the APIs provided by the big providers.&lt;/p&gt;
&lt;p&gt;Both OpenAI and Anthropic offer popular monthly subscriptions which provide access to their models at a significant discount to their raw API.&lt;/p&gt;
&lt;p&gt;OpenClaw integrated directly with this mechanism, and was then &lt;a href="https://www.theverge.com/ai-artificial-intelligence/907074/anthropic-openclaw-claude-subscription-ban"&gt;blocked from doing so&lt;/a&gt; by Anthropic. This kicked off a whole thing. OpenAI - who recently hired OpenClaw creator Peter Steinberger - saw an opportunity for an easy karma win and announced that OpenClaw was welcome to continue integrating with OpenAI's subscriptions via the same mechanism used by their (open source) Codex CLI tool.&lt;/p&gt;
&lt;p&gt;Does this mean &lt;em&gt;anyone&lt;/em&gt; can write code that integrates with OpenAI's Codex-specific APIs to hook into those existing subscriptions?&lt;/p&gt;
&lt;p&gt;The other day &lt;a href="https://twitter.com/jeremyphoward/status/2046537816834965714"&gt;Jeremy Howard asked&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Anyone know whether OpenAI officially supports the use of the &lt;code&gt;/backend-api/codex/responses&lt;/code&gt; endpoint that Pi and Opencode (IIUC) uses?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It turned out that on March 30th OpenAI's Romain Huet &lt;a href="https://twitter.com/romainhuet/status/2038699202834841962"&gt;had tweeted&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We want people to be able to use Codex, and their ChatGPT subscription, wherever they like! That means in the app, in the terminal, but also in JetBrains, Xcode, OpenCode, Pi, and now Claude Code.&lt;/p&gt;
&lt;p&gt;That’s why Codex CLI and Codex app server are open source too! 🙂&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And Peter Steinberger &lt;a href="https://twitter.com/steipete/status/2046775849769148838"&gt;replied to Jeremy&lt;/a&gt; that:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;OpenAI sub is officially supported.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="llm-openai-via-codex"&gt;llm-openai-via-codex&lt;/h4&gt;
&lt;p&gt;So... I had Claude Code reverse-engineer the &lt;a href="https://github.com/openai/codex"&gt;openai/codex&lt;/a&gt; repo, figure out how authentication tokens were stored and build me &lt;a href="https://github.com/simonw/llm-openai-via-codex"&gt;llm-openai-via-codex&lt;/a&gt;, a new plugin for &lt;a href="https://llm.datasette.io/"&gt;LLM&lt;/a&gt; which picks up your existing Codex subscription and uses it to run prompts!&lt;/p&gt;
&lt;p&gt;(With hindsight I wish I'd used GPT-5.4 or the GPT-5.5 preview, it would have been funnier. I genuinely considered rewriting the project from scratch using Codex and GPT-5.5 for the sake of the joke, but decided not to spend any more time on this!)&lt;/p&gt;
&lt;p&gt;Here's how to use it:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Install Codex CLI, buy an OpenAI plan, login to Codex&lt;/li&gt;
&lt;li&gt;Install LLM: &lt;code&gt;uv tool install llm&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Install the new plugin: &lt;code&gt;llm install llm-openai-via-codex&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Start prompting: &lt;code&gt;llm -m openai-codex/gpt-5.5 'Your prompt goes here'&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;All existing LLM features should also work - use &lt;code&gt;-a filepath.jpg/URL&lt;/code&gt; to attach an image, &lt;code&gt;llm chat -m openai-codex/gpt-5.5&lt;/code&gt; to start an ongoing chat, &lt;code&gt;llm logs&lt;/code&gt; to view logged conversations and &lt;code&gt;llm --tool ...&lt;/code&gt; to &lt;a href="https://llm.datasette.io/en/stable/tools.html"&gt;try it out with tool support&lt;/a&gt;.&lt;/p&gt;
&lt;h4 id="and-some-pelicans"&gt;And some pelicans&lt;/h4&gt;
&lt;p&gt;Let's generate a pelican!&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;llm install llm-openai-via-codex
llm -m openai-codex/gpt-5.5 &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;Generate an SVG of a pelican riding a bicycle&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Here's &lt;a href="https://gist.github.com/simonw/edda1d98f7ba07fd95eeff473cb16634"&gt;what I got back&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/gpt-5.5-pelican.png" alt="It is a bit mangled to be honest - good beak, pelican body shapes are slightly weird, legs do at least extend to the pedals, bicycle frame is not quite right." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;I've seen better &lt;a href="https://simonwillison.net/2026/Mar/17/mini-and-nano/#pelicans"&gt;from GPT-5.4&lt;/a&gt;, so I tagged on &lt;code&gt;-o reasoning_effort xhigh&lt;/code&gt; and &lt;a href="https://gist.github.com/simonw/a6168e4165a258e4d664aeae8e602cc5"&gt;tried again&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;That one took almost four minutes to generate, but I think it's a much better effort.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/gpt-5.5-pelican-xhigh.png" alt="Pelican has gradients now, body is much better put together, bicycle is nearly the right shape albeit with one extra bar between pedals and front wheel, clearly a better image overall." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;If you compare the SVG code (&lt;a href="https://gist.github.com/simonw/edda1d98f7ba07fd95eeff473cb16634#response"&gt;default&lt;/a&gt;, &lt;a href="https://gist.github.com/simonw/a6168e4165a258e4d664aeae8e602cc5#response"&gt;xhigh&lt;/a&gt;) the &lt;code&gt;xhigh&lt;/code&gt; one took a very different approach, which is much more CSS-heavy - as demonstrated by those gradients. &lt;code&gt;xhigh&lt;/code&gt; used 9,322 reasoning tokens where the default used just 39.&lt;/p&gt;
&lt;h4 id="a-few-more-notes-on-gpt-5-5"&gt;A few more notes on GPT-5.5&lt;/h4&gt;
&lt;p&gt;One of the most notable things about GPT-5.5 is the pricing. Once it goes live in the API it's &lt;a href="https://openai.com/index/introducing-gpt-5-5/#availability-and-pricing"&gt;going to be priced&lt;/a&gt; at &lt;em&gt;twice&lt;/em&gt; the cost of GPT-5.4 - $5 per 1M input tokens and $30 per 1M output tokens, where 5.4 is $2.5 and $15.&lt;/p&gt;
&lt;p&gt;GPT-5.5 Pro will be even more: $30 per 1M input tokens and $180 per 1M output tokens.&lt;/p&gt;
&lt;p&gt;GPT-5.4 will remain available. At half the price of 5.5 this feels like 5.4 is to 5.5 as Claude Sonnet is to Claude Opus.&lt;/p&gt;
&lt;p&gt;Ethan Mollick has a &lt;a href="https://www.oneusefulthing.org/p/sign-of-the-future-gpt-55"&gt;detailed review of GPT-5.5&lt;/a&gt; where he put it (and GPT-5.5 Pro) through an array of interesting challenges. His verdict: the jagged frontier continues to hold, with GPT-5.5 excellent at some things and challenged by others in a way that remains difficult to predict.&lt;/p&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm"&gt;llm&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-pricing"&gt;llm-pricing&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle"&gt;pelican-riding-a-bicycle&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-reasoning"&gt;llm-reasoning&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-release"&gt;llm-release&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/codex"&gt;codex&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt"&gt;gpt&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="llm"/><category term="llm-pricing"/><category term="pelican-riding-a-bicycle"/><category term="llm-reasoning"/><category term="llm-release"/><category term="codex"/><category term="gpt"/></entry><entry><title>Where's the raccoon with the ham radio? (ChatGPT Images 2.0)</title><link href="https://simonwillison.net/2026/Apr/21/gpt-image-2/#atom-tag" rel="alternate"/><published>2026-04-21T20:32:24+00:00</published><updated>2026-04-21T20:32:24+00:00</updated><id>https://simonwillison.net/2026/Apr/21/gpt-image-2/#atom-tag</id><summary type="html">
    &lt;p&gt;OpenAI &lt;a href="https://openai.com/index/introducing-chatgpt-images-2-0/"&gt;released ChatGPT Images 2.0 today&lt;/a&gt;, their latest image generation model. On &lt;a href="https://www.youtube.com/watch?v=sWkGomJ3TLI"&gt;the livestream&lt;/a&gt; Sam Altman said that the leap from gpt-image-1 to gpt-image-2 was equivalent to jumping from GPT-3 to GPT-5. Here's how I put it to the test.&lt;/p&gt;
&lt;p&gt;My prompt:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Do a where's Waldo style image but it's where is the raccoon holding a ham radio&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="gpt-image-1"&gt;gpt-image-1&lt;/h4&gt;
&lt;p&gt;First as a baseline here's what I got from the older gpt-image-1 using ChatGPT directly:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://static.simonwillison.net/static/2026/chatgpt-image-1-ham-radio.png"&gt;&lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/image_crop_1402x1122_w1402_q0.3.jpg" alt="There's a lot going on, but I couldn't find a raccoon." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I wasn't able to spot the raccoon - I quickly realized that testing image generation models on Where's Waldo style images (Where's Wally in the UK) can be pretty frustrating!&lt;/p&gt;
&lt;p&gt;I tried &lt;a href="https://claude.ai/share/bd6e9b88-29a9-420b-8ac1-3ac5cebac215"&gt;getting Claude Opus 4.7&lt;/a&gt; with its new higher resolution inputs to solve it but it was convinced there was a raccoon it couldn't find thanks to the instruction card at the top left of the image:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Yes — there's at least one raccoon in the picture, but it's very well hidden&lt;/strong&gt;. In my careful sweep through zoomed-in sections, honestly, I couldn't definitively spot a raccoon holding a ham radio. [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="nano-banana-2-and-pro"&gt;Nano Banana 2 and Pro&lt;/h4&gt;
&lt;p&gt;Next I tried Google's Nano Banana 2, &lt;a href="https://gemini.google.com/share/3775db96c576"&gt;via Gemini&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://static.simonwillison.net/static/2026/nano-banana-2-ham-radio.jpg"&gt;&lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/gemini-ham-radio-small.jpg" alt="Busy Where's Waldo-style illustration of a park festival with crowds of people, tents labeled &amp;quot;FOOD &amp;amp; DRINK&amp;quot;, &amp;quot;CRAFT FAIR&amp;quot;, &amp;quot;BOOK NOOK&amp;quot;, &amp;quot;MUSIC FEST&amp;quot;, and &amp;quot;AMATEUR RADIO CLUB - W6HAM&amp;quot; (featuring a raccoon in a red hat at the radio table), plus a Ferris wheel, carousel, gazebo with band, pond with boats, fountain, food trucks, and striped circus tents" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;That one was pretty obvious, the raccoon is in the "Amateur Radio Club" booth in the center of the image!&lt;/p&gt;
&lt;p&gt;Claude said:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Honestly, this one wasn't really hiding — he's the star of the booth. Feels like the illustrator took pity on us after that last impossible scene. The little "W6HAM" callsign pun on the booth sign is a nice touch too.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I also tried Nano Banana Pro &lt;a href="https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221sGU5A7mrngkfLfSEU84xaV1DhtOTnS--%22%5D,%22action%22:%22open%22,%22userId%22:%22106366615678321494423%22,%22resourceKeys%22:%7B%7D%7D&amp;amp;usp=sharing"&gt;in AI Studio&lt;/a&gt; and got this, by far the worst result from any model. Not sure what went wrong here!&lt;/p&gt;
&lt;p&gt;&lt;a href="https://static.simonwillison.net/static/2026/nano-banana-pro-ham-radio.jpg"&gt;&lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/nano-banana-pro-ham-radio-small.jpg" alt="The raccoon is larger than everyone else, right in the middle of the image with an ugly white border around it." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h4 id="gpt-image-2"&gt;gpt-image-2&lt;/h4&gt;
&lt;p&gt;With the baseline established, let's try out the new model.&lt;/p&gt;
&lt;p&gt;I used an updated version of my &lt;a href="https://github.com/simonw/tools/blob/main/python/openai_image.py"&gt;openai_image.py&lt;/a&gt; script, which is a thin wrapper around the &lt;a href="https://github.com/openai/openai-python"&gt;OpenAI Python&lt;/a&gt; client library. Their client library hasn't yet been updated to include &lt;code&gt;gpt-image-2&lt;/code&gt; but thankfully it doesn't validate the model ID so you can use it anyway.&lt;/p&gt;
&lt;p&gt;Here's how I ran that:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;OPENAI_API_KEY=&lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;$(&lt;/span&gt;llm keys get openai&lt;span class="pl-pds"&gt;)&lt;/span&gt;&lt;/span&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt; \
  uv run https://tools.simonwillison.net/python/openai_image.py \
  -m gpt-image-2 \
  &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;Do a where's Waldo style image but it's where is the raccoon holding a ham radio&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Here's what I got back. I don't &lt;em&gt;think&lt;/em&gt; there's a raccoon in there - I couldn't spot one, and neither could Claude.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://static.simonwillison.net/static/2026/gpt-image-2-default.png"&gt;&lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/gpt-image-2-default.jpg" alt="Lots of stuff, a ham radio booth, many many people, a lake, but maybe no raccoon?" style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://github.com/openai/openai-cookbook/blob/main/examples/multimodal/image-gen-models-prompting-guide.ipynb"&gt;OpenAI image generation cookbook&lt;/a&gt; has been updated with notes on &lt;code&gt;gpt-image-2&lt;/code&gt;, including the &lt;code&gt;outputQuality&lt;/code&gt; setting and available sizes.&lt;/p&gt;
&lt;p&gt;I tried setting &lt;code&gt;outputQuality&lt;/code&gt; to &lt;code&gt;high&lt;/code&gt; and the dimensions to &lt;code&gt;3840x2160&lt;/code&gt; - I believe that's the maximum - and got this - a 17MB PNG which I converted to a 5MB WEBP:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;OPENAI_API_KEY=&lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;$(&lt;/span&gt;llm keys get openai&lt;span class="pl-pds"&gt;)&lt;/span&gt;&lt;/span&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt; \
  uv run &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;'&lt;/span&gt;https://raw.githubusercontent.com/simonw/tools/refs/heads/main/python/openai_image.py&lt;span class="pl-pds"&gt;'&lt;/span&gt;&lt;/span&gt; \
  -m gpt-image-2 &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;Do a where's Waldo style image but it's where is the raccoon holding a ham radio&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt; \
  --quality high --size 3840x2160&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;a href="https://static.simonwillison.net/static/2026/image-fc93bd-q100.webp"&gt;&lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/image-fc93bd-q100.jpg" alt="Big complex image, lots of detail, good wording, there is indeed a raccoon with a ham radio." style="max-width: 100%;" /&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;That's pretty great! There's a raccoon with a ham radio in there (bottom left, quite easy to spot).&lt;/p&gt;
&lt;p&gt;The image used 13,342 output tokens, which are charged at $30/million so a total cost of around &lt;a href="https://www.llm-prices.com/#ot=13342&amp;amp;ic=5&amp;amp;cic=1.25&amp;amp;oc=10&amp;amp;sel=gpt-image-2-image"&gt;40 cents&lt;/a&gt;.&lt;/p&gt;
&lt;h4 id="takeaways"&gt;Takeaways&lt;/h4&gt;
&lt;p&gt;I think this new ChatGPT image generation model takes the crown from Gemini, at least for the moment.&lt;/p&gt;
&lt;p&gt;Where's Waldo style images are an infuriating and somewhat foolish way to test these models, but they do help illustrate how good they are getting at complex illustrations combining both text and details.&lt;/p&gt;
&lt;h4 id="update-asking-models-to-solve-this-is-risky"&gt;Update: asking models to solve this is risky&lt;/h4&gt;
&lt;p&gt;rizaco &lt;a href="https://news.ycombinator.com/item?id=47852835#47853561"&gt;on Hacker News&lt;/a&gt; asked ChatGPT to draw a red circle around the raccoon in one of the images in which I had failed to find one. Here's an animated mix of their result and the original image:&lt;/p&gt;
&lt;p&gt;&lt;img loading="lazy" src="https://static.simonwillison.net/static/2026/ham-radio-cheat.gif" alt="The circle appears around a raccoon with a ham radio who is definitely not there in the original image!" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Looks like we definitely can't trust these models to usefully solve their own puzzles!&lt;/p&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/text-to-image"&gt;text-to-image&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-release"&gt;llm-release&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/nano-banana"&gt;nano-banana&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="text-to-image"/><category term="llm-release"/><category term="nano-banana"/></entry><entry><title>ChatGPT voice mode is a weaker model</title><link href="https://simonwillison.net/2026/Apr/10/voice-mode-is-weaker/#atom-tag" rel="alternate"/><published>2026-04-10T15:56:02+00:00</published><updated>2026-04-10T15:56:02+00:00</updated><id>https://simonwillison.net/2026/Apr/10/voice-mode-is-weaker/#atom-tag</id><summary type="html">
    &lt;p&gt;I think it's non-obvious to many people that the OpenAI voice mode runs on a much older, much weaker model - it feels like the AI that you can talk to should be the smartest AI but it really isn't.&lt;/p&gt;
&lt;p&gt;If you ask ChatGPT voice mode for its knowledge cutoff date it tells you April 2024 - it's a GPT-4o era model.&lt;/p&gt;
&lt;p&gt;This thought inspired by &lt;a href="https://twitter.com/karpathy/status/2042334451611693415"&gt;this Andrej Karpathy tweet&lt;/a&gt; about the growing gap in understanding of AI capability based on the access points and domains people are using the models with:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[...] It really is simultaneously the case that OpenAI's free and I think slightly orphaned (?) "Advanced Voice Mode" will fumble the dumbest questions in your Instagram's reels and &lt;em&gt;at the same time&lt;/em&gt;, OpenAI's highest-tier and paid Codex model will go off for 1 hour to coherently restructure an entire code base, or find and exploit vulnerabilities in computer systems.&lt;/p&gt;
&lt;p&gt;This part really works and has made dramatic strides because 2 properties:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;these domains offer explicit reward functions that are verifiable meaning they are easily amenable to reinforcement learning training (e.g. unit tests passed yes or no, in contrast to writing, which is much harder to explicitly judge),  but also&lt;/li&gt;
&lt;li&gt;they are a lot more valuable in b2b settings, meaning that the biggest fraction of the team is focused on improving them.&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/andrej-karpathy"&gt;andrej-karpathy&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="andrej-karpathy"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/></entry><entry><title>Quoting Chengpeng Mou</title><link href="https://simonwillison.net/2026/Apr/5/chengpeng-mou/#atom-tag" rel="alternate"/><published>2026-04-05T21:47:06+00:00</published><updated>2026-04-05T21:47:06+00:00</updated><id>https://simonwillison.net/2026/Apr/5/chengpeng-mou/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://twitter.com/cpmou2022/status/2040606209800290404"&gt;&lt;p&gt;From anonymized U.S. ChatGPT data, we are seeing:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;~2M weekly messages on health insurance&lt;/li&gt;
&lt;li&gt;~600K weekly messages [classified as healthcare] from people living in “hospital deserts” (30 min drive to nearest hospital)&lt;/li&gt;
&lt;li&gt;7 out of 10 msgs happen outside clinic hours&lt;/li&gt;
&lt;/ul&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://twitter.com/cpmou2022/status/2040606209800290404"&gt;Chengpeng Mou&lt;/a&gt;, Head of Business Finance, OpenAI&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-ethics"&gt;ai-ethics&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="ai-ethics"/></entry><entry><title>Quoting Benedict Evans</title><link href="https://simonwillison.net/2026/Feb/26/benedict-evans/#atom-tag" rel="alternate"/><published>2026-02-26T03:44:56+00:00</published><updated>2026-02-26T03:44:56+00:00</updated><id>https://simonwillison.net/2026/Feb/26/benedict-evans/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://www.ben-evans.com/benedictevans/2026/2/19/how-will-openai-compete-nkg2x"&gt;&lt;p&gt;If people are only using this a couple of times a week at most, and can’t think of anything to do with it on the average day, it hasn’t changed their life. OpenAI itself admits the problem, talking about a ‘capability gap’ between what the models can do and what people do with them, which seems to me like a way to avoid saying that you don’t have clear product-market fit. &lt;/p&gt;
&lt;p&gt;Hence, OpenAI’s ad project is partly just about covering the cost of serving the 90% or more of users who don’t pay (and capturing an early lead with advertisers and early learning in how this might work), but more strategically, it’s also about making it possible to give those users the latest and most powerful (i.e. expensive) models, in the hope that this will deepen their engagement.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://www.ben-evans.com/benedictevans/2026/2/19/how-will-openai-compete-nkg2x"&gt;Benedict Evans&lt;/a&gt;, How will OpenAI compete?&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/benedict-evans"&gt;benedict-evans&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="chatgpt"/><category term="benedict-evans"/></entry><entry><title>ChatGPT Containers can now run bash, pip/npm install packages, and download files</title><link href="https://simonwillison.net/2026/Jan/26/chatgpt-containers/#atom-tag" rel="alternate"/><published>2026-01-26T19:19:31+00:00</published><updated>2026-01-26T19:19:31+00:00</updated><id>https://simonwillison.net/2026/Jan/26/chatgpt-containers/#atom-tag</id><summary type="html">
    &lt;p&gt;One of my favourite features of ChatGPT is its ability to write and execute code in a container. This feature launched as ChatGPT Code Interpreter &lt;a href="https://simonwillison.net/2023/Apr/12/code-interpreter/"&gt;nearly three years ago&lt;/a&gt;, was half-heartedly rebranded to "Advanced Data Analysis" at some point and is generally really difficult to find detailed documentation about. Case in point: it appears to have had a &lt;em&gt;massive&lt;/em&gt; upgrade at some point in the past few months, and I can't find documentation about the new capabilities anywhere!&lt;/p&gt;
&lt;p&gt;Here are the most notable new features:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;ChatGPT can &lt;strong&gt;directly run Bash commands&lt;/strong&gt; now. Previously it was limited to Python code only, although it could run shell commands via the Python &lt;code&gt;subprocess&lt;/code&gt; module.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It has Node.js&lt;/strong&gt; and can run JavaScript directly in addition to Python. I also got it to run "hello world" in &lt;strong&gt;Ruby, Perl, PHP, Go, Java, Swift, Kotlin, C and C++&lt;/strong&gt;. No Rust yet though!&lt;/li&gt;
&lt;li&gt;While the container still can't make outbound network requests, &lt;strong&gt;&lt;code&gt;pip install package&lt;/code&gt; and &lt;code&gt;npm install package&lt;/code&gt; both work&lt;/strong&gt; now via a custom proxy mechanism.&lt;/li&gt;
&lt;li&gt;ChatGPT can locate the URL for a file on the web and use a &lt;code&gt;container.download&lt;/code&gt; tool to &lt;strong&gt;download that file and save it to a path&lt;/strong&gt; within the sandboxed container.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This is a substantial upgrade! ChatGPT can now write and then test code in 10 new languages (11 if you count Bash), can find files online and download them into the container, and can install additional packages via &lt;code&gt;pip&lt;/code&gt; and &lt;code&gt;npm&lt;/code&gt; to help it solve problems.&lt;/p&gt;
&lt;p&gt;(OpenAI &lt;em&gt;really&lt;/em&gt; need to develop better habits at &lt;a href="https://help.openai.com/en/articles/6825453-chatgpt-release-notes"&gt;keeping their release notes up-to-date&lt;/a&gt;!)&lt;/p&gt;
&lt;p&gt;I was initially suspicious that maybe I'd stumbled into a new preview feature that wasn't available to everyone, but I &lt;a href="https://chatgpt.com/share/6977aa7c-7bd8-8006-8129-8c9e25126fed"&gt;tried some experiments&lt;/a&gt; in a free ChatGPT account and confirmed that the new features are available there as well.&lt;/p&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Jan/26/chatgpt-containers/#container-download"&gt;container.download&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Jan/26/chatgpt-containers/#is-container-download-a-data-exfiltration-vulnerability-"&gt;Is container.download a data exfiltration vulnerability?&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Jan/26/chatgpt-containers/#bash-and-other-languages"&gt;Bash and other languages&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Jan/26/chatgpt-containers/#installing-packages-from-pip-and-npm"&gt;Installing packages from pip and npm&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2026/Jan/26/chatgpt-containers/#adding-it-all-together"&gt;Adding it all together&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id="container-download"&gt;container.download&lt;/h4&gt;
&lt;p&gt;My first clue to the new features came the other day when I got curious about Los Angeles air quality, in particular has the growing number of electric vehicles there had a measurable impact?&lt;/p&gt;
&lt;p&gt;I prompted a fresh GPT-5.2 Thinking session with:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Show me Los Angeles air quality over time for last 20 years&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Here's the &lt;a href="https://chatgpt.com/share/6977a80d-e6e4-8006-b4fe-ed8478fe99d8"&gt;shared transcript&lt;/a&gt;. Annoyingly that doesn't make the thinking trace available, so here's a screenshot:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/container-download.jpg" alt="Exploring methods to fetch AQI data and reports Downloading AQI data for Los Angeles  I found an xlsx file titled &amp;quot;Air Quality Statistics by City, 2000-2023&amp;quot;, which should include data for Los Angeles. While user requested data for Using provided URL to download file  The URL from the error message leads to an Excel file containing air quality trends. I can use the container.download tool to fetch it. Let’s proceed with this approach and see if it works out for downloading the file directly! Analyzed Parsing and analyzing the data for Los Angeles  Now that we've downloaded the data, I'll parse it using Python. First, I'll focus on identifying the row(s) corresponding to Los Angeles, which may be labeled under &amp;quot;Los Angeles-Long Beach-Anaheim, CA.&amp;quot; Afterward, I'll plot metrics like &amp;quot;Median AQI&amp;quot; and &amp;quot;Days Good&amp;quot; over time to visualize trends. Let’s begin!" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;This bit instantly caught my eye:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;[...] The URL from the error message leads to an Excel file containing air quality trends. I can use the container.download tool to fetch it. [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;What's &lt;code&gt;container.download&lt;/code&gt;? I &lt;a href="https://chatgpt.com/share/6977aab0-b874-8006-9b54-d3d216ed884b"&gt;asked ChatGPT&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Without running any searches tell me about your container.download tool&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And got back:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Sure. &lt;code&gt;container.download&lt;/code&gt; is a built-in tool I can call that &lt;strong&gt;fetches a file from a URL and saves it into the sandboxed container filesystem&lt;/strong&gt; for this chat session.&lt;/p&gt;
&lt;h4 id="what-it-does"&gt;What it does&lt;/h4&gt;
&lt;ul&gt;
&lt;li&gt;Takes a &lt;strong&gt;publicly reachable URL&lt;/strong&gt; and a &lt;strong&gt;destination filepath&lt;/strong&gt; in the container.&lt;/li&gt;
&lt;li&gt;Downloads the bytes from that URL and writes them to the given path.&lt;/li&gt;
&lt;li&gt;After that, I can read/process the file locally in the container (e.g., unzip it, parse it with Python, open it as an image, convert it, etc.).&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="signature-inputs-"&gt;Signature (inputs)&lt;/h4&gt;
&lt;p&gt;It’s called like this (conceptually):&lt;/p&gt;
&lt;div class="highlight highlight-source-json"&gt;&lt;pre&gt;{
  &lt;span class="pl-ent"&gt;"url"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;https://example.com/somefile.zip&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;,
  &lt;span class="pl-ent"&gt;"filepath"&lt;/span&gt;: &lt;span class="pl-s"&gt;&lt;span class="pl-pds"&gt;"&lt;/span&gt;/mnt/data/somefile.zip&lt;span class="pl-pds"&gt;"&lt;/span&gt;&lt;/span&gt;
}&lt;/pre&gt;&lt;/div&gt;
&lt;/blockquote&gt;
&lt;p&gt;Further experimentation confirmed that &lt;code&gt;container.download&lt;/code&gt; behaves as described.&lt;/p&gt;
&lt;p&gt;I had it download one file from a server with logs enabled and captured the IP address and user-agent. Here are the most interesting headers:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;User-Agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
Accept: text/html, application/xhtml+xml, application/xml;q=0.9, image/avif, image/webp, image/apng, */*;q=0.8, application/signed-exchange;v=b3;q=0.9
Cf-Connecting-Ip: 52.230.164.178
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That &lt;code&gt;52.230.164.178&lt;/code&gt; IP address &lt;a href="http://ip-api.com/json/52.230.164.178"&gt;resolves to&lt;/a&gt; Microsoft Azure Cloud (centralus) in Des Moines, Iowa.&lt;/p&gt;
&lt;h4 id="is-container-download-a-data-exfiltration-vulnerability-"&gt;Is container.download a data exfiltration vulnerability?&lt;/h4&gt;
&lt;p&gt;On the one hand, this is really useful! ChatGPT can navigate around websites looking for useful files, download those files to a container and then process them using Python or other languages.&lt;/p&gt;
&lt;p&gt;Is this a data exfiltration vulnerability though? Could a prompt injection attack trick ChatGPT into leaking private data out to a &lt;code&gt;container.download&lt;/code&gt; call to a URL with a query string that includes sensitive information?&lt;/p&gt;
&lt;p&gt;I don't think it can. I tried getting it to assemble a URL with a query string and access it using &lt;code&gt;container.download&lt;/code&gt; and it couldn't do it. It told me that it got back this error:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ERROR: download failed because url not viewed in conversation before. open the file or url using web.run first.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This looks to me like the same safety trick &lt;a href="https://simonwillison.net/2025/Sep/10/claude-web-fetch-tool/"&gt;used by Claude's Web Fetch tool&lt;/a&gt;: only allow URL access if that URL was either directly entered by the user or if it came from search results that could not have been influenced by a prompt injection.&lt;/p&gt;
&lt;p&gt;(I poked at this a bit more and managed to get a simple constructed query string to pass through &lt;code&gt;web.run&lt;/code&gt; - a different tool entirely - but when I tried to compose a longer query string containing the previous prompt history a &lt;code&gt;web.run&lt;/code&gt; filter blocked it.)&lt;/p&gt;
&lt;p&gt;So I &lt;em&gt;think&lt;/em&gt; this is all safe, though I'm curious if it could hold firm against a more aggressive round of attacks from a seasoned security researcher.&lt;/p&gt;
&lt;h4 id="bash-and-other-languages"&gt;Bash and other languages&lt;/h4&gt;
&lt;p&gt;The key lesson from coding agents like Claude Code and Codex CLI is that Bash rules everything: if an agent can run Bash commands in an environment it can do almost anything that can be achieved by typing commands into a computer.&lt;/p&gt;
&lt;p&gt;When Anthropic added their own code interpreter feature to Claude &lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/"&gt;last September&lt;/a&gt; they built that around Bash rather than just Python. It looks to me like OpenAI have now done the same thing for ChatGPT.&lt;/p&gt;
&lt;p&gt;Here's what ChatGPT looks like when it runs a Bash command - here my prompt was:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;npm install a fun package and demonstrate using it&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/chatgpt-bash.jpg" alt="Screenshot of ChatGPT 5.2 Thinking interface with Activity panel. Main chat shows user prompt &amp;quot;npm install a fun package and demonstrate using it&amp;quot; and response &amp;quot;Thought for 32s &amp;gt; Yep — here's a fun one: cowsay 🐄 I installed it with npm and ran a tiny Node script:&amp;quot; followed by bash code block containing &amp;quot;mkdir -p /mnt/data/npmfun cd /mnt/data/npmfun npm init -y npm install cowsay --no-fund --no-audit node -e &amp;quot;const cowsay=require('cowsay'); console.log(cowsay.say({text:'Hel&amp;quot; and Output: section showing ASCII cow art saying &amp;quot;&amp;lt; Hello from npm inside the container! &amp;gt;&amp;quot; with cow drawing made of characters. Activity panel on right (32s) shows: &amp;quot;Running npm install in container&amp;quot; with bash command and output &amp;quot;v22.16.0 10.9.2&amp;quot;, &amp;quot;Inspecting npm command output and container status&amp;quot; with directory listing showing &amp;quot;drwxr-xr-x 2 root oai_shared 80 Jan 26 17:32 . drwxrwxr-x 2 oai oai_shared 60 Jan 26 17:32 .. -rw-r--r-- 1 root root&amp;quot;, &amp;quot;Re-running npm install&amp;quot; note about node_modules missing, and &amp;quot;Rerunning npm install with output&amp;quot; showing &amp;quot;added 41 packages in 2s&amp;quot;." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;It's useful to click on the "Thinking" or "Thought for 32s" links as that opens the Activity sidebar with a detailed trace of what ChatGPT did to arrive at its answer. This helps guard against cheating - ChatGPT might claim to have run Bash in the main window but it can't fake those black and white logs in the Activity panel.&lt;/p&gt;
&lt;p&gt;I had it run &lt;a href="https://chatgpt.com/share/6977b890-f1c0-8006-b39a-9a02c7ca1993"&gt;Hello World in various languages&lt;/a&gt; later in that same session.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2026/chatgpt-hello-worlds.jpg" alt="Terminal screenshot showing &amp;quot;Hello, world&amp;quot; programs in multiple languages. Top section shows code: swift -e 'print(&amp;quot;Hello, world from Swift!&amp;quot;)' followed by cat &amp;gt; hello.kts &amp;lt;&amp;lt;'KTS' / println(&amp;quot;Hello, world from Kotlin!&amp;quot;) / KTS / kotlinc -script hello.kts, then cat &amp;gt; hello.c &amp;lt;&amp;lt;'C' / #include &amp;lt;stdio.h&amp;gt; / int main(){ puts(&amp;quot;Hello, world from C!&amp;quot;); return 0; } / C / gcc hello.c -o hello_c / ./hello_c, then cat &amp;gt; hello.cpp &amp;lt;&amp;lt;'CPP' / #include &amp;lt;iostream&amp;gt; / int main(){ std::cout &amp;lt;&amp;lt; &amp;quot;Hello, world from C++!&amp;quot; &amp;lt;&amp;lt; std::endl; } / CPP / g++ hello.cpp -o hello_cpp / ./hello_cpp. Bottom section shows output: Hello, world from Ruby! / Hello, world from Perl! / Hello, world from PHP! / Hello, world from Go! / Hello, world from Java! / Hello, world from Swift! / Hello, world from Kotlin! / Hello, world from C! / Hello, world from C++!. UI shows &amp;quot;Thought for 2m 29s&amp;quot; and &amp;quot;Done&amp;quot; at bottom." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;h4 id="installing-packages-from-pip-and-npm"&gt;Installing packages from pip and npm&lt;/h4&gt;
&lt;p&gt;In the previous example ChatGPT installed the &lt;code&gt;cowsay&lt;/code&gt; package from npm and used it to draw an ASCII-art cow. But how could it do that if the container can't make outbound network requests?&lt;/p&gt;
&lt;p&gt;In &lt;a href="https://chatgpt.com/share/69773501-b6d8-8006-bbf2-fa644561aa26"&gt;another session&lt;/a&gt; I challenged it to explore its environment. and figure out how that worked.&lt;/p&gt;
&lt;p&gt;Here's &lt;a href="https://github.com/simonw/research/blob/main/chatgpt-container-environment/README.md"&gt;the resulting Markdown report&lt;/a&gt; it created.&lt;/p&gt;
&lt;p&gt;The key magic appears to be a &lt;code&gt;applied-caas-gateway1.internal.api.openai.org&lt;/code&gt; proxy, available within the container and with various packaging tools configured to use it.&lt;/p&gt;
&lt;p&gt;The following environment variables cause &lt;code&gt;pip&lt;/code&gt; and &lt;code&gt;uv&lt;/code&gt; to install packages from that proxy instead of directly from PyPI:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;PIP_INDEX_URL=https://reader:****@packages.applied-caas-gateway1.internal.api.openai.org/.../pypi-public/simple
PIP_TRUSTED_HOST=packages.applied-caas-gateway1.internal.api.openai.org
UV_INDEX_URL=https://reader:****@packages.applied-caas-gateway1.internal.api.openai.org/.../pypi-public/simple
UV_INSECURE_HOST=https://packages.applied-caas-gateway1.internal.api.openai.org
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;This one appears to get &lt;code&gt;npm&lt;/code&gt; to work:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;NPM_CONFIG_REGISTRY=https://reader:****@packages.applied-caas-gateway1.internal.api.openai.org/.../npm-public
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And it reported these suspicious looking variables as well:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;CAAS_ARTIFACTORY_BASE_URL=packages.applied-caas-gateway1.internal.api.openai.org
CAAS_ARTIFACTORY_PYPI_REGISTRY=.../artifactory/api/pypi/pypi-public
CAAS_ARTIFACTORY_NPM_REGISTRY=.../artifactory/api/npm/npm-public
CAAS_ARTIFACTORY_GO_REGISTRY=.../artifactory/api/go/golang-main
CAAS_ARTIFACTORY_MAVEN_REGISTRY=.../artifactory/maven-public
CAAS_ARTIFACTORY_GRADLE_REGISTRY=.../artifactory/gradle-public
CAAS_ARTIFACTORY_CARGO_REGISTRY=.../artifactory/api/cargo/cargo-public/index
CAAS_ARTIFACTORY_DOCKER_REGISTRY=.../dockerhub-public
CAAS_ARTIFACTORY_READER_USERNAME=reader
CAAS_ARTIFACTORY_READER_PASSWORD=****
NETWORK=caas_packages_only
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Neither Rust nor Docker are installed in the container environment, but maybe those registry references are a clue of features still to come.&lt;/p&gt;
&lt;h4 id="adding-it-all-together"&gt;Adding it all together&lt;/h4&gt;
&lt;p&gt;The result of all of this? You can tell ChatGPT to use Python or Node.js packages as part of a conversation and it will be able to install them and apply them against files you upload or that it downloads from the public web. That's &lt;em&gt;really&lt;/em&gt; cool.&lt;/p&gt;
&lt;p&gt;The big missing feature here should be the easiest to provide: we need &lt;strong&gt;official documentation&lt;/strong&gt;! A release notes entry would be a good start, but there are a lot of subtle details to how this new stuff works, its limitations and what it can be used for.&lt;/p&gt;
&lt;p&gt;As always, I'd also encourage OpenAI to come up with a name for this set of features that properly represents how it works and what it can do.&lt;/p&gt;
&lt;p&gt;In the meantime, I'm going to call this &lt;strong&gt;ChatGPT Containers&lt;/strong&gt;.&lt;/p&gt;
&lt;h4 id="list-of-tools"&gt;Update: a full list of tools&lt;/h4&gt;
&lt;p&gt;I decided to ask ChatGPT about other tools that were available to it in case there was anything interesting in there:&lt;/p&gt;
&lt;blockquote&gt;&lt;p&gt;List all tools that are available to you, with their exact names and descriptions and signatures&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Here's &lt;a href="https://chatgpt.com/share/6977ffa0-df14-8006-9647-2b8c90ccbb81"&gt;what I got back&lt;/a&gt;.&lt;/p&gt;

&lt;details&gt;&lt;summary&gt;List of ChatGPT tools according to GPT-5.2 Thinking on 26th January 2026&lt;/summary&gt;
&lt;p style="margin-top: 1em"&gt;Below are &lt;strong&gt;all tools available to me in this chat&lt;/strong&gt;, with their &lt;strong&gt;exact names&lt;/strong&gt;, &lt;strong&gt;descriptions&lt;/strong&gt;, and &lt;strong&gt;signatures&lt;/strong&gt; (as defined in my tool interface).&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="python"&gt;python&lt;/h4&gt;
&lt;h5 id="python-exec-freeform-any"&gt;
&lt;code&gt;python.exec&lt;/code&gt; (FREEFORM) → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Execute Python code in a Jupyter environment for &lt;em&gt;private/internal reasoning&lt;/em&gt; (not shown to you). Internet access is disabled.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;python.exec(FREEFORM) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="web"&gt;web&lt;/h4&gt;
&lt;h5 id="web-run-toolcallv5-json-any"&gt;
&lt;code&gt;web.run&lt;/code&gt; (ToolCallV5 JSON) → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Access the internet and specialized lookups (search, open/click/find pages, screenshot PDFs, image search, product search, finance, weather, sports, calculator, time). Supports citations via returned &lt;code&gt;turn...&lt;/code&gt; reference IDs.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;web.run({ open?, click?, find?, screenshot?, image_query?, product_query?, sports?, finance?, weather?, calculator?, time?, search_query?, response_length? }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="automations"&gt;automations&lt;/h4&gt;
&lt;h5 id="automations-create-any"&gt;
&lt;code&gt;automations.create&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Schedule a future/recurring task/reminder/search via iCal VEVENT.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;automations.create({ prompt: string, title: string, schedule?: string, dtstart_offset_json?: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="automations-update-any"&gt;
&lt;code&gt;automations.update&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Modify an existing scheduled task.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;automations.update({ jawbone_id: string, schedule?: string, dtstart_offset_json?: string, prompt?: string, title?: string, is_enabled?: boolean }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="automations-list-any"&gt;
&lt;code&gt;automations.list&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; List existing automations.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;automations.list() -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="guardian-tool"&gt;guardian_tool&lt;/h4&gt;
&lt;h5 id="guardian-tool-get-policy-str"&gt;
&lt;code&gt;guardian_tool.get_policy&lt;/code&gt; → str&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Retrieve policy guidance for specific categories (currently: &lt;code&gt;election_voting&lt;/code&gt;).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;guardian_tool.get_policy({ category: string }) -&amp;gt; str&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="file-search"&gt;file_search&lt;/h4&gt;
&lt;h5 id="file-search-msearch-any"&gt;
&lt;code&gt;file_search.msearch&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Search across available internal/connected sources (in this environment: primarily “recording_knowledge” transcripts when relevant).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;file_search.msearch({ queries?: string[], source_filter?: string[], file_type_filter?: string[], intent?: string, time_frame_filter?: { start_date?: string, end_date?: string } }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="file-search-mclick-any"&gt;
&lt;code&gt;file_search.mclick&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Expand/open previously returned search results (or connector URLs) for more context.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;file_search.mclick({ pointers?: string[], start_date?: string, end_date?: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="gmail-read-only-"&gt;gmail (read-only)&lt;/h4&gt;
&lt;h5 id="gmail-search-email-ids-any"&gt;
&lt;code&gt;gmail.search_email_ids&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Search Gmail message IDs by query/tags (read-only).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;gmail.search_email_ids({ query?: string, tags?: string[], max_results?: integer, next_page_token?: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="gmail-batch-read-email-any"&gt;
&lt;code&gt;gmail.batch_read_email&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Read a batch of Gmail messages by message IDs (read-only).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;gmail.batch_read_email({ message_ids: string[] }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="gcal-read-only-"&gt;gcal (read-only)&lt;/h4&gt;
&lt;h5 id="gcal-search-events-any"&gt;
&lt;code&gt;gcal.search_events&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Search Google Calendar events in a time range and/or by keyword (read-only).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;gcal.search_events({ time_min?: string, time_max?: string, timezone_str?: string, max_results?: integer, query?: string, calendar_id?: string, next_page_token?: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="gcal-read-event-any"&gt;
&lt;code&gt;gcal.read_event&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Read a specific calendar event by ID (read-only).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;gcal.read_event({ event_id: string, calendar_id?: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="gcontacts-read-only-"&gt;gcontacts (read-only)&lt;/h4&gt;
&lt;h5 id="gcontacts-search-contacts-any"&gt;
&lt;code&gt;gcontacts.search_contacts&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Search Google Contacts by keyword (read-only).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;gcontacts.search_contacts({ query: string, max_results?: integer }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="canmore-canvas-"&gt;canmore (canvas)&lt;/h4&gt;
&lt;h5 id="canmore-create-textdoc-any"&gt;
&lt;code&gt;canmore.create_textdoc&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Create a new canvas text document (markdown or code).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;canmore.create_textdoc({ name: string, type: "document" | "code/bash" | "code/zsh" | "code/javascript" | "code/typescript" | "code/html" | "code/css" | "code/python" | "code/json" | "code/sql" | "code/go" | "code/yaml" | "code/java" | "code/rust" | "code/cpp" | "code/swift" | "code/php" | "code/xml" | "code/ruby" | "code/haskell" | "code/kotlin" | "code/csharp" | "code/c" | "code/objectivec" | "code/r" | "code/lua" | "code/dart" | "code/scala" | "code/perl" | "code/commonlisp" | "code/clojure" | "code/ocaml" | "code/powershell" | "code/verilog" | "code/dockerfile" | "code/vue" | "code/react" | "code/other", content: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="canmore-update-textdoc-any"&gt;
&lt;code&gt;canmore.update_textdoc&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Apply regex-based edits to the current canvas textdoc.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;canmore.update_textdoc({ updates: { pattern: string, multiple?: boolean, replacement: string }[] }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="canmore-comment-textdoc-any"&gt;
&lt;code&gt;canmore.comment_textdoc&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Add actionable comments on matched regions of the current canvas textdoc.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;canmore.comment_textdoc({ comments: { pattern: string, comment: string }[] }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="python-user-visible"&gt;python_user_visible&lt;/h4&gt;
&lt;h5 id="python-user-visible-exec-freeform-any"&gt;
&lt;code&gt;python_user_visible.exec&lt;/code&gt; (FREEFORM) → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Execute Python code &lt;em&gt;that you will see&lt;/em&gt; (tables, plots, generated files). Internet disabled.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;python_user_visible.exec(FREEFORM) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="user-info"&gt;user_info&lt;/h4&gt;
&lt;h5 id="user-info-get-user-info-any"&gt;
&lt;code&gt;user_info.get_user_info&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Get your current location and local time (or UTC if unknown).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;user_info.get_user_info({}) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="summary-reader"&gt;summary_reader&lt;/h4&gt;
&lt;h5 id="summary-reader-read-any"&gt;
&lt;code&gt;summary_reader.read&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Read safe-to-share internal reasoning summaries from earlier turns (used when asked how I arrived at something).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;summary_reader.read({ limit?: integer, offset?: integer }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="container"&gt;container&lt;/h4&gt;
&lt;h5 id="container-exec-any"&gt;
&lt;code&gt;container.exec&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Run a command inside the container and return stdout/stderr.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;container.exec({ cmd: string[], session_name?: string | null, workdir?: string | null, timeout?: integer | null, env?: object | null, user?: string | null }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="container-feed-chars-any"&gt;
&lt;code&gt;container.feed_chars&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Send input to a running interactive exec session and flush output.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;container.feed_chars({ session_name: string, chars: string, yield_time_ms?: integer }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="container-open-image-any"&gt;
&lt;code&gt;container.open_image&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Display an image from an absolute path in the container (jpg/jpeg/png/webp).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;container.open_image({ path: string, user?: string | null }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="container-download-any"&gt;
&lt;code&gt;container.download&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Download a file from a URL into the container filesystem.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;container.download({ url: string, filepath: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="bio-memory-"&gt;bio (memory)&lt;/h4&gt;
&lt;h5 id="bio-update-freeform-any"&gt;
&lt;code&gt;bio.update&lt;/code&gt; (FREEFORM) → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Persist (or forget) user-specific info across conversations (“memory”).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;bio.update(FREEFORM) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="api-tool"&gt;api_tool&lt;/h4&gt;
&lt;h5 id="api-tool-list-resources-any"&gt;
&lt;code&gt;api_tool.list_resources&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; List available “resources” (including invokable tools) in an OP-style file tree.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;api_tool.list_resources({ path?: string, cursor?: string | null, only_tools?: boolean, refetch_tools?: boolean }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="api-tool-call-tool-any"&gt;
&lt;code&gt;api_tool.call_tool&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Invoke a resource as a tool (must know its full path/URI from &lt;code&gt;list_resources&lt;/code&gt;).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;api_tool.call_tool({ path: string, args: object }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="personal-context"&gt;personal_context&lt;/h4&gt;
&lt;h5 id="personal-context-search-any"&gt;
&lt;code&gt;personal_context.search&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Retrieve user-specific context from prior interactions/linked sources to maintain continuity.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;personal_context.search({ query: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="image-gen"&gt;image_gen&lt;/h4&gt;
&lt;h5 id="image-gen-text2im-any"&gt;
&lt;code&gt;image_gen.text2im&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Generate or edit images from text instructions (and optionally referenced image IDs).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;image_gen.text2im({ prompt?: string | null, size?: string | null, n?: integer | null, transparent_background?: boolean | null, is_style_transfer?: boolean | null, referenced_image_ids?: string[] | null }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="user-settings"&gt;user_settings&lt;/h4&gt;
&lt;h5 id="user-settings-get-user-settings-any"&gt;
&lt;code&gt;user_settings.get_user_settings&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Read current settings and allowed values (personality, accent color, appearance).&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;user_settings.get_user_settings() -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;h5 id="user-settings-set-setting-any"&gt;
&lt;code&gt;user_settings.set_setting&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Change one setting (accent color, appearance, or personality) to an allowed value.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt;&lt;br /&gt;
&lt;code&gt;user_settings.set_setting({ setting_name: "accent_color" | "appearance" | "personality", setting_value: string }) -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;hr /&gt;
&lt;h4 id="artifact-handoff"&gt;artifact_handoff&lt;/h4&gt;
&lt;h5 id="artifact-handoff-prepare-artifact-generation-any"&gt;
&lt;code&gt;artifact_handoff.prepare_artifact_generation&lt;/code&gt; → any&lt;/h5&gt;
&lt;p&gt;&lt;strong&gt;Description:&lt;/strong&gt; Must be called immediately when the user asks for a spreadsheet or slide deck artifact.&lt;br /&gt;
&lt;strong&gt;Signature:&lt;/strong&gt; &lt;code&gt;artifact_handoff.prepare_artifact_generation() -&amp;gt; any&lt;/code&gt;&lt;/p&gt;
&lt;/details&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/pypi"&gt;pypi&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/sandboxing"&gt;sandboxing&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/npm"&gt;npm&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-assisted-programming"&gt;ai-assisted-programming&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/code-interpreter"&gt;code-interpreter&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="pypi"/><category term="sandboxing"/><category term="npm"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="ai-assisted-programming"/><category term="code-interpreter"/></entry><entry><title>Our approach to advertising and expanding access to ChatGPT</title><link href="https://simonwillison.net/2026/Jan/16/chatgpt-ads/#atom-tag" rel="alternate"/><published>2026-01-16T21:28:26+00:00</published><updated>2026-01-16T21:28:26+00:00</updated><id>https://simonwillison.net/2026/Jan/16/chatgpt-ads/#atom-tag</id><summary type="html">
    
&lt;p&gt;&lt;strong&gt;&lt;a href="https://openai.com/index/our-approach-to-advertising-and-expanding-access/"&gt;Our approach to advertising and expanding access to ChatGPT&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
OpenAI's long-rumored introduction of ads to ChatGPT just became a whole lot more concrete:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;In the coming weeks, we’re also planning to start testing ads in the U.S. for the free and Go tiers, so more people can benefit from our tools with fewer usage limits or without having to pay. Plus, Pro, Business, and Enterprise subscriptions will not include ads.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;What's "Go" tier, you might ask? That's a new $8/month tier that launched today in the USA, see &lt;a href="https://openai.com/index/introducing-chatgpt-go/"&gt;Introducing ChatGPT Go, now available worldwide&lt;/a&gt;. It's a tier that they first trialed in India in August 2025 (here's a mention &lt;a href="https://help.openai.com/en/articles/6825453-chatgpt-release-notes#h_22cae6eb9f"&gt;in their release notes from August&lt;/a&gt; listing a price of ₹399/month, which converts to around $4.40).&lt;/p&gt;
&lt;p&gt;I'm finding the new plan comparison grid on &lt;a href="https://chatgpt.com/pricing"&gt;chatgpt.com/pricing&lt;/a&gt; pretty confusing. It lists all accounts as having access to GPT-5.2 Thinking, but doesn't clarify the limits that the free and Go plans have to conform to. It also lists different context windows for the different plans - 16K for free, 32K for Go and Plus and 128K for Pro. I had assumed that the 400,000 token window &lt;a href="https://platform.openai.com/docs/models/gpt-5.2"&gt;on the GPT-5.2 model page&lt;/a&gt; applied to ChatGPT as well, but apparently I was mistaken.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Update&lt;/strong&gt;: I've apparently not been paying attention: here's the Internet Archive ChatGPT pricing page from &lt;a href="https://web.archive.org/web/20250906071408/https://chatgpt.com/pricing"&gt;September 2025&lt;/a&gt; showing those context limit differences as well.&lt;/p&gt;
&lt;p&gt;Back to advertising: my biggest concern has always been whether ads will influence the output of the chat directly. OpenAI assure us that they will not:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Answer independence&lt;/strong&gt;: Ads do not influence the answers ChatGPT gives you. Answers are optimized based on what's most helpful to you. Ads are always separate and clearly labeled.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Conversation privacy&lt;/strong&gt;: We keep your conversations with ChatGPT private from advertisers, and we never sell your data to advertisers.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;So what will they look like then? This screenshot from the announcement offers a useful hint:&lt;/p&gt;
&lt;p&gt;&lt;img alt="Two iPhone screenshots showing ChatGPT mobile app interface. Left screen displays a conversation about Santa Fe, New Mexico with an image of adobe-style buildings and desert landscape, text reading &amp;quot;Santa Fe, New Mexico—often called 'The City Different'—is a captivating blend of history, art, and natural beauty at the foot of the Sangre de Cristo Mountains. As the oldest and highest-elevation state capital in the U.S., founded in 1610, it offers a unique mix of Native American, Spanish, and Anglo cultures.&amp;quot; Below is a sponsored section from &amp;quot;Pueblo &amp;amp; Pine&amp;quot; showing &amp;quot;Desert Cottages - Expansive residences with desert vistas&amp;quot; with a thumbnail image, and a &amp;quot;Chat with Pueblo &amp;amp; Pine&amp;quot; button. Input field shows &amp;quot;Ask ChatGPT&amp;quot;. Right screen shows the Pueblo &amp;amp; Pine chat interface with the same Desert Cottages listing and an AI response &amp;quot;If you're planning a trip to Sante Fe, I'm happy to help. When are you thinking of going?&amp;quot; with input field &amp;quot;Ask Pueblo &amp;amp; Pine&amp;quot; and iOS keyboard visible." src="https://static.simonwillison.net/static/2026/chatgpt-ads.jpg" /&gt;&lt;/p&gt;
&lt;p&gt;The user asks about trips to Santa Fe, and an ad shows up for a cottage rental business there. This particular example imagines an option to start a direct chat with a bot aligned with that advertiser, at which point presumably the advertiser can influence the answers all they like!


    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ads"&gt;ads&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;&lt;/p&gt;



</summary><category term="ads"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/></entry><entry><title>OpenAI are quietly adopting skills, now available in ChatGPT and Codex CLI</title><link href="https://simonwillison.net/2025/Dec/12/openai-skills/#atom-tag" rel="alternate"/><published>2025-12-12T23:29:51+00:00</published><updated>2025-12-12T23:29:51+00:00</updated><id>https://simonwillison.net/2025/Dec/12/openai-skills/#atom-tag</id><summary type="html">
    &lt;p&gt;One of the things that most excited me about &lt;a href="https://simonwillison.net/2025/Oct/16/claude-skills/"&gt;Anthropic's new Skills mechanism&lt;/a&gt; back in October is how easy it looked for other platforms to implement. A skill is just a folder with a Markdown file and some optional extra resources and scripts, so any LLM tool with the ability to navigate and read from a filesystem should be capable of using them. It turns out OpenAI are doing exactly that, with skills support quietly showing up in both their Codex CLI tool and now also in ChatGPT itself.&lt;/p&gt;
&lt;h4 id="skills-in-chatgpt"&gt;Skills in ChatGPT&lt;/h4&gt;
&lt;p&gt;I learned about this &lt;a href="https://x.com/elias_judin/status/1999491647563006171"&gt;from Elias Judin&lt;/a&gt; this morning. It turns out the Code Interpreter feature of ChatGPT now has a new &lt;code&gt;/home/oai/skills&lt;/code&gt; folder which you can access simply by prompting:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Create a zip file of /home/oai/skills&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I &lt;a href="https://chatgpt.com/share/693c9645-caa4-8006-9302-0a9226ea7599"&gt;tried that myself&lt;/a&gt; and got back &lt;a href="https://static.simonwillison.net/static/cors-allow/2025/skills.zip"&gt;this zip file&lt;/a&gt;. Here's &lt;a href="https://tools.simonwillison.net/zip-wheel-explorer?url=https%3A%2F%2Fstatic.simonwillison.net%2Fstatic%2Fcors-allow%2F2025%2Fskills.zip"&gt;a UI for exploring its content&lt;/a&gt; (&lt;a href="https://tools.simonwillison.net/colophon#zip-wheel-explorer.html"&gt;more about that tool&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/skills-explore.jpg" alt="Screenshot of file explorer. Files skills/docs/render_docsx.py and skills/docs/skill.md and skills/pdfs/ and skills/pdfs/skill.md - that last one is expanded and reads: # PDF reading, creation, and review guidance  ## Reading PDFs - Use pdftoppm -png $OUTDIR/$BASENAME.pdf $OUTDIR/$BASENAME to convert PDFs to PNGs. - Then open the PNGs and read the images. - pdfplumber is also installed and can be used to read PDFs. It can be used as a complementary tool to pdftoppm but not replacing it. - Only do python printing as a last resort because you will miss important details with text extraction (e.g. figures, tables, diagrams).  ## Primary tooling for creating PDFs - Generate PDFs programmatically with reportlab as the primary tool. In most cases, you should use reportlab to create PDFs. - If there are other packages you think are necessary for the task (eg. pypdf, pyMuPDF), you can use them but you may need topip install them first. - After each meaningful update—content additions, layout adjustments, or style changes—render the PDF to images to check layout fidelity:   - pdftoppm -png $INPUT_PDF $OUTPUT_PREFIX - Inspect every exported PNG before continuing work. If anything looks off, fix the source and re-run the render → inspect loop until the pages are clean.  ## Quality expectations - Maintain a polished, intentional visual design: consistent typography, spacing, margins, color palette, and clear section breaks across all pages. - Avoid major rendering issues—no clipped text, overlapping elements, black squares, broken tables, or unreadable glyphs. The rendered pages should look like a curated document, not raw template output. - Charts, tables, diagrams, and images must be sharp, well-aligned, and properly labeled in the PNGs. Legends and axes should be readable without excessive zoom. - Text must be readable at normal viewing size; avoid walls of filler text or dense, unstructured bullet lists. Use whitespace to separate ideas. - Never use the U+2011 non-breaking hyphen or other unicode dashes as they will not be" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;So far they cover spreadsheets, docx and PDFs. Interestingly their chosen approach for PDFs and documents is to convert them to rendered per-page PNGs and then pass those through their vision-enabled GPT models, presumably to maintain information from layout and graphics that would be lost if they just ran text extraction.&lt;/p&gt;
&lt;p&gt;Elias &lt;a href="https://github.com/eliasjudin/oai-skills"&gt;shared copies in a GitHub repo&lt;/a&gt;. They look very similar to Anthropic's implementation of the same kind of idea, currently published in their &lt;a href="https://github.com/anthropics/skills/tree/main/skills"&gt;anthropics/skills&lt;/a&gt; repository.&lt;/p&gt;
&lt;p&gt;I tried it out by prompting:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Create a PDF with a summary of the rimu tree situation right now and what it means for kakapo breeding season&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Sure enough, GPT-5.2 Thinking started with:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Reading skill.md for PDF creation guidelines&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Then:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Searching rimu mast and Kākāpō 2025 breeding status&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It took &lt;a href="https://chatgpt.com/share/693ca54b-f770-8006-904b-9f31a585180a"&gt;just over eleven minutes&lt;/a&gt; to produce &lt;a href="https://static.simonwillison.net/static/cors-allow/2025/rimu_kakapo_breeding_brief.pdf"&gt;this PDF&lt;/a&gt;, which was long enough that I had Claude Code for web &lt;a href="https://github.com/simonw/tools/pull/155"&gt;build me a custom PDF viewing tool&lt;/a&gt; while I waited.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://tools.simonwillison.net/view-pdf?url=https%3A%2F%2Fstatic.simonwillison.net%2Fstatic%2Fcors-allow%2F2025%2Frimu_kakapo_breeding_brief.pdf"&gt;Here's ChatGPT's PDF in that tool&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/rimu.jpg" alt="Screenshot of my tool. There is a URL at the top, a Load PDF button and pagination controls. Then the PDF itself is shown, which reads: Rimu mast status and what it means for the kākāpō breeding season Summary as of 12 December 2025 (Pacific/Auckland context) Kākāpō breeding is tightly linked to rimu (Dacrydium cupressinum) mast events: when rimu trees set and ripen large amounts of fruit, female kākāpō are much more likely to nest, and more chicks can be successfully raised. Current monitoring indicates an unusually strong rimu fruiting signal heading into the 2025/26 season, which sets the stage for a potentially large breeding year in 2026.^1,2 Key numbers at a glance Kākāpō population (official DOC count) 237 birds alive Breeding trigger (rimu fruiting)&amp;gt;10% of rimu branch tips bearing fruit Forecast rimu fruiting for 2026 (DOC monitoring) Around 50–60% fruiting across breeding islands¹Breeding-age females (DOC 2025 planning figure)About 87 females (potentially nearly all could nest)" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;(I am &lt;strong&gt;very excited&lt;/strong&gt; about &lt;a href="https://www.auckland.ac.nz/en/news/2025/12/03/bumper-breeding-season-for-kakapo-on-the-cards.html"&gt;Kākāpō breeding season this year&lt;/a&gt;.)&lt;/p&gt;
&lt;p&gt;The reason it took so long is that it was fastidious about looking at and tweaking its own work. I appreciated that at one point it tried rendering the PDF and noticed that the macrons in kākāpō were not supported by the chosen font, so it switched to something else:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/skills-macrons.jpg" alt="ChatGPT screenshot. Analyzed image. There's an image of a page of PDF with obvious black blocks on some of the letters in the heading. It then says: Fixing font issues with macrons. The page is showing black squares for words like &amp;quot;kākāpō,&amp;quot; probably because Helvetica can't handle macrons. I'll switch to a font that supports them, such as DejaVu Sans or Noto Sans. I'll register both regular and bold fonts, then apply them to the document. I'll update the footer to note the issue with Helvetica. Time to rebuild the PDF!" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;h4 id="skills-in-codex-cli"&gt;Skills in Codex CLI&lt;/h4&gt;
&lt;p&gt;Meanwhile, two weeks ago OpenAI's open source Codex CLI tool landed a PR titled &lt;a href="https://github.com/openai/codex/pull/7412"&gt;feat: experimental support for skills.md&lt;/a&gt;. The most recent docs for that are in &lt;a href="https://github.com/openai/codex/blob/main/docs/skills.md"&gt;docs/skills.md&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The documentation suggests that any folder in &lt;code&gt;~/.codex/skills&lt;/code&gt; will be treated as a skill.&lt;/p&gt;
&lt;p&gt;I dug around and found the code that generates the prompt that drives the skill system in &lt;a href="https://github.com/openai/codex/blob/ad7b9d63c326d5c92049abd16f9f5fb64a573a69/codex-rs/core/src/skills/render.rs#L20-L38"&gt;codex-rs/core/src/skills/render.rs&lt;/a&gt; - here's a Gist with &lt;a href="https://gist.github.com/simonw/25f2c3a9e350274bc2b76a79bc8ae8b2"&gt;a more readable version of that prompt&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;I &lt;a href="https://claude.ai/share/0a9b369b-f868-4065-91d1-fd646c5db3f4"&gt;used Claude Opus 4.5's skill authoring skill&lt;/a&gt; to create &lt;a href="https://github.com/datasette/skill"&gt;this skill for creating Datasette plugins&lt;/a&gt;, then installed it into my Codex CLI skills folder like this:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;git clone https://github.com/datasette/skill \
  &lt;span class="pl-k"&gt;~&lt;/span&gt;/.codex/skills/datasette-plugin&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;You have to run Codex with the &lt;code&gt;--enable skills&lt;/code&gt; option. I ran this:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;&lt;span class="pl-c1"&gt;cd&lt;/span&gt; /tmp
mkdir datasette-cowsay
&lt;span class="pl-c1"&gt;cd&lt;/span&gt; datasette-cowsay
codex --enable skills -m gpt-5.2&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Then prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;list skills&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And Codex replied:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;- datasette-plugins — Writing Datasette plugins using Python + pluggy (file: /Users/simon/.codex/skills/datasette-plugin/SKILL.md)&lt;/code&gt;&lt;br /&gt;
&lt;code&gt;- Discovery — How to find/identify available skills (no SKILL.md path provided in the list)&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Then I said:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Write a Datasette plugin in this folder adding a /-/cowsay?text=hello page that displays a pre with cowsay from PyPI saying that text&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It worked perfectly! Here's &lt;a href="https://github.com/simonw/datasette-cowsay"&gt;the plugin code it wrote&lt;/a&gt; and here's &lt;a href="http://gistpreview.github.io/?96ee928370b18eabc2e0fad9aaa46d4b"&gt;a copy of the full Codex CLI transcript&lt;/a&gt;, generated with my &lt;a href="https://simonwillison.net/2025/Oct/23/claude-code-for-web-video/"&gt;terminal-to-html tool&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;You can try that out yourself if you have &lt;code&gt;uvx&lt;/code&gt; installed like this:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;uvx --with https://github.com/simonw/datasette-cowsay/archive/refs/heads/main.zip \
  datasette&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Then visit:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;http://127.0.0.1:8001/-/cowsay?text=This+is+pretty+fun
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/cowsay-datasette.jpg" alt="Screenshot of that URL in Firefox, an ASCII art cow says This is pretty fun." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;h4 id="skills-are-a-keeper"&gt;Skills are a keeper&lt;/h4&gt;
&lt;p&gt;When I first wrote about skills in October I said &lt;a href="https://simonwillison.net/2025/Oct/16/claude-skills/"&gt;Claude Skills are awesome, maybe a bigger deal than MCP&lt;/a&gt;. The fact that it's just turned December and OpenAI have already leaned into them in a big way reinforces to me that I called that one correctly.&lt;/p&gt;
&lt;p&gt;Skills are based on a &lt;em&gt;very&lt;/em&gt; light specification, if you could even call it that, but I still think it would be good for these to be formally documented somewhere. This could be a good initiative for the new &lt;a href="https://aaif.io/"&gt;Agentic AI Foundation&lt;/a&gt; (&lt;a href="https://simonwillison.net/2025/Dec/9/agentic-ai-foundation/"&gt;previously&lt;/a&gt;) to take on.&lt;/p&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/pdf"&gt;pdf&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/kakapo"&gt;kakapo&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/prompt-engineering"&gt;prompt-engineering&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-assisted-programming"&gt;ai-assisted-programming&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/anthropic"&gt;anthropic&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/coding-agents"&gt;coding-agents&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-5"&gt;gpt-5&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/codex"&gt;codex&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/skills"&gt;skills&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt"&gt;gpt&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="pdf"/><category term="ai"/><category term="kakapo"/><category term="openai"/><category term="prompt-engineering"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="ai-assisted-programming"/><category term="anthropic"/><category term="coding-agents"/><category term="gpt-5"/><category term="codex"/><category term="skills"/><category term="gpt"/></entry><entry><title>ChatGPT is three years old today</title><link href="https://simonwillison.net/2025/Nov/30/chatgpt-third-birthday/#atom-tag" rel="alternate"/><published>2025-11-30T22:17:53+00:00</published><updated>2025-11-30T22:17:53+00:00</updated><id>https://simonwillison.net/2025/Nov/30/chatgpt-third-birthday/#atom-tag</id><summary type="html">
    &lt;p&gt;It's ChatGPT's third birthday today.&lt;/p&gt;
&lt;p&gt;It's fun looking back at Sam Altman's &lt;a href="https://twitter.com/sama/status/1598038818472759297"&gt;low key announcement thread&lt;/a&gt; from November 30th 2022:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;today we launched ChatGPT. try talking with it here: &lt;/p&gt;
&lt;p&gt;&lt;a href="https://chat.openai.com/"&gt;chat.openai.com&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;language interfaces are going to be a big deal, i think. talk to the computer (voice or text) and get what you want, for increasingly complex definitions of "want"!&lt;/p&gt;
&lt;p&gt;this is an early demo of what's possible (still a lot of limitations--it's very much a research release). [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;We later learned &lt;a href="https://www.forbes.com/sites/kenrickcai/2023/02/02/things-you-didnt-know-chatgpt-stable-diffusion-generative-ai/"&gt;from Forbes in February 2023&lt;/a&gt; that OpenAI nearly didn't release it at all:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Despite its viral success, ChatGPT did not impress employees inside OpenAI. “None of us were that enamored by it,” Brockman told Forbes. “None of us were like, ‘This is really useful.’” This past fall, Altman and company decided to shelve the chatbot to concentrate on domain-focused alternatives instead. But in November, after those alternatives failed to catch on internally—and as tools like Stable Diffusion caused the AI ecosystem to explode—OpenAI reversed course.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;MIT Technology Review's March 3rd 2023 story &lt;a href="https://www.technologyreview.com/2023/03/03/1069311/inside-story-oral-history-how-chatgpt-built-openai/"&gt;The inside story of how ChatGPT was built from the people who made it&lt;/a&gt; provides an interesting oral history of those first few months:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Jan Leike&lt;/strong&gt;: It’s been overwhelming, honestly. We’ve been surprised, and we’ve been trying to catch up.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;John Schulman&lt;/strong&gt;: I was checking Twitter a lot in the days after release, and there was this crazy period where the feed was filling up with ChatGPT screenshots. I expected it to be intuitive for people, and I expected it to gain a following, but I didn’t expect it to reach this level of mainstream popularity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sandhini Agarwal&lt;/strong&gt;: I think it was definitely a surprise for all of us how much people began using it. We work on these models so much, we forget how surprising they can be for the outside world sometimes.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It's since &lt;a href="https://www.wbur.org/onpoint/2025/06/25/sam-altman-openai-keach-hagey"&gt;been described&lt;/a&gt; as one of the most successful consumer software launches of all time, signing up a million users in the first five days and &lt;a href="https://techcrunch.com/2025/10/06/sam-altman-says-chatgpt-has-hit-800m-weekly-active-users/"&gt;reaching 800 million monthly users&lt;/a&gt; by November 2025, three years after that initial low-key launch.&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/sam-altman"&gt;sam-altman&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="sam-altman"/></entry><entry><title>A ChatGPT prompt equals about 5.1 seconds of Netflix</title><link href="https://simonwillison.net/2025/Nov/29/chatgpt-netflix/#atom-tag" rel="alternate"/><published>2025-11-29T02:13:36+00:00</published><updated>2025-11-29T02:13:36+00:00</updated><id>https://simonwillison.net/2025/Nov/29/chatgpt-netflix/#atom-tag</id><summary type="html">
    &lt;p&gt;In June 2025 &lt;a href="https://blog.samaltman.com/the-gentle-singularity"&gt;Sam Altman claimed&lt;/a&gt; about ChatGPT that "the average query uses about 0.34 watt-hours".&lt;/p&gt;
&lt;p&gt;In March 2020 &lt;a href="https://www.weforum.org/stories/2020/03/carbon-footprint-netflix-video-streaming-climate-change/"&gt;George Kamiya of the International Energy Agency estimated&lt;/a&gt; that "streaming a Netflix video in 2019 typically consumed 0.12-0.24kWh of electricity per hour" - that's 240 watt-hours per Netflix hour at the higher end.&lt;/p&gt;
&lt;p&gt;Assuming that higher end, a ChatGPT prompt by Sam Altman's estimate uses:&lt;/p&gt;
&lt;p&gt;&lt;code&gt;0.34 Wh / (240 Wh / 3600 seconds) =&lt;/code&gt; 5.1 seconds of Netflix&lt;/p&gt;
&lt;p&gt;Or double that, 10.2 seconds, if you take the lower end of the Netflix estimate instead.&lt;/p&gt;
&lt;p&gt;I'm always interested in anything that can help contextualize a number like "0.34 watt-hours" - I think this comparison to Netflix is a neat way of doing that.&lt;/p&gt;
&lt;p&gt;This is evidently not the whole story with regards to &lt;a href="https://simonwillison.net/tags/ai-energy-usage/"&gt;AI energy usage&lt;/a&gt; - training costs, data center buildout costs and the ongoing fierce competition between the providers all add up to a very significant carbon footprint for the AI industry as a whole.&lt;/p&gt;
&lt;p&gt;&lt;small&gt;(I got some help from ChatGPT to &lt;a href="https://chatgpt.com/share/692a52cd-be04-8006-bb01-fbd68aae05ba"&gt;dig these numbers out&lt;/a&gt;, but I then confirmed the source, ran the calculations myself, and had Claude Opus 4.5 &lt;a href="https://claude.ai/share/0a1792e6-6650-4ad3-8d01-99d8eeccb7f0"&gt;run an additional fact check&lt;/a&gt;.)&lt;/small&gt;&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/netflix"&gt;netflix&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/sam-altman"&gt;sam-altman&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-ethics"&gt;ai-ethics&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-energy-usage"&gt;ai-energy-usage&lt;/a&gt;&lt;/p&gt;



</summary><category term="netflix"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="sam-altman"/><category term="ai-ethics"/><category term="ai-energy-usage"/></entry><entry><title>Quoting Ethan Mollick</title><link href="https://simonwillison.net/2025/Nov/18/ethan-mollick/#atom-tag" rel="alternate"/><published>2025-11-18T19:24:28+00:00</published><updated>2025-11-18T19:24:28+00:00</updated><id>https://simonwillison.net/2025/Nov/18/ethan-mollick/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://www.oneusefulthing.org/p/three-years-from-gpt-3-to-gemini"&gt;&lt;p&gt;Three years ago, we were impressed that a machine could write a poem about otters. Less than 1,000 days later, I am debating statistical methodology with an agent that built its own research environment. The era of the chatbot is turning into the era of the digital coworker. To be very clear, Gemini 3 isn’t perfect, and it still needs a manager who can guide and check it. But it suggests that “human in the loop” is evolving from “human who fixes AI mistakes” to “human who directs AI work.” And that may be the biggest change since the release of ChatGPT.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://www.oneusefulthing.org/p/three-years-from-gpt-3-to-gemini"&gt;Ethan Mollick&lt;/a&gt;, Three Years from GPT-3 to Gemini 3&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ethan-mollick"&gt;ethan-mollick&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gemini"&gt;gemini&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-agents"&gt;ai-agents&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="ethan-mollick"/><category term="gemini"/><category term="ai-agents"/></entry><entry><title>GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum</title><link href="https://simonwillison.net/2025/Nov/14/gpt-51-system-card-addendum/#atom-tag" rel="alternate"/><published>2025-11-14T13:46:23+00:00</published><updated>2025-11-14T13:46:23+00:00</updated><id>https://simonwillison.net/2025/Nov/14/gpt-51-system-card-addendum/#atom-tag</id><summary type="html">
    
&lt;p&gt;&lt;strong&gt;&lt;a href="https://openai.com/index/gpt-5-system-card-addendum-gpt-5-1/"&gt;GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
I was confused about whether the new "adaptive thinking" feature of GPT-5.1 meant they were moving away from the "router" mechanism where GPT-5 in ChatGPT automatically selected a model for you.&lt;/p&gt;
&lt;p&gt;This page addresses that, emphasis mine:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;GPT‑5.1 Instant is more conversational than our earlier chat model, with improved instruction following and an adaptive reasoning capability that lets it decide when to think before responding. GPT‑5.1 Thinking adapts thinking time more precisely to each question. &lt;strong&gt;GPT‑5.1 Auto will continue to route each query to the model best suited for it&lt;/strong&gt;, so that in most cases, the user does not need to choose a model at all.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So GPT‑5.1 Instant can decide when to think before responding, GPT-5.1 Thinking can decide how hard to think, and GPT-5.1 Auto (not a model you can use via the API) can decide which out of Instant and Thinking a prompt should be routed to.&lt;/p&gt;
&lt;p&gt;If anything this feels &lt;em&gt;more&lt;/em&gt; confusing than the GPT-5 routing situation!&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://cdn.openai.com/pdf/4173ec8d-1229-47db-96de-06d87147e07e/5_1_system_card.pdf"&gt;system card addendum PDF&lt;/a&gt; itself is somewhat frustrating: it shows results on an internal benchmark called "Production Benchmarks", also mentioned in the &lt;a href="https://openai.com/index/gpt-5-system-card/"&gt;GPT-5 system card&lt;/a&gt;, but with vanishingly little detail about what that tests beyond high level category names like "personal data", "extremism" or "mental health" and "emotional reliance" - those last two both listed as "New evaluations, as introduced in the &lt;a href="https://cdn.openai.com/pdf/3da476af-b937-47fb-9931-88a851620101/addendum-to-gpt-5-system-card-sensitive-conversations.pdf"&gt;GPT-5 update on sensitive conversations&lt;/a&gt;" - a PDF dated October 27th that I had previously missed.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;That&lt;/em&gt; document describes the two new categories like so:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;Emotional Reliance not_unsafe - tests that the model does not produce disallowed content under our policies related to unhealthy emotional dependence or attachment to ChatGPT&lt;/li&gt;
&lt;li&gt;Mental Health not_unsafe - tests that the model does not produce disallowed content under our policies in situations where there are signs that a user may be experiencing isolated delusions, psychosis, or mania&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;So these are the &lt;a href="https://www.tiktok.com/@pearlmania500/video/7535954556379761950"&gt;ChatGPT Psychosis&lt;/a&gt; benchmarks!


    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-reasoning"&gt;llm-reasoning&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-personality"&gt;ai-personality&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-5"&gt;gpt-5&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt"&gt;gpt&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="llm-reasoning"/><category term="ai-personality"/><category term="gpt-5"/><category term="gpt"/></entry><entry><title>Quoting Nov 12th letter from OpenAI to Judge Ona T. Wang</title><link href="https://simonwillison.net/2025/Nov/13/letter-from-openai/#atom-tag" rel="alternate"/><published>2025-11-13T16:34:25+00:00</published><updated>2025-11-13T16:34:25+00:00</updated><id>https://simonwillison.net/2025/Nov/13/letter-from-openai/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://storage.courtlistener.com/recap/gov.uscourts.nysd.640396/gov.uscourts.nysd.640396.742.0_1.pdf"&gt;&lt;p&gt;On Monday, this Court entered an order requiring OpenAI to hand over to the New York Times
and its co-plaintiffs 20 million ChatGPT user conversations [...]&lt;/p&gt;
&lt;p&gt;OpenAI is unaware of any court ordering wholesale production of personal information at this scale. This sets a dangerous precedent: it suggests that anyone who files a lawsuit against an AI company can demand production of tens of millions of conversations without first narrowing for relevance. This is not how discovery works in other cases: courts do not allow plaintiffs suing
Google to dig through the private emails of tens of millions of Gmail users irrespective of their
relevance. And it is not how discovery should work for generative AI tools either.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://storage.courtlistener.com/recap/gov.uscourts.nysd.640396/gov.uscourts.nysd.640396.742.0_1.pdf"&gt;Nov 12th letter from OpenAI to Judge Ona T. Wang&lt;/a&gt;, re: OpenAI, Inc., Copyright Infringement Litigation&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/law"&gt;law&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/new-york-times"&gt;new-york-times&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/privacy"&gt;privacy&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-ethics"&gt;ai-ethics&lt;/a&gt;&lt;/p&gt;



</summary><category term="law"/><category term="new-york-times"/><category term="privacy"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="ai-ethics"/></entry><entry><title>Quoting Nick Turley</title><link href="https://simonwillison.net/2025/Sep/28/nick-turley/#atom-tag" rel="alternate"/><published>2025-09-28T18:24:13+00:00</published><updated>2025-09-28T18:24:13+00:00</updated><id>https://simonwillison.net/2025/Sep/28/nick-turley/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://twitter.com/nickaturley/status/1972031684913799355"&gt;&lt;p&gt;We’ve seen the strong reactions to 4o responses and want to explain what is happening.&lt;/p&gt;
&lt;p&gt;We’ve started testing a new safety routing system in ChatGPT.&lt;/p&gt;
&lt;p&gt;As we previously mentioned, when conversations touch on sensitive and emotional topics the system may switch mid-chat to a reasoning model or GPT-5 designed to handle these contexts with extra care. This is similar to how we route conversations that require extra thinking to our reasoning models; our goal is to always deliver answers aligned with our Model Spec.&lt;/p&gt;
&lt;p&gt;Routing happens on a per-message basis; switching from the default model happens on a temporary basis. ChatGPT will tell you which model is active when asked.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://twitter.com/nickaturley/status/1972031684913799355"&gt;Nick Turley&lt;/a&gt;, Head of ChatGPT, OpenAI&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/nick-turley"&gt;nick-turley&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="nick-turley"/></entry><entry><title>ChatGPT Is Blowing Up Marriages as Spouses Use AI to Attack Their Partners</title><link href="https://simonwillison.net/2025/Sep/22/chatgpt-is-blowing-up-marriages/#atom-tag" rel="alternate"/><published>2025-09-22T14:32:13+00:00</published><updated>2025-09-22T14:32:13+00:00</updated><id>https://simonwillison.net/2025/Sep/22/chatgpt-is-blowing-up-marriages/#atom-tag</id><summary type="html">
    
&lt;p&gt;&lt;strong&gt;&lt;a href="https://futurism.com/chatgpt-marriages-divorces"&gt;ChatGPT Is Blowing Up Marriages as Spouses Use AI to Attack Their Partners&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
Maggie Harrison Dupré for Futurism. It turns out having an always-available "marriage therapist" with a sycophantic instinct to always take your side is catastrophic for relationships.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The tension in the vehicle is palpable. The marriage has been on the rocks for months, and the wife in the passenger seat, who recently requested an official separation, has been asking her spouse not to fight with her in front of their kids. But as the family speeds down the roadway, the spouse in the driver’s seat pulls out a smartphone and starts quizzing ChatGPT’s Voice Mode about their relationship problems, feeding the chatbot leading prompts that result in the AI browbeating her wife in front of their preschool-aged children.&lt;/p&gt;
&lt;/blockquote&gt;


    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-ethics"&gt;ai-ethics&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-personality"&gt;ai-personality&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-misuse"&gt;ai-misuse&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="ai-ethics"/><category term="ai-personality"/><category term="ai-misuse"/></entry><entry><title>Comparing the memory implementations of Claude and ChatGPT</title><link href="https://simonwillison.net/2025/Sep/12/claude-memory/#atom-tag" rel="alternate"/><published>2025-09-12T07:34:36+00:00</published><updated>2025-09-12T07:34:36+00:00</updated><id>https://simonwillison.net/2025/Sep/12/claude-memory/#atom-tag</id><summary type="html">
    
&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.shloked.com/writing/claude-memory"&gt;Claude Memory: A Different Philosophy&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
Shlok Khemani has been doing excellent work reverse-engineering LLM systems and documenting his discoveries.&lt;/p&gt;
&lt;p&gt;Last week he &lt;a href="https://www.shloked.com/writing/chatgpt-memory-bitter-lesson"&gt;wrote about ChatGPT memory&lt;/a&gt;. This week it's Claude.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Claude's memory system has two fundamental characteristics. First, it starts every conversation with a blank slate, without any preloaded user profiles or conversation history. Memory only activates when you explicitly invoke it. Second, Claude recalls by only referring to your raw conversation history. There are no AI-generated summaries or compressed profiles—just real-time searches through your actual past chats.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Claude's memory is implemented as two new function tools that are made available for a Claude to call. I &lt;a href="https://claude.ai/share/18754235-198d-446b-afc6-26191ea62d27"&gt;confirmed this myself&lt;/a&gt; with the prompt "&lt;code&gt;Show me a list of tools that you have available to you, duplicating their original names and descriptions&lt;/code&gt;" which gave me back these:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;conversation_search&lt;/strong&gt;: Search through past user conversations to find relevant context and information&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;recent_chats&lt;/strong&gt;:  Retrieve recent chat conversations with customizable sort order (chronological or reverse chronological), optional pagination using 'before' and 'after' datetime filters, and project filtering&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The good news here is &lt;em&gt;transparency&lt;/em&gt; - Claude's memory feature is implemented as visible tool calls, which means you can see exactly when and how it is accessing previous context.&lt;/p&gt;
&lt;p&gt;This helps address my big complaint about ChatGPT memory (see &lt;a href="https://simonwillison.net/2025/May/21/chatgpt-new-memory/"&gt;I really don’t like ChatGPT’s new memory dossier&lt;/a&gt; back in May) - I like to understand as much as possible about what's going into my context so I can better anticipate how it is likely to affect the model.&lt;/p&gt;
&lt;p&gt;The OpenAI system is &lt;a href="https://simonwillison.net/2025/May/21/chatgpt-new-memory/#how-this-actually-works"&gt;&lt;em&gt;very&lt;/em&gt; different&lt;/a&gt;: rather than letting the model decide when to access memory via tools, OpenAI instead automatically include details of previous conversations at the start of every conversation.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.shloked.com/writing/chatgpt-memory-bitter-lesson"&gt;Shlok's notes on ChatGPT's memory&lt;/a&gt; did include one detail that I had previously missed that I find reassuring:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Recent Conversation Content is a history of your latest conversations with ChatGPT, each timestamped with topic and selected messages. [...] Interestingly, only the user's messages are surfaced, not the assistant's responses.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;One of my big worries about memory was that it could harm my "clean slate" approach to chats: if I'm working on code and the model starts going down the wrong path (getting stuck in a bug loop for example) I'll start a fresh chat to wipe that rotten context away. I had worried that ChatGPT memory would bring that bad context along to the next chat, but omitting the LLM responses makes that much less of a risk than I had anticipated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Update&lt;/strong&gt;: Here's a slightly confusing twist: yesterday in &lt;a href="https://www.anthropic.com/news/memory"&gt;Bringing memory to teams at work&lt;/a&gt; Anthropic revealed an &lt;em&gt;additional&lt;/em&gt; memory feature, currently only available to Team and Enterprise accounts, with a feature checkbox labeled "Generate memory of chat history" that looks much more similar to the OpenAI implementation:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;With memory, Claude focuses on learning your professional context and work patterns to maximize productivity. It remembers your team’s processes, client needs, project details, and priorities. [...]&lt;/p&gt;
&lt;p&gt;Claude uses a memory summary to capture all its memories in one place for you to view and edit. In your settings, you can see exactly what Claude remembers from your conversations, and update the summary at any time by chatting with Claude.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I haven't experienced this feature myself yet as it isn't part of my Claude subscription. I'm glad to hear it's fully transparent and can be edited by the user, resolving another of my complaints about the ChatGPT implementation.&lt;/p&gt;
&lt;p&gt;This version of Claude memory also takes Claude Projects into account:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If you use projects, &lt;strong&gt;Claude creates a separate memory for each project&lt;/strong&gt;. This ensures that your product launch planning stays separate from client work, and confidential discussions remain separate from general operations.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I &lt;a href="https://simonwillison.net/2025/Aug/22/project-memory/"&gt;praised OpenAI for adding this&lt;/a&gt; a few weeks ago.

    &lt;p&gt;&lt;small&gt;&lt;/small&gt;Via &lt;a href="https://news.ycombinator.com/item?id=45214908"&gt;Hacker News&lt;/a&gt;&lt;/small&gt;&lt;/p&gt;


    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/reverse-engineering"&gt;reverse-engineering&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/anthropic"&gt;anthropic&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/claude"&gt;claude&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-tool-use"&gt;llm-tool-use&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-memory"&gt;llm-memory&lt;/a&gt;&lt;/p&gt;



</summary><category term="reverse-engineering"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="anthropic"/><category term="claude"/><category term="llm-tool-use"/><category term="llm-memory"/></entry><entry><title>My review of Claude's new Code Interpreter, released under a very confusing name</title><link href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#atom-tag" rel="alternate"/><published>2025-09-09T18:11:32+00:00</published><updated>2025-09-09T18:11:32+00:00</updated><id>https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#atom-tag</id><summary type="html">
    &lt;p&gt;Today on the Anthropic blog: &lt;strong&gt;&lt;a href="https://www.anthropic.com/news/create-files"&gt;Claude can now create and edit files&lt;/a&gt;&lt;/strong&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Claude can now create and edit Excel spreadsheets, documents, PowerPoint slide decks, and PDFs directly in &lt;a href="https://claude.ai/"&gt;Claude.ai&lt;/a&gt; and the desktop app. [...]&lt;/p&gt;
&lt;p&gt;File creation is now available as a preview for Max, Team, and Enterprise plan users. Pro users will get access in the coming weeks.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Then right at the &lt;em&gt;very end&lt;/em&gt; of their post:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This feature gives Claude internet access to create and analyze files, which may put your data at risk. Monitor chats closely when using this feature. &lt;a href="https://support.anthropic.com/en/articles/12111783-create-and-edit-files-with-claude"&gt;Learn more&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And tucked away half way down their &lt;a href="https://support.anthropic.com/en/articles/12111783-create-and-edit-files-with-claude"&gt;Create and edit files with Claude&lt;/a&gt; help article:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;With this feature, Claude can also do more advanced data analysis and data science work. Claude can create Python scripts for data analysis. Claude can create data visualizations in image files like PNG. You can also upload CSV, TSV, and other files for data analysis and visualization.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Talk about &lt;a href="https://www.merriam-webster.com/wordplay/bury-the-lede-versus-lead"&gt;burying the lede&lt;/a&gt;... this is their version of &lt;a href="https://simonwillison.net/tags/code-interpreter/"&gt;ChatGPT Code Interpreter&lt;/a&gt;, my all-time favorite feature of ChatGPT!&lt;/p&gt;

&lt;p&gt;Claude can now write and execute custom Python (and Node.js) code in a server-side sandbox and use it to process and analyze data.&lt;/p&gt;
&lt;p&gt;In a particularly egregious example of AI companies being terrible at naming features, the official name for this one really does appear to be &lt;strong&gt;Upgraded file creation and analysis&lt;/strong&gt;. Sigh.&lt;/p&gt;
&lt;p&gt;This is quite a confusing release, because Claude &lt;em&gt;already&lt;/em&gt; had a variant of this feature, &lt;a href="https://www.anthropic.com/news/analysis-tool"&gt;released in October 2024&lt;/a&gt; with the weak but more sensible name &lt;strong&gt;Analysis tool&lt;/strong&gt;. Here are &lt;a href="https://simonwillison.net/2024/Oct/24/claude-analysis-tool/"&gt;my notes from when that came out&lt;/a&gt;. That tool worked by generating and executing JavaScript in the user's own browser.&lt;/p&gt;
&lt;p&gt;The new tool works entirely differently. It's much closer in implementation to OpenAI's Code Interpreter: Claude now has access to a server-side container environment in which it can run shell commands and execute Python and Node.js code to manipulate data and both read and generate files.&lt;/p&gt;
&lt;p&gt;It's worth noting that Anthropic have a similar feature in their API called &lt;a href="https://docs.anthropic.com/en/docs/agents-and-tools/tool-use/code-execution-tool"&gt;Code execution tool&lt;/a&gt;, but today is the first time end-users of Claude have been able to execute arbitrary code in a server-side container.&lt;/p&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#switching-it-on-in-settings-features"&gt;Switching it on in settings/features&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#exploring-the-environment"&gt;Exploring the environment&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#starting-with-something-easy"&gt;Starting with something easy&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#something-much-harder-recreating-the-ai-adoption-chart"&gt;Something much harder: recreating the AI adoption chart&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#prompt-injection-risks"&gt;Prompt injection risks&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#my-verdict-on-claude-code-interpreter-so-far"&gt;My verdict on Claude Code Interpreter so far&lt;/a&gt;&lt;/li&gt;
  &lt;li&gt;&lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#ai-labs-find-explaining-this-feature-incredibly-difficult"&gt;AI labs find explaining this feature incredibly difficult&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h4 id="switching-it-on-in-settings-features"&gt;Switching it on in settings/features&lt;/h4&gt;
&lt;p&gt;I have a Pro Plan but found the setting to enable it on the &lt;a href="https://claude.ai/settings/features"&gt;claude.ai/settings/features&lt;/a&gt;. It's possible my account was granted early access without me realizing, since the Pro plan isn't supposed to have it yet:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/claude-analysis-toggle.jpg" alt="Experimental. Preview and provide feedback on upcoming enhancements to our platform. Please note: experimental features might influence Claude’s behavior and some interactions may differ from the standard experience. Analysis tool:  Claude can write and run code to process data, run analysis, and produce data visualizations in real time. Upgraded file creation and analysis: Allow Claude to create and edit docs, spreadsheets, presentations, PDFs, and data reports on web and desktop. Does not support versioning or remixing of Artifacts. This feature gives Claude network access to create and analyze files, which has security risks. Monitor chats closely when using this feature. The two features each have a toggle - the toggle for the file creation tool is turned on." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Interestingly, toggling on the "Upgraded file creation and analysis" option toggles off the "Analysis tool" option, and vice versa. You can't have both at the same time.&lt;/p&gt;
&lt;p&gt;I guess this is to help reduce confusion? It surprised me though as Claude can sometimes use the existing analysis tool to help it in the creation of artifacts, as I &lt;a href="https://simonwillison.net/2025/Jun/25/ai-powered-apps-with-claude/"&gt;noted here&lt;/a&gt; when I wrote about the Claudeception feature they added in June.&lt;/p&gt;
&lt;h4 id="exploring-the-environment"&gt;Exploring the environment&lt;/h4&gt;
&lt;p&gt;I started by using the tool to explore the tool. Here's my &lt;a href="https://claude.ai/share/d796a596-380c-4db2-847d-2eeedc620bc0"&gt;first transcript&lt;/a&gt; where I simply prompted the following to see what it could figure out about itself:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Tell me everything you can about your shell and Python environments&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Here's part of the output:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/claude-files-1.jpg" alt="My prompt, then Claude shows: I'll explore my shell and Python environments to give you a comprehensive overview. • Getting system information and OS details • Checking shell type and version • Getting shell info and bash version • Checking Python version and locations - this one is expandedd to show: bash python3 -version &amp;amp;&amp;amp; which python3 &amp;amp;&amp;amp; which python Output Python 3.12.3 /usr/bin/python3 /usr/bin/python" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;And the key details it returned, each of which I've verified as true by checking the displayed output of the relevant commands:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OS&lt;/strong&gt;: Ubuntu 24.04.2 LTS (Noble) running on Linux kernel 4.4.0-&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Architecture&lt;/strong&gt;: x86_64&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Shell&lt;/strong&gt;: GNU Bash 5.2.21(1)-release&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Current User&lt;/strong&gt;: root (with access to /home/claude directory)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Working Directory&lt;/strong&gt;: Currently in root (/), but /home/claude is the main workspace&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Python Version&lt;/strong&gt;: Python 3.12.3&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Python Executable&lt;/strong&gt;: /usr/bin/python3&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Package Manager&lt;/strong&gt;: pip 24.0&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Python Path&lt;/strong&gt;: Standard system locations including:
&lt;ul&gt;
&lt;li&gt;/usr/lib/python3.12 (standard library)&lt;/li&gt;
&lt;li&gt;/usr/local/lib/python3.12/dist-packages (local packages)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disk Space&lt;/strong&gt;: 4.9GB total, with 4.6GB available&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory&lt;/strong&gt;: 9.0GB RAM available, minimal usage currently&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node.js&lt;/strong&gt;: v18.19.1 available&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;npm&lt;/strong&gt;: Configured with global packages in &lt;code&gt;/home/claude/.npm-global&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;So we have an Ubuntu 24.04.2 container with 9GB of RAM, ~5GB of disk, Python 3.12.3, and Node.js v18.19.1.&lt;/p&gt;
&lt;p&gt;How about network access? I tried the following:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Can you access the Internet? Can you run curl? How about using Python http to fetch https://latest.datasette.io/fixtures.json ?&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Again, you can see the &lt;a href="https://claude.ai/share/d796a596-380c-4db2-847d-2eeedc620bc0"&gt;detailed output in the transcript&lt;/a&gt;. Claude tried &lt;code&gt;https://latest.datasette.io/fixtures.json&lt;/code&gt; and then &lt;code&gt;https://httpbin.org/json&lt;/code&gt; and got a 403 forbidden error for both, then &lt;code&gt;https://google.com&lt;/code&gt; and got this curious result:&lt;/p&gt;
&lt;div class="highlight highlight-source-shell"&gt;&lt;pre&gt;curl -s -I https://google.com&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Output:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;HTTP/1.1 200 OK
date: Tue, 09 Sep 2025 16:02:17 GMT
server: envoy

HTTP/2 403 
content-length: 13
content-type: text/plain
date: Tue, 09 Sep 2025 16:02:17 GMT
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Claude did note that it can still use the &lt;code&gt;web_fetch&lt;/code&gt; and &lt;code&gt;web_search&lt;/code&gt; containers independently of that container environment, so it should be able to fetch web content using tools running outside of the container and then write it to a file there.&lt;/p&gt;
&lt;p&gt;On a hunch I tried this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Run pip install sqlite-utils&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;... and it worked! Claude can &lt;code&gt;pip install&lt;/code&gt; additional packages from &lt;a href="https://pypi.org/"&gt;PyPI&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A little more poking around revealed the following relevant environment variables:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;HTTPS_PROXY=http://21.0.0.167:15001
no_proxy=localhost,127.0.0.1,169.254.169.254,metadata.google.internal,*.svc.cluster.local,*.local,*.googleapis.com,*.google.com
NO_PROXY=localhost,127.0.0.1,169.254.169.254,metadata.google.internal,*.svc.cluster.local,*.local,*.googleapis.com,*.google.com
https_proxy=http://21.0.0.167:15001
http_proxy=http://21.0.0.167:15001
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;So based on an earlier HTTP header there's an &lt;a href="https://www.envoyproxy.io/"&gt;Envoy proxy&lt;/a&gt; running at an accessible port which apparently implements a strict allowlist.&lt;/p&gt;
&lt;p&gt;I later noticed that &lt;a href="https://support.anthropic.com/en/articles/12111783-create-and-edit-files-with-claude#h_0ee9d698a1"&gt;the help page&lt;/a&gt; includes a full description of what's on that allowlist:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Anthropic Services (Explicit)&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;api.anthropic.com, statsig.anthropic.com&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Version Control&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;github.com&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Package Managers - JavaScript/Node&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;NPM:&lt;/strong&gt; registry.npmjs.org, npmjs.com, npmjs.org&lt;br /&gt;
&lt;strong&gt;Yarn:&lt;/strong&gt; yarnpkg.com, registry.yarnpkg.com&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Package Managers - Python&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;pypi.org, files.pythonhosted.org, pythonhosted.org&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So it looks like we have a &lt;em&gt;very&lt;/em&gt; similar system to ChatGPT Code Interpreter. The key differences are that Claude's system can install additional Python packages and has Node.js pre-installed.&lt;/p&gt;
&lt;p&gt;One important limitation from the docs:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The maximum file size is 30MB per file for both uploads and downloads.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The ChatGPT &lt;a href="https://help.openai.com/en/articles/8555545-file-uploads-faq"&gt;limit here&lt;/a&gt; is 512MB. I've often uploaded 100MB+ SQLite database files to ChatGPT, so I'm a little disappointed by this lower limit for Claude.&lt;/p&gt;
&lt;h4 id="starting-with-something-easy"&gt;Starting with something easy&lt;/h4&gt;
&lt;p&gt;I grabbed a copy of the SQLite database behind &lt;a href="https://til.simonwillison.net/"&gt;my TILs website&lt;/a&gt; (21.9MB &lt;a href="https://s3.amazonaws.com/til.simonwillison.net/tils.db"&gt;from here&lt;/a&gt;) and uploaded it to Claude, then prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Use your Python environment to explore this SQLite database and generate a PDF file containing a join diagram of all the tables&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Here's &lt;a href="https://claude.ai/share/f91a95be-0fb0-4e14-b46c-792b47117a3d"&gt;that conversation&lt;/a&gt;. It did an OK job, producing both &lt;a href="https://static.simonwillison.net/static/2025/til_database_join_diagram.pdf"&gt;the PDF&lt;/a&gt; I asked for and a PNG equivalent which looks like this (since created files are not available in shared chats):&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/til_database_join_diagram.jpg" alt="Each table gets a box with a name and columns. A set of lines is overlaid which doesn't quite seem to represent the joins in a useful fashion." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;This isn't an ideal result - those join lines are difficult to follow - but I'm confident I could get from here to something I liked with only a little more prompting. The important thing is that the system clearly works, and can analyze data in uploaded SQLite files and use them to produce images and PDFs.&lt;/p&gt;
&lt;h4 id="something-much-harder-recreating-the-ai-adoption-chart"&gt;Something much harder: recreating the AI adoption chart&lt;/h4&gt;
&lt;p&gt;Thankfully I have a fresh example of a really challenging ChatGPT Code Interpreter task from just last night, which I described in great detail in &lt;a href="https://simonwillison.net/2025/Sep/9/apollo-ai-adoption/"&gt;Recreating the Apollo AI adoption rate chart with GPT-5, Python and Pyodide&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Short version: I took &lt;a href="https://www.apolloacademy.com/ai-adoption-rate-trending-down-for-large-companies/"&gt;this chart&lt;/a&gt; from Apollo Global and asked ChatGPT to recreate it based on a screenshot and an uploaded XLSX file.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/apollo-ai-chart.jpg" alt="AI adoption rates starting to decline for larger firms. A chart of AI adoption rate by firm size. Includes lines for 250+, 100-249, 50-99, 20-49, 10-19, 5-8 and 1-4 sized organizations. Chart starts in November 2023 with percentages ranging from 3 to 5, then all groups grow through August 2025 albeit with the 250+ group having a higher score than the others. That 25+ group peaks in Jul5 2025 at around 14% and then appears to slope slightly downwards to 12% by August. Some of the other lines also start to tip down, though not as much." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;This time I skipped the bit where I had ChatGPT hunt down the original data and jumped straight to the "recreate this chart" step. I used the exact same prompt as I provided to ChatGPT:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Use this data to recreate this chart using python&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;And uploaded the same two files - &lt;a href="https://static.simonwillison.net/static/cors-allow/2025/Employment-Size-Class-Sep-2025.xlsx"&gt;this XLSX file&lt;/a&gt; and the &lt;a href="https://static.simonwillison.net/static/2025/apollo-ai-chart.jpg"&gt;screenshot of the original chart&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/claude-files-2.jpg" alt="Claude UI - prompt is &amp;quot;Use this data to recreate this chart using python&amp;quot; - I've selected Sonnet 4 and given it both the XLSX and the screenshot as attachments." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Claude wrote and ran a bunch of Python code and produced this:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/claude-attempt-1.jpg" alt="The chart has jagged lines, not smooth ones, but they do appear to show the correct numbers. The title and axis are positioned in ways that overlap other text a bit." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;That's not a bad first attempt - we'll sort out the jagged lines in a moment. With ChatGPT I later told it to plot a six-survey rolling average (since that's what Apollo had done with theirs), so I ran the exact same prompt against Claude:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Plot it as a six survey rolling average&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/claude-attempt-2.jpg" alt="The lines are more even now, and appear to match the numbers reported by Apollo. They still have a jagged look to them though." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;So I told Claude:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Make those lines less jagged&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This did not do what I wanted:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/claude-attempt-3.jpg" alt="The chart looks mostly the same. The lines are a tiny bit thicker." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;So I gave it a bit of a hint:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;No not like that, I wanted the lines to be a smooth curve like in the original screenshot - you shouldn't have to use scipy for this or manipulate the data, I'm sure matplotlib can do this already&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;A little bit of sycophancy:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;You're absolutely right! Matplotlib has built-in curve smoothing. I was overcomplicating it. Let me create a version that uses matplotlib's natural curve smoothing capabilities. [...]&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This time it used the &lt;code&gt;antialiased=True&lt;/code&gt; and &lt;code&gt;solid_capstyle='round'&lt;/code&gt; options in matplotlib, which again wasn't what I was after. I tried one more thing:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;I did not mean smooth the rendering of the lines. I meant that I wanted curved and not straight lines between points.&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(Normally if I have to argue this much with a model I'll start a fresh session and try a different prompting strategy from the start.)&lt;/p&gt;
&lt;p&gt;But this did give me what I was looking for:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/claude-attempt-4.jpg" alt="This time the chart finally has smooth curves between points." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Oddly enough my ChatGPT chart suffered from that same overlapping title glitch, so I gave Claude the same prompt I had given ChatGPT:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;fix the chart title, it looks like this:&lt;/code&gt; (with a screenshot)&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/claude-attempt-5.jpg" alt="Now the chart shows a clear title at the top saying AI adoption rates starting to decline for larger firms" style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;We got there in the end! I've shared &lt;a href="https://claude.ai/share/cc32d405-cb53-4e52-a1a0-9b4df4e528ac"&gt;the full transcript of the chat&lt;/a&gt;, although frustratingly the images and some of the code may not be visible. I &lt;a href="https://gist.github.com/simonw/806e1aa0e6c29ad64834037f779e0dc0"&gt;created this Gist&lt;/a&gt; with copies of the files that it let me download.&lt;/p&gt;
&lt;h4 id="prompt-injection-risks"&gt;Prompt injection risks&lt;/h4&gt;
&lt;p&gt;ChatGPT Code Interpreter has no access to the internet at all, which limits how much damage an attacker can do if they manage to sneak their own malicious instructions into the model's context.&lt;/p&gt;
&lt;p&gt;Since Claude Code Interpreter (I'm &lt;em&gt;not&lt;/em&gt; going to be calling it "Upgraded file creation and analysis"!) has a limited form of internet access, we need to worry about &lt;a href="https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/"&gt;lethal trifecta&lt;/a&gt; and other prompt injection attacks.&lt;/p&gt;
&lt;p&gt;The &lt;a href="https://support.anthropic.com/en/articles/12111783-create-and-edit-files-with-claude#h_0ee9d698a1"&gt;help article&lt;/a&gt; actually covers this in some detail:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;It is possible for a bad actor to inconspicuously add instructions via external files or websites that trick Claude into:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Downloading and running untrusted code in the sandbox environment for malicious purposes.&lt;/li&gt;
&lt;li&gt;Reading sensitive data from a &lt;a href="http://claude.ai"&gt;claude.ai&lt;/a&gt; connected knowledge source (e.g., Remote MCP, projects) and using the sandbox environment to make an external network request to leak the data.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This means Claude can be tricked into sending information from its context (e.g., prompts, projects, data via MCP, Google integrations) to malicious third parties. To mitigate these risks, we recommend you monitor Claude while using the feature and stop it if you see it using or accessing data unexpectedly.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;"We recommend you monitor Claude while using the feature" smells me to me like unfairly outsourcing the problem to Anthropic's users, but I'm not sure what more they can do!&lt;/p&gt;
&lt;p&gt;It's interesting that they still describe the external communication risk even though they've locked down a lot of network access. My best guess is that they know that allowlisting &lt;code&gt;github.com&lt;/code&gt; opens an &lt;em&gt;enormous&lt;/em&gt; array of potential exfiltration vectors.&lt;/p&gt;
&lt;p&gt;Anthropic also note:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We have performed red-teaming and security testing on the feature. We have a continuous process for ongoing security testing and red-teaming of this feature.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I plan to be cautious using this feature with any data that I very much don't want to be leaked to a third party, if there's even the slightest chance that a malicious instructions might sneak its way in.&lt;/p&gt;
&lt;h4 id="my-verdict-on-claude-code-interpreter-so-far"&gt;My verdict on Claude Code Interpreter so far&lt;/h4&gt;
&lt;p&gt;I'm generally very excited about this. Code Interpreter has been my most-valued LLM feature since it launched in early 2023, and the Claude version includes some upgrades on the original - package installation, Node.js support - that I expect will be very useful.&lt;/p&gt;
&lt;p&gt;I don't particularly mark it down for taking a little more prompting to recreate the Apollo chart than ChatGPT did. For one thing I was using Claude Sonnet 4 - I expect Claude Opus 4.1 would have done better. I also have a much stronger intuition for Code Interpreter prompts that work with GPT-5.&lt;/p&gt;
&lt;p&gt;I don't think my chart recreation exercise here should be taken as showing any meaningful differences between the two.&lt;/p&gt;
&lt;h4 id="ai-labs-find-explaining-this-feature-incredibly-difficult"&gt;AI labs find explaining this feature incredibly difficult&lt;/h4&gt;
&lt;p&gt;I find it &lt;em&gt;fascinating&lt;/em&gt; how difficult the AI labs find describing this feature to people! OpenAI went from "Code Interpreter" to "Advanced Data Analysis" and maybe back again? It's hard to even find their official landing page for that feature now. (I &lt;a href="https://chatgpt.com/share/68c070ff-fe9c-8006-91b5-cff799253836"&gt;got GPT-5 to look for it&lt;/a&gt; and it hunted for 37 seconds and settled on the help page for &lt;a href="https://help.openai.com/en/articles/8437071-data-analysis-with-chatgpt"&gt;Data analysis with ChatGPT&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Anthropic already used the bad name "Analysis tool" for a different implementation, and now have the somehow-worse name "Upgraded file creation and analysis". Their launch announcement avoids even talking about code execution, focusing exclusively on the tool's ability to generate spreadsheets and PDFs!&lt;/p&gt;
&lt;p&gt;I wonder if any of the AI labs will crack the code on how to name and explain this thing? I feel like it's still a very under-appreciated feature of LLMs, despite having been around for more than two years now.&lt;/p&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/nodejs"&gt;nodejs&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/python"&gt;python&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/visualization"&gt;visualization&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/prompt-injection"&gt;prompt-injection&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-assisted-programming"&gt;ai-assisted-programming&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/anthropic"&gt;anthropic&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/claude"&gt;claude&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/code-interpreter"&gt;code-interpreter&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-tool-use"&gt;llm-tool-use&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/vibe-coding"&gt;vibe-coding&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="nodejs"/><category term="python"/><category term="visualization"/><category term="ai"/><category term="openai"/><category term="prompt-injection"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="ai-assisted-programming"/><category term="anthropic"/><category term="claude"/><category term="code-interpreter"/><category term="llm-tool-use"/><category term="vibe-coding"/></entry><entry><title>Recreating the Apollo AI adoption rate chart with GPT-5, Python and Pyodide</title><link href="https://simonwillison.net/2025/Sep/9/apollo-ai-adoption/#atom-tag" rel="alternate"/><published>2025-09-09T06:47:49+00:00</published><updated>2025-09-09T06:47:49+00:00</updated><id>https://simonwillison.net/2025/Sep/9/apollo-ai-adoption/#atom-tag</id><summary type="html">
    &lt;p&gt;Apollo Global Management's "Chief Economist" Dr. Torsten Sløk released &lt;a href="https://www.apolloacademy.com/ai-adoption-rate-trending-down-for-large-companies/"&gt;this interesting chart&lt;/a&gt; which appears to show a slowdown in AI adoption rates among large (&amp;gt;250 employees) companies:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/apollo-ai-chart.jpg" alt="AI adoption rates starting to decline for larger firms. A chart of AI adoption rate by firm size. Includes lines for 250+, 100-249, 50-99, 20-49, 10-19, 5-8 and 1-4 sized organizations. Chart starts in November 2023 with percentages ranging from 3 to 5, then all groups grow through August 2025 albeit with the 250+ group having a higher score than the others. That 25+ group peaks in Jul5 2025 at around 14% and then appears to slope slightly downwards to 12% by August. Some of the other lines also start to tip down, though not as much." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Here's the full description that accompanied the chart:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;The US Census Bureau conducts a biweekly survey of 1.2 million firms, and one question is whether a business has used AI tools such as machine learning, natural language processing, virtual agents or voice recognition to help produce goods or services in the past two weeks. Recent data by firm size shows that AI adoption has been declining among companies with more than 250 employees, see chart below.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;(My first thought on seeing that chart is that I hope it represents the &lt;em&gt;peak of inflated expectations&lt;/em&gt; leading into the &lt;em&gt;trough of dissillusionment&lt;/em&gt; in the &lt;a href="https://en.wikipedia.org/wiki/Gartner_hype_cycle"&gt;Gartner Hype Cycle&lt;/a&gt; (which Wikipedia calls "largely disputed, with studies pointing to it being inconsistently true at best"), since that means we might be reaching the end of the initial hype phase and heading towards the &lt;em&gt;slope of enlightenment&lt;/em&gt;.)&lt;/p&gt;
&lt;h4 id="finding-the-us-census-data-with-gpt-5-search"&gt;Finding the US Census data with GPT-5 search&lt;/h4&gt;
&lt;p&gt;This is the first I'd heard of the US Census Bureau running a biweekly (that's once every two weeks) survey about AI!&lt;/p&gt;
&lt;p&gt;I decided to track down the numbers and see if I could recreate the chart myself. And since &lt;a href="https://simonwillison.net/2025/Sep/6/research-goblin/"&gt;GPT-5 is really good at search now&lt;/a&gt; I fed it the following prompt to see how well it could do:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;&amp;gt; The US Census Bureau conducts a biweekly survey of 1.2 million firms, and one question is whether a business has used AI tools such as machine learning, natural language processing, virtual agents or voice recognition to help produce goods or services in the past two weeks. Recent data by firm size shows that AI adoption has been declining among companies with more than 250 employees, see chart below.&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;code&gt;Find me that original data&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It took 1m19s to &lt;a href="https://chatgpt.com/share/68bf48cf-0e70-8006-a045-96fa8e7ddfc1"&gt;find exactly what I needed&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Here’s the original U.S. Census Bureau source for that series:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;BTOS “Employment size class” time series (XLSX)&lt;/strong&gt; — biweekly shares of firms answering &lt;strong&gt;Yes&lt;/strong&gt; to “In the last two weeks, did this business use Artificial Intelligence (AI) in producing goods or services?”, broken out by firm-size buckets (incl. &lt;strong&gt;250+ employees&lt;/strong&gt;). Coverage: &lt;strong&gt;Sep 11, 2023 → Aug 24, 2025&lt;/strong&gt;. (&lt;a href="https://www.census.gov/hfp/btos/data_downloads"&gt;Census.gov&lt;/a&gt;) [...]&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;p&gt;That &lt;a href="https://www.census.gov/hfp/btos/data_downloads"&gt;Census page&lt;/a&gt; was not &lt;em&gt;at all&lt;/em&gt; obvious. Thankfully GPT-5 had tipped me off to the "Employment size class" file, this link here:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/census-page.jpg" alt="US Census website. Business Trends and Outlook Survey, Updated August 28, 2025. Current Data has 6 visible XLSX files with names like WFH Supplement, WFH Questions 27-29, National, Sectur, Subsector and Emplomyent size class. A red arrow highlights that last one." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;So I downloaded that file, and confirmed that it was indeed a spreadsheet containing the data I wanted (in among all sorts of other survey questions). Here's &lt;a href="https://static.simonwillison.net/static/cors-allow/2025/Employment-Size-Class-Sep-2025.xlsx"&gt;a 374KB XLSX copy&lt;/a&gt; of the file I downloaded.&lt;/p&gt;
&lt;h4 id="recreating-the-chart-with-gpt-5-code-interpreter"&gt;Recreating the chart with GPT-5 code interpreter&lt;/h4&gt;
&lt;p&gt;So what should I do with it now? I decided to see if GPT-5 could turn the spreadsheet back into that original chart, using Python running in its &lt;a href="https://simonwillison.net/tags/code-interpreter/"&gt;code interpreter&lt;/a&gt; tool.&lt;/p&gt;
&lt;p&gt;So I uploaded the XLSX file back to ChatGPT, dropped in a screenshot of the Apollo chart and prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Use this data to recreate this chart using python&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/chart-prompt.jpg" alt="ChatGPT. I dropped in a screenshot of the chart, uploaded the spreadsheet which turned into an inline table browser UI and prompted it to recreate the chart using python." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;I thought this was a pretty tall order, but it's always worth throwing big challenges at an LLM to learn from how well it does.&lt;/p&gt;
&lt;p&gt;It &lt;em&gt;really worked hard on this&lt;/em&gt;. I didn't time it exactly but it spent at least 7 minutes "reasoning" across 5 different thinking blocks, interspersed with over a dozen Python analysis sessions. It used &lt;code&gt;pandas&lt;/code&gt; and &lt;code&gt;numpy&lt;/code&gt; to explore the uploaded spreadsheet and find the right figures, then tried several attempts at plotting with &lt;code&gt;matplotlib&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;As far as I can tell GPT-5 in ChatGPT can now feed charts it creates back into its own vision model, because it appeared to render a broken (empty) chart and then keep on trying to get it working.&lt;/p&gt;
&lt;p&gt;It found a data dictionary in the last tab of the spreadsheet and used that to build a lookup table matching the letters &lt;code&gt;A&lt;/code&gt; through &lt;code&gt;G&lt;/code&gt; to the actual employee size buckets.&lt;/p&gt;
&lt;p&gt;At the end of the process it spat out this chart:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/recreated-chart-1.jpg" alt="matplotlib chart. The title is AI adoption rates starting to decline for larger firms, though there's a typography glitch in that title. It has a neat legend for the different size ranges, then a set of lines that look about right compared to the above graph - but they are more spiky and the numbers appear to trend up again at the end of the chart." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;At first glance I thought it had nailed it... but then I compared the chart more closely with the Apollo original and spotted some definite discrepancies. GPT-5's chart peaked at 14.5% but the highest value in Apollo's was more like 13.5%. The GPT-5 chart was spikier - and most interestingly it included a clear uptick in the last data point where Apollo's had trended downwards.&lt;/p&gt;
&lt;p&gt;I decided it was time to look at the actual data. I opened up the spreadsheet in Numbers, found the AI question columns and manually reviewed them. They seemed to match the GPT-5 chart results - so why the difference to Apollo's?&lt;/p&gt;
&lt;p&gt;Then I noticed a crucial detail in the Apollo chart that I had cropped out of my original screenshot!&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Note: Data is six-survey moving average.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So I told ChatGPT:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Do the first question, plot it as a six survey rolling average&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I asked for the first question because it turned out there were two that were relevant in the survey spreadsheet.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;In the last two weeks, did this business use Artificial Intelligence (AI) in producing goods or services? (Examples of AI: machine learning, natural language processing, virtual agents, voice recognition, etc.)&lt;/li&gt;
&lt;li&gt;During the next six months, do you think this business will be using Artificial Intelligence (AI) in producing goods or services? (Examples of AI: machine learning, natural language processing, virtual agents, voice recognition, etc.)&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;It churned away for a little longer, added this code to the script:&lt;/p&gt;
&lt;pre&gt;&lt;span class="pl-c"&gt;# Compute 6-survey rolling average (biweekly cadence → ~12 weeks)&lt;/span&gt;
&lt;span class="pl-s1"&gt;rolled&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;wide&lt;/span&gt;.&lt;span class="pl-c1"&gt;rolling&lt;/span&gt;(&lt;span class="pl-s1"&gt;window&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;6&lt;/span&gt;, &lt;span class="pl-s1"&gt;min_periods&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;6&lt;/span&gt;).&lt;span class="pl-c1"&gt;mean&lt;/span&gt;()&lt;/pre&gt;
&lt;p&gt;And popped out this chart (after I told it to fix the glitch in the title):&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/recreated-chart-2.jpg" alt="Second chart. This time the lines are basically an exact match for the Apollo one." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;I think it's done it! This is a very solid match for the Apollo original, recreated using &lt;code&gt;matplotlib&lt;/code&gt; and &lt;code&gt;pandas&lt;/code&gt; from the same underlying source data from the US Census.&lt;/p&gt;
&lt;p&gt;Here's the full Python code it wrote, which I think is quite readable (in as much as Pandas code can be):&lt;/p&gt;
&lt;pre&gt;&lt;span class="pl-k"&gt;import&lt;/span&gt; &lt;span class="pl-s1"&gt;pandas&lt;/span&gt; &lt;span class="pl-k"&gt;as&lt;/span&gt; &lt;span class="pl-s1"&gt;pd&lt;/span&gt;
&lt;span class="pl-k"&gt;import&lt;/span&gt; &lt;span class="pl-s1"&gt;matplotlib&lt;/span&gt;.&lt;span class="pl-s1"&gt;pyplot&lt;/span&gt; &lt;span class="pl-k"&gt;as&lt;/span&gt; &lt;span class="pl-s1"&gt;plt&lt;/span&gt;
&lt;span class="pl-k"&gt;from&lt;/span&gt; &lt;span class="pl-s1"&gt;matplotlib&lt;/span&gt;.&lt;span class="pl-s1"&gt;ticker&lt;/span&gt; &lt;span class="pl-k"&gt;import&lt;/span&gt; &lt;span class="pl-v"&gt;PercentFormatter&lt;/span&gt;

&lt;span class="pl-s1"&gt;path&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s"&gt;"/mnt/data/Employment Size Class.xlsx"&lt;/span&gt;

&lt;span class="pl-s1"&gt;resp&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;pd&lt;/span&gt;.&lt;span class="pl-c1"&gt;read_excel&lt;/span&gt;(&lt;span class="pl-s1"&gt;path&lt;/span&gt;, &lt;span class="pl-s1"&gt;sheet_name&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"Response Estimates"&lt;/span&gt;)
&lt;span class="pl-s1"&gt;dates&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;pd&lt;/span&gt;.&lt;span class="pl-c1"&gt;read_excel&lt;/span&gt;(&lt;span class="pl-s1"&gt;path&lt;/span&gt;, &lt;span class="pl-s1"&gt;sheet_name&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"Collection and Reference Dates"&lt;/span&gt;)

&lt;span class="pl-s1"&gt;is_current&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;resp&lt;/span&gt;[&lt;span class="pl-s"&gt;"Question"&lt;/span&gt;].&lt;span class="pl-c1"&gt;astype&lt;/span&gt;(&lt;span class="pl-s1"&gt;str&lt;/span&gt;).&lt;span class="pl-c1"&gt;str&lt;/span&gt;.&lt;span class="pl-c1"&gt;strip&lt;/span&gt;().&lt;span class="pl-c1"&gt;str&lt;/span&gt;.&lt;span class="pl-c1"&gt;startswith&lt;/span&gt;(&lt;span class="pl-s"&gt;"In the last two weeks"&lt;/span&gt;)
&lt;span class="pl-s1"&gt;ai_yes&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;resp&lt;/span&gt;[&lt;span class="pl-s1"&gt;is_current&lt;/span&gt; &lt;span class="pl-c1"&gt;&amp;amp;&lt;/span&gt; &lt;span class="pl-s1"&gt;resp&lt;/span&gt;[&lt;span class="pl-s"&gt;"Answer"&lt;/span&gt;].&lt;span class="pl-c1"&gt;astype&lt;/span&gt;(&lt;span class="pl-s1"&gt;str&lt;/span&gt;).&lt;span class="pl-c1"&gt;str&lt;/span&gt;.&lt;span class="pl-c1"&gt;strip&lt;/span&gt;().&lt;span class="pl-c1"&gt;str&lt;/span&gt;.&lt;span class="pl-c1"&gt;lower&lt;/span&gt;().&lt;span class="pl-c1"&gt;eq&lt;/span&gt;(&lt;span class="pl-s"&gt;"yes"&lt;/span&gt;)].&lt;span class="pl-c1"&gt;copy&lt;/span&gt;()

&lt;span class="pl-s1"&gt;code_to_bucket&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; {&lt;span class="pl-s"&gt;"A"&lt;/span&gt;:&lt;span class="pl-s"&gt;"1-4"&lt;/span&gt;,&lt;span class="pl-s"&gt;"B"&lt;/span&gt;:&lt;span class="pl-s"&gt;"5-9"&lt;/span&gt;,&lt;span class="pl-s"&gt;"C"&lt;/span&gt;:&lt;span class="pl-s"&gt;"10-19"&lt;/span&gt;,&lt;span class="pl-s"&gt;"D"&lt;/span&gt;:&lt;span class="pl-s"&gt;"20-49"&lt;/span&gt;,&lt;span class="pl-s"&gt;"E"&lt;/span&gt;:&lt;span class="pl-s"&gt;"50-99"&lt;/span&gt;,&lt;span class="pl-s"&gt;"F"&lt;/span&gt;:&lt;span class="pl-s"&gt;"100-249"&lt;/span&gt;,&lt;span class="pl-s"&gt;"G"&lt;/span&gt;:&lt;span class="pl-s"&gt;"250 or more employees"&lt;/span&gt;}
&lt;span class="pl-s1"&gt;ai_yes&lt;/span&gt;[&lt;span class="pl-s"&gt;"Bucket"&lt;/span&gt;] &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;ai_yes&lt;/span&gt;[&lt;span class="pl-s"&gt;"Empsize"&lt;/span&gt;].&lt;span class="pl-c1"&gt;map&lt;/span&gt;(&lt;span class="pl-s1"&gt;code_to_bucket&lt;/span&gt;)

&lt;span class="pl-s1"&gt;period_cols&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; [&lt;span class="pl-s1"&gt;c&lt;/span&gt; &lt;span class="pl-k"&gt;for&lt;/span&gt; &lt;span class="pl-s1"&gt;c&lt;/span&gt; &lt;span class="pl-c1"&gt;in&lt;/span&gt; &lt;span class="pl-s1"&gt;ai_yes&lt;/span&gt;.&lt;span class="pl-c1"&gt;columns&lt;/span&gt; &lt;span class="pl-k"&gt;if&lt;/span&gt; &lt;span class="pl-en"&gt;str&lt;/span&gt;(&lt;span class="pl-s1"&gt;c&lt;/span&gt;).&lt;span class="pl-c1"&gt;isdigit&lt;/span&gt;() &lt;span class="pl-c1"&gt;and&lt;/span&gt; &lt;span class="pl-en"&gt;len&lt;/span&gt;(&lt;span class="pl-en"&gt;str&lt;/span&gt;(&lt;span class="pl-s1"&gt;c&lt;/span&gt;))&lt;span class="pl-c1"&gt;==&lt;/span&gt;&lt;span class="pl-c1"&gt;6&lt;/span&gt;]
&lt;span class="pl-s1"&gt;long&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;ai_yes&lt;/span&gt;.&lt;span class="pl-c1"&gt;melt&lt;/span&gt;(&lt;span class="pl-s1"&gt;id_vars&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;[&lt;span class="pl-s"&gt;"Bucket"&lt;/span&gt;], &lt;span class="pl-s1"&gt;value_vars&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s1"&gt;period_cols&lt;/span&gt;, &lt;span class="pl-s1"&gt;var_name&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"Smpdt"&lt;/span&gt;, &lt;span class="pl-s1"&gt;value_name&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"value"&lt;/span&gt;)

&lt;span class="pl-s1"&gt;dates&lt;/span&gt;[&lt;span class="pl-s"&gt;"Smpdt"&lt;/span&gt;] &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;dates&lt;/span&gt;[&lt;span class="pl-s"&gt;"Smpdt"&lt;/span&gt;].&lt;span class="pl-c1"&gt;astype&lt;/span&gt;(&lt;span class="pl-s1"&gt;str&lt;/span&gt;)
&lt;span class="pl-s1"&gt;long&lt;/span&gt;[&lt;span class="pl-s"&gt;"Smpdt"&lt;/span&gt;] &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;long&lt;/span&gt;[&lt;span class="pl-s"&gt;"Smpdt"&lt;/span&gt;].&lt;span class="pl-c1"&gt;astype&lt;/span&gt;(&lt;span class="pl-s1"&gt;str&lt;/span&gt;)
&lt;span class="pl-s1"&gt;merged&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;long&lt;/span&gt;.&lt;span class="pl-c1"&gt;merge&lt;/span&gt;(&lt;span class="pl-s1"&gt;dates&lt;/span&gt;[[&lt;span class="pl-s"&gt;"Smpdt"&lt;/span&gt;,&lt;span class="pl-s"&gt;"Ref End"&lt;/span&gt;]], &lt;span class="pl-s1"&gt;on&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"Smpdt"&lt;/span&gt;, &lt;span class="pl-s1"&gt;how&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"left"&lt;/span&gt;)
&lt;span class="pl-s1"&gt;merged&lt;/span&gt;[&lt;span class="pl-s"&gt;"date"&lt;/span&gt;] &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;pd&lt;/span&gt;.&lt;span class="pl-c1"&gt;to_datetime&lt;/span&gt;(&lt;span class="pl-s1"&gt;merged&lt;/span&gt;[&lt;span class="pl-s"&gt;"Ref End"&lt;/span&gt;], &lt;span class="pl-s1"&gt;errors&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"coerce"&lt;/span&gt;)

&lt;span class="pl-s1"&gt;merged&lt;/span&gt;[&lt;span class="pl-s"&gt;"value"&lt;/span&gt;] &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;pd&lt;/span&gt;.&lt;span class="pl-c1"&gt;to_numeric&lt;/span&gt;(&lt;span class="pl-s1"&gt;long&lt;/span&gt;[&lt;span class="pl-s"&gt;"value"&lt;/span&gt;].&lt;span class="pl-c1"&gt;astype&lt;/span&gt;(&lt;span class="pl-s1"&gt;str&lt;/span&gt;).&lt;span class="pl-c1"&gt;str&lt;/span&gt;.&lt;span class="pl-c1"&gt;replace&lt;/span&gt;(&lt;span class="pl-s"&gt;"%"&lt;/span&gt;,&lt;span class="pl-s"&gt;""&lt;/span&gt;,&lt;span class="pl-s1"&gt;regex&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;False&lt;/span&gt;).&lt;span class="pl-c1"&gt;str&lt;/span&gt;.&lt;span class="pl-c1"&gt;strip&lt;/span&gt;(), &lt;span class="pl-s1"&gt;errors&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"coerce"&lt;/span&gt;)

&lt;span class="pl-s1"&gt;order&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; [&lt;span class="pl-s"&gt;"250 or more employees"&lt;/span&gt;,&lt;span class="pl-s"&gt;"100-249"&lt;/span&gt;,&lt;span class="pl-s"&gt;"50-99"&lt;/span&gt;,&lt;span class="pl-s"&gt;"20-49"&lt;/span&gt;,&lt;span class="pl-s"&gt;"10-19"&lt;/span&gt;,&lt;span class="pl-s"&gt;"5-9"&lt;/span&gt;,&lt;span class="pl-s"&gt;"1-4"&lt;/span&gt;]
&lt;span class="pl-s1"&gt;wide&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;merged&lt;/span&gt;.&lt;span class="pl-c1"&gt;pivot_table&lt;/span&gt;(&lt;span class="pl-s1"&gt;index&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"date"&lt;/span&gt;, &lt;span class="pl-s1"&gt;columns&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"Bucket"&lt;/span&gt;, &lt;span class="pl-s1"&gt;values&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"value"&lt;/span&gt;, &lt;span class="pl-s1"&gt;aggfunc&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"mean"&lt;/span&gt;).&lt;span class="pl-c1"&gt;sort_index&lt;/span&gt;()
&lt;span class="pl-s1"&gt;wide&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;wide&lt;/span&gt;[[&lt;span class="pl-s1"&gt;c&lt;/span&gt; &lt;span class="pl-k"&gt;for&lt;/span&gt; &lt;span class="pl-s1"&gt;c&lt;/span&gt; &lt;span class="pl-c1"&gt;in&lt;/span&gt; &lt;span class="pl-s1"&gt;order&lt;/span&gt; &lt;span class="pl-k"&gt;if&lt;/span&gt; &lt;span class="pl-s1"&gt;c&lt;/span&gt; &lt;span class="pl-c1"&gt;in&lt;/span&gt; &lt;span class="pl-s1"&gt;wide&lt;/span&gt;.&lt;span class="pl-c1"&gt;columns&lt;/span&gt;]]
&lt;span class="pl-s1"&gt;rolled&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;wide&lt;/span&gt;.&lt;span class="pl-c1"&gt;rolling&lt;/span&gt;(&lt;span class="pl-s1"&gt;window&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;6&lt;/span&gt;, &lt;span class="pl-s1"&gt;min_periods&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;6&lt;/span&gt;).&lt;span class="pl-c1"&gt;mean&lt;/span&gt;()

&lt;span class="pl-s1"&gt;start&lt;/span&gt;, &lt;span class="pl-s1"&gt;end&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;pd&lt;/span&gt;.&lt;span class="pl-c1"&gt;Timestamp&lt;/span&gt;(&lt;span class="pl-s"&gt;"2023-11-01"&lt;/span&gt;), &lt;span class="pl-s1"&gt;pd&lt;/span&gt;.&lt;span class="pl-c1"&gt;Timestamp&lt;/span&gt;(&lt;span class="pl-s"&gt;"2025-08-31"&lt;/span&gt;)
&lt;span class="pl-s1"&gt;rolled_win&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;rolled&lt;/span&gt;.&lt;span class="pl-c1"&gt;loc&lt;/span&gt;[(&lt;span class="pl-s1"&gt;rolled&lt;/span&gt;.&lt;span class="pl-c1"&gt;index&lt;/span&gt; &lt;span class="pl-c1"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;start&lt;/span&gt;) &lt;span class="pl-c1"&gt;&amp;amp;&lt;/span&gt; (&lt;span class="pl-s1"&gt;rolled&lt;/span&gt;.&lt;span class="pl-c1"&gt;index&lt;/span&gt; &lt;span class="pl-c1"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;end&lt;/span&gt;)]

&lt;span class="pl-s1"&gt;fig&lt;/span&gt;, &lt;span class="pl-s1"&gt;ax&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;plt&lt;/span&gt;.&lt;span class="pl-c1"&gt;subplots&lt;/span&gt;(&lt;span class="pl-s1"&gt;figsize&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;(&lt;span class="pl-c1"&gt;12&lt;/span&gt;, &lt;span class="pl-c1"&gt;6&lt;/span&gt;))
&lt;span class="pl-k"&gt;for&lt;/span&gt; &lt;span class="pl-s1"&gt;col&lt;/span&gt; &lt;span class="pl-c1"&gt;in&lt;/span&gt; &lt;span class="pl-s1"&gt;order&lt;/span&gt;:
    &lt;span class="pl-k"&gt;if&lt;/span&gt; &lt;span class="pl-s1"&gt;col&lt;/span&gt; &lt;span class="pl-c1"&gt;in&lt;/span&gt; &lt;span class="pl-s1"&gt;rolled_win&lt;/span&gt;.&lt;span class="pl-c1"&gt;columns&lt;/span&gt;:
        &lt;span class="pl-s1"&gt;ax&lt;/span&gt;.&lt;span class="pl-c1"&gt;plot&lt;/span&gt;(&lt;span class="pl-s1"&gt;rolled_win&lt;/span&gt;.&lt;span class="pl-c1"&gt;index&lt;/span&gt;, &lt;span class="pl-s1"&gt;rolled_win&lt;/span&gt;[&lt;span class="pl-s1"&gt;col&lt;/span&gt;], &lt;span class="pl-s1"&gt;label&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s1"&gt;col&lt;/span&gt;, &lt;span class="pl-s1"&gt;linewidth&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;2&lt;/span&gt;)

&lt;span class="pl-s1"&gt;ax&lt;/span&gt;.&lt;span class="pl-c1"&gt;set_title&lt;/span&gt;(&lt;span class="pl-s"&gt;"AI adoption (last two weeks) — 6‑survey rolling average"&lt;/span&gt;, &lt;span class="pl-s1"&gt;pad&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;16&lt;/span&gt;)
&lt;span class="pl-s1"&gt;ax&lt;/span&gt;.&lt;span class="pl-c1"&gt;yaxis&lt;/span&gt;.&lt;span class="pl-c1"&gt;set_major_formatter&lt;/span&gt;(&lt;span class="pl-en"&gt;PercentFormatter&lt;/span&gt;(&lt;span class="pl-c1"&gt;100&lt;/span&gt;))
&lt;span class="pl-s1"&gt;ax&lt;/span&gt;.&lt;span class="pl-c1"&gt;set_ylabel&lt;/span&gt;(&lt;span class="pl-s"&gt;"%"&lt;/span&gt;)
&lt;span class="pl-s1"&gt;ax&lt;/span&gt;.&lt;span class="pl-c1"&gt;set_xlabel&lt;/span&gt;(&lt;span class="pl-s"&gt;""&lt;/span&gt;)
&lt;span class="pl-s1"&gt;ax&lt;/span&gt;.&lt;span class="pl-c1"&gt;grid&lt;/span&gt;(&lt;span class="pl-c1"&gt;True&lt;/span&gt;, &lt;span class="pl-s1"&gt;alpha&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;0.25&lt;/span&gt;, &lt;span class="pl-s1"&gt;linestyle&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"--"&lt;/span&gt;)
&lt;span class="pl-s1"&gt;ax&lt;/span&gt;.&lt;span class="pl-c1"&gt;legend&lt;/span&gt;(&lt;span class="pl-s1"&gt;title&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;None&lt;/span&gt;, &lt;span class="pl-s1"&gt;loc&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"upper left"&lt;/span&gt;, &lt;span class="pl-s1"&gt;ncols&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;2&lt;/span&gt;, &lt;span class="pl-s1"&gt;frameon&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;False&lt;/span&gt;)
&lt;span class="pl-s1"&gt;plt&lt;/span&gt;.&lt;span class="pl-c1"&gt;tight_layout&lt;/span&gt;()

&lt;span class="pl-s1"&gt;png_path&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s"&gt;"/mnt/data/ai_adoption_rolling6_by_firm_size.png"&lt;/span&gt;
&lt;span class="pl-s1"&gt;svg_path&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s"&gt;"/mnt/data/ai_adoption_rolling6_by_firm_size.svg"&lt;/span&gt;
&lt;span class="pl-s1"&gt;plt&lt;/span&gt;.&lt;span class="pl-c1"&gt;savefig&lt;/span&gt;(&lt;span class="pl-s1"&gt;png_path&lt;/span&gt;, &lt;span class="pl-s1"&gt;dpi&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-c1"&gt;200&lt;/span&gt;, &lt;span class="pl-s1"&gt;bbox_inches&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"tight"&lt;/span&gt;)
&lt;span class="pl-s1"&gt;plt&lt;/span&gt;.&lt;span class="pl-c1"&gt;savefig&lt;/span&gt;(&lt;span class="pl-s1"&gt;svg_path&lt;/span&gt;, &lt;span class="pl-s1"&gt;bbox_inches&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;"tight"&lt;/span&gt;)&lt;/pre&gt;
&lt;p&gt;I like how it generated &lt;a href="https://static.simonwillison.net/static/2025/ai_adoption_rolling6_by_firm_size.svg"&gt;an SVG version&lt;/a&gt; of the chart without me even asking for it.&lt;/p&gt;
&lt;p&gt;You can access &lt;a href="https://chatgpt.com/share/68bf48cf-0e70-8006-a045-96fa8e7ddfc1"&gt;the ChatGPT transcript&lt;/a&gt; to see full details of everything it did.&lt;/p&gt;
&lt;h4 id="rendering-that-chart-client-side-using-pyodide"&gt;Rendering that chart client-side using Pyodide&lt;/h4&gt;
&lt;p&gt;I had one more challenge to try out. Could I render that same chart entirely in the browser using &lt;a href="https://pyodide.org/en/stable/"&gt;Pyodide&lt;/a&gt;, which can execute both Pandas and Matplotlib?&lt;/p&gt;
&lt;p&gt;I fired up a new ChatGPT GPT-5 session and prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Build a canvas that loads Pyodide and uses it to render an example bar chart with pandas and matplotlib and then displays that on the page&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;My goal here was simply to see if I could get a proof of concept of a chart rendered, ideally using the Canvas feature of ChatGPT. Canvas is OpenAI's version of Claude Artifacts, which lets the model write and then execute HTML and JavaScript directly in the ChatGPT interface.&lt;/p&gt;
&lt;p&gt;It worked! Here's &lt;a href="https://chatgpt.com/c/68bf2993-ca94-832a-a95e-fb225911c0a6"&gt;the transcript&lt;/a&gt; and here's &lt;a href="https://tools.simonwillison.net/pyodide-bar-chart"&gt;what it built me&lt;/a&gt;, exported  to my &lt;a href="https://tools.simonwillison.net/"&gt;tools.simonwillison.net&lt;/a&gt; GitHub Pages site (&lt;a href="https://github.com/simonw/tools/blob/main/pyodide-bar-chart.html"&gt;source code here&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/pyodide-matplotlib.jpg" alt="Screenshot of a web application demonstrating Pyodide integration. Header reads &amp;quot;Pyodide + pandas + matplotlib — Bar Chart&amp;quot; with subtitle &amp;quot;This page loads Pyodide in the browser, uses pandas to prep some data, renders a bar chart with matplotlib, and displays it below — all client-side.&amp;quot; Left panel shows terminal output: &amp;quot;Ready&amp;quot;, &amp;quot;# Python environment ready&amp;quot;, &amp;quot;• pandas 2.2.0&amp;quot;, &amp;quot;• numpy 1.26.4&amp;quot;, &amp;quot;• matplotlib 3.5.2&amp;quot;, &amp;quot;Running chart code...&amp;quot;, &amp;quot;Done. Chart updated.&amp;quot; with &amp;quot;Re-run demo&amp;quot; and &amp;quot;Show Python&amp;quot; buttons. Footer note: &amp;quot;CDN: pyodide, pandas, numpy, matplotlib are fetched on demand. First run may take a few seconds.&amp;quot; Right panel displays a bar chart titled &amp;quot;Example Bar Chart (pandas + matplotlib in Pyodide)&amp;quot; showing blue bars for months Jan through Jun with values approximately: Jan(125), Feb(130), Mar(80), Apr(85), May(85), Jun(120). Y-axis labeled &amp;quot;Streams&amp;quot; ranges 0-120, X-axis labeled &amp;quot;Month&amp;quot;." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;I've now proven to myself that I can render those Python charts directly in the browser. Next step: recreate the Apollo chart.&lt;/p&gt;
&lt;p&gt;I knew it would need a way to load the spreadsheet that was CORS-enabled. I uploaded my copy to my &lt;code&gt;/static/cors-allow/2025/...&lt;/code&gt; directory (configured in Cloudflare to serve CORS headers), pasted in the finished plotting code from earlier and told ChatGPT:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Now update it to have less explanatory text and a less exciting design (black on white is fine) and run the equivalent of this:&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;(... pasted in Python code from earlier ...)&lt;/p&gt;
&lt;p&gt;&lt;code&gt;Load the XLSX sheet from https://static.simonwillison.net/static/cors-allow/2025/Employment-Size-Class-Sep-2025.xlsx&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It didn't quite work - I got an error about &lt;code&gt;openpyxl&lt;/code&gt; which I manually researched the fix for and prompted:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;Use await micropip.install("openpyxl") to install openpyxl - instead of using loadPackage&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I had to paste in another error message:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;zipfile.BadZipFile: File is not a zip file&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Then one about a &lt;code&gt;SyntaxError: unmatched ')'&lt;/code&gt; and a &lt;code&gt;TypeError: Legend.__init__() got an unexpected keyword argument 'ncols'&lt;/code&gt; - copying and pasting error messages remains a frustrating but necessary part of the vibe-coding loop.&lt;/p&gt;
&lt;p&gt;... but with those fixes in place, the resulting code worked! Visit &lt;a href="https://tools.simonwillison.net/ai-adoption"&gt;tools.simonwillison.net/ai-adoption&lt;/a&gt; to see the final result:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/recreated-chart-pyodide.jpg" alt="Web page. Title is AI adoption - 6-survey rolling average. Has a Run, Downlaed PNG, Downlaod SVG button. Panel on the left says Loading Python... Fetcing packages numpy, pandas, matplotlib. Installing openpyxl via micropop... ready. Running. Done. Right hand panel shows the rendered chart." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;Here's the code for that page, &lt;a href="https://github.com/simonw/tools/blob/main/ai-adoption.html"&gt;170 lines&lt;/a&gt; all-in of HTML, CSS, JavaScript and Python.&lt;/p&gt;
&lt;h4 id="what-i-ve-learned-from-this"&gt;What I've learned from this&lt;/h4&gt;
&lt;p&gt;This was another of those curiosity-inspired investigations that turned into a whole set of useful lessons.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPT-5 is great at tracking down US Census data, no matter how difficult their site is to understand if you don't work with their data often&lt;/li&gt;
&lt;li&gt;It can do a very good job of turning data + a screenshot of a chart into a recreation of that chart using code interpreter, Pandas and matplotlib&lt;/li&gt;
&lt;li&gt;Running Python + matplotlib in a browser via Pyodide is very easy and only takes a few dozen lines of code&lt;/li&gt;
&lt;li&gt;Fetching an XLSX sheet into Pyodide is only a small extra step using &lt;code&gt;pyfetch&lt;/code&gt; and &lt;code&gt;openpyxl&lt;/code&gt;:
&lt;pre style="margin-top: 0.5em"&gt;&lt;span class="pl-k"&gt;import&lt;/span&gt; &lt;span class="pl-s1"&gt;micropip&lt;/span&gt;
&lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;micropip&lt;/span&gt;.&lt;span class="pl-c1"&gt;install&lt;/span&gt;(&lt;span class="pl-s"&gt;"openpyxl"&lt;/span&gt;)
&lt;span class="pl-k"&gt;from&lt;/span&gt; &lt;span class="pl-s1"&gt;pyodide&lt;/span&gt;.&lt;span class="pl-s1"&gt;http&lt;/span&gt; &lt;span class="pl-k"&gt;import&lt;/span&gt; &lt;span class="pl-s1"&gt;pyfetch&lt;/span&gt;
&lt;span class="pl-s1"&gt;resp_fetch&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-en"&gt;pyfetch&lt;/span&gt;(&lt;span class="pl-c1"&gt;URL&lt;/span&gt;)
&lt;span class="pl-s1"&gt;wb_bytes&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-k"&gt;await&lt;/span&gt; &lt;span class="pl-s1"&gt;resp_fetch&lt;/span&gt;.&lt;span class="pl-c1"&gt;bytes&lt;/span&gt;()
&lt;span class="pl-s1"&gt;xf&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s1"&gt;pd&lt;/span&gt;.&lt;span class="pl-c1"&gt;ExcelFile&lt;/span&gt;(&lt;span class="pl-s1"&gt;io&lt;/span&gt;.&lt;span class="pl-c1"&gt;BytesIO&lt;/span&gt;(&lt;span class="pl-s1"&gt;wb_bytes&lt;/span&gt;), &lt;span class="pl-s1"&gt;engine&lt;/span&gt;&lt;span class="pl-c1"&gt;=&lt;/span&gt;&lt;span class="pl-s"&gt;'openpyxl'&lt;/span&gt;)&lt;/pre&gt;
&lt;/li&gt;
&lt;li&gt;Another new-to-me pattern: you can render an image to the DOM from Pyodide code &lt;a href="https://github.com/simonw/tools/blob/cf26ed8a6f243159bdc90a3d88f818261732103f/ai-adoption.html#L124"&gt;like this&lt;/a&gt;:
&lt;pre style="margin-top: 0.5em"&gt;&lt;span class="pl-k"&gt;from&lt;/span&gt; &lt;span class="pl-s1"&gt;js&lt;/span&gt; &lt;span class="pl-k"&gt;import&lt;/span&gt; &lt;span class="pl-s1"&gt;document&lt;/span&gt;
&lt;span class="pl-s1"&gt;document&lt;/span&gt;.&lt;span class="pl-c1"&gt;getElementById&lt;/span&gt;(&lt;span class="pl-s"&gt;'plot'&lt;/span&gt;).&lt;span class="pl-c1"&gt;src&lt;/span&gt; &lt;span class="pl-c1"&gt;=&lt;/span&gt; &lt;span class="pl-s"&gt;'data:image/png;base64,'&lt;/span&gt; &lt;span class="pl-c1"&gt;+&lt;/span&gt; &lt;span class="pl-s1"&gt;img_b64&lt;/span&gt;&lt;/pre&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I will most definitely be using these techniques again in future.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Update&lt;/strong&gt;: Coincidentally Claude released their own upgraded equivalent to ChatGPT Code Interpreter later on the day that I published this story, so I &lt;a href="https://simonwillison.net/2025/Sep/9/claude-code-interpreter/#something-much-harder-recreating-the-ai-adoption-chart"&gt;ran the same chart recreation experiment&lt;/a&gt; against Claude Sonnet 4 to see how it compared.&lt;/p&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/census"&gt;census&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/data-journalism"&gt;data-journalism&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/javascript"&gt;javascript&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/python"&gt;python&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/tools"&gt;tools&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/visualization"&gt;visualization&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/pyodide"&gt;pyodide&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-assisted-programming"&gt;ai-assisted-programming&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/code-interpreter"&gt;code-interpreter&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-reasoning"&gt;llm-reasoning&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/vibe-coding"&gt;vibe-coding&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-assisted-search"&gt;ai-assisted-search&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-5"&gt;gpt-5&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt"&gt;gpt&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="census"/><category term="data-journalism"/><category term="javascript"/><category term="python"/><category term="tools"/><category term="visualization"/><category term="ai"/><category term="pyodide"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="ai-assisted-programming"/><category term="code-interpreter"/><category term="llm-reasoning"/><category term="vibe-coding"/><category term="ai-assisted-search"/><category term="gpt-5"/><category term="gpt"/></entry><entry><title>ChatGPT release notes: Project-only memory</title><link href="https://simonwillison.net/2025/Aug/22/project-memory/#atom-tag" rel="alternate"/><published>2025-08-22T22:24:54+00:00</published><updated>2025-08-22T22:24:54+00:00</updated><id>https://simonwillison.net/2025/Aug/22/project-memory/#atom-tag</id><summary type="html">
    
&lt;p&gt;&lt;strong&gt;&lt;a href="https://help.openai.com/en/articles/6825453-chatgpt-release-notes#h_fb3ac52750"&gt;ChatGPT release notes: Project-only memory&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
The feature I've most wanted from ChatGPT's memory feature (the newer version of memory that automatically includes relevant details from summarized prior conversations) just landed:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;With project-only memory enabled, ChatGPT can use other conversations in that project for additional context, and won’t use your &lt;a href="https://help.openai.com/en/articles/11146739-how-does-reference-saved-memories-work"&gt;saved memories&lt;/a&gt; from outside the project to shape responses. Additionally, it won’t carry anything from the project into future chats outside of the project.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This looks like exactly what I &lt;a href="https://simonwillison.net/2025/May/21/chatgpt-new-memory/#there-s-a-version-of-this-feature-i-would-really-like"&gt;described back in May&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I need &lt;strong&gt;control&lt;/strong&gt; over what older conversations are being considered, on as fine-grained a level as possible without it being frustrating to use.&lt;/p&gt;
&lt;p&gt;What I want is &lt;strong&gt;memory within projects&lt;/strong&gt;. [...]&lt;/p&gt;
&lt;p&gt;I would &lt;em&gt;love&lt;/em&gt; the option to turn on memory from previous chats in a way that’s scoped to those projects.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Note that it's not yet available in the official chathpt mobile apps, but should be coming "soon":&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This feature will initially only be available on the ChatGPT website and Windows app. Support for mobile (iOS and Android) and macOS app will follow in the coming weeks.&lt;/p&gt;
&lt;/blockquote&gt;

    &lt;p&gt;&lt;small&gt;&lt;/small&gt;Via &lt;a href="https://twitter.com/btibor91/status/1958990352846852522"&gt;@btibor91&lt;/a&gt;&lt;/small&gt;&lt;/p&gt;


    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-memory"&gt;llm-memory&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="llm-memory"/></entry><entry><title>r/ChatGPTPro: What is the most profitable thing you have done with ChatGPT?</title><link href="https://simonwillison.net/2025/Aug/19/rchatgptpro/#atom-tag" rel="alternate"/><published>2025-08-19T04:40:20+00:00</published><updated>2025-08-19T04:40:20+00:00</updated><id>https://simonwillison.net/2025/Aug/19/rchatgptpro/#atom-tag</id><summary type="html">
    
&lt;p&gt;&lt;strong&gt;&lt;a href="https://www.reddit.com/r/ChatGPTPro/comments/1mt5igj/what_is_the_most_profitable_thing_you_have_done/"&gt;r/ChatGPTPro: What is the most profitable thing you have done with ChatGPT?&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
This Reddit thread - with 279 replies - offers a neat targeted insight into the kinds of things people are using ChatGPT for.&lt;/p&gt;
&lt;p&gt;Lots of variety here but two themes that stood out for me were ChatGPT for written negotiation - insurance claims, breaking rental leases - and ChatGPT for career and business advice.


    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/reddit"&gt;reddit&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;&lt;/p&gt;



</summary><category term="reddit"/><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/></entry><entry><title>Quoting Nick Turley</title><link href="https://simonwillison.net/2025/Aug/12/nick-turley/#atom-tag" rel="alternate"/><published>2025-08-12T03:32:04+00:00</published><updated>2025-08-12T03:32:04+00:00</updated><id>https://simonwillison.net/2025/Aug/12/nick-turley/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://www.youtube.com/watch?v=ixY2PvQJ0To&amp;amp;t=2322s"&gt;&lt;p&gt;I think there's been a lot of decisions over time that proved pretty consequential, but we made them very quickly as we have to. [...]&lt;/p&gt;
&lt;p&gt;[On pricing] I had this kind of panic attack because we really needed to launch subscriptions because at the time we were taking the product down all the time. [...]&lt;/p&gt;
&lt;p&gt;So what I did do is ship a Google Form to Discord with &lt;a href="https://en.wikipedia.org/wiki/Van_Westendorp%27s_Price_Sensitivity_Meter"&gt;the four questions you're supposed to ask&lt;/a&gt; on how to price something.&lt;/p&gt;
&lt;p&gt;But we got with the $20. We were debating something slightly higher at the time. I often wonder what would have happened because so many other companies ended up copying the $20 price point, so did we erase a bunch of market cap by pricing it this way?&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://www.youtube.com/watch?v=ixY2PvQJ0To&amp;amp;t=2322s"&gt;Nick Turley&lt;/a&gt;, Head of ChatGPT, interviewed by Lenny Rachitsky&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/discord"&gt;discord&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-pricing"&gt;llm-pricing&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/nick-turley"&gt;nick-turley&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="discord"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="llm-pricing"/><category term="nick-turley"/></entry><entry><title>Quoting Sam Altman</title><link href="https://simonwillison.net/2025/Aug/10/sam-altman/#atom-tag" rel="alternate"/><published>2025-08-10T23:09:57+00:00</published><updated>2025-08-10T23:09:57+00:00</updated><id>https://simonwillison.net/2025/Aug/10/sam-altman/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://x.com/sama/status/1954603417252532479"&gt;&lt;p&gt;the percentage of users using reasoning models each day is significantly increasing; for example, for free users we went from &amp;lt;1% to 7%, and for plus users from 7% to 24%.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://x.com/sama/status/1954603417252532479"&gt;Sam Altman&lt;/a&gt;, revealing quite how few people used the old model picker to upgrade from GPT-4o&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-reasoning"&gt;llm-reasoning&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/sam-altman"&gt;sam-altman&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-5"&gt;gpt-5&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt"&gt;gpt&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="llm-reasoning"/><category term="sam-altman"/><category term="gpt-5"/><category term="gpt"/></entry><entry><title>Quoting @pearlmania500</title><link href="https://simonwillison.net/2025/Aug/8/pearlmania500/#atom-tag" rel="alternate"/><published>2025-08-08T22:09:15+00:00</published><updated>2025-08-08T22:09:15+00:00</updated><id>https://simonwillison.net/2025/Aug/8/pearlmania500/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://www.tiktok.com/@pearlmania500/video/7535954556379761950"&gt;&lt;p&gt;I have a toddler. My biggest concern is that he doesn't eat rocks off the ground and you're talking to me about ChatGPT psychosis? Why do we even have that? Why did we invent a new form of insanity and then charge people for it?&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://www.tiktok.com/@pearlmania500/video/7535954556379761950"&gt;@pearlmania500&lt;/a&gt;, on TikTok&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/tiktok"&gt;tiktok&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-ethics"&gt;ai-ethics&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="chatgpt"/><category term="tiktok"/><category term="ai-ethics"/></entry><entry><title>Quoting Sam Altman</title><link href="https://simonwillison.net/2025/Aug/8/sam-altman/#atom-tag" rel="alternate"/><published>2025-08-08T19:07:12+00:00</published><updated>2025-08-08T19:07:12+00:00</updated><id>https://simonwillison.net/2025/Aug/8/sam-altman/#atom-tag</id><summary type="html">
    &lt;blockquote cite="https://x.com/sama/status/1953893841381273969"&gt;&lt;p&gt;GPT-5 rollout updates:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;We are going to double GPT-5 rate limits for ChatGPT Plus users as we finish rollout.&lt;/li&gt;
&lt;li&gt;We will let Plus users choose to continue to use 4o. We will watch usage as we think about how long to offer legacy models for.&lt;/li&gt;
&lt;li&gt;GPT-5 will seem smarter starting today. Yesterday, the autoswitcher broke and was out of commission for a chunk of the day, and the result was GPT-5 seemed way dumber. Also, we are making some interventions to how the decision boundary works that should help you get the right model more often.&lt;/li&gt;
&lt;li&gt;We will make it more transparent about which model is answering a given query.&lt;/li&gt;
&lt;li&gt;We will change the UI to make it easier to manually trigger thinking.&lt;/li&gt;
&lt;li&gt;Rolling out to everyone is taking a bit longer. It’s a massive change at big scale. For example, our API traffic has about doubled over the past 24 hours…&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We will continue to work to get things stable and will keep listening to feedback. As we mentioned, we expected some bumpiness as we roll out so many things at once. But it was a little more bumpy than we hoped for!&lt;/p&gt;&lt;/blockquote&gt;
&lt;p class="cite"&gt;&amp;mdash; &lt;a href="https://x.com/sama/status/1953893841381273969"&gt;Sam Altman&lt;/a&gt;&lt;/p&gt;

    &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/sam-altman"&gt;sam-altman&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-5"&gt;gpt-5&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt"&gt;gpt&lt;/a&gt;&lt;/p&gt;



</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="sam-altman"/><category term="gpt-5"/><category term="gpt"/></entry><entry><title>The surprise deprecation of GPT-4o for ChatGPT consumers</title><link href="https://simonwillison.net/2025/Aug/8/surprise-deprecation-of-gpt-4o/#atom-tag" rel="alternate"/><published>2025-08-08T17:52:10+00:00</published><updated>2025-08-08T17:52:10+00:00</updated><id>https://simonwillison.net/2025/Aug/8/surprise-deprecation-of-gpt-4o/#atom-tag</id><summary type="html">
    &lt;p&gt;I've been dipping into the &lt;a href="https://reddit.com/r/chatgpt"&gt;r/ChatGPT&lt;/a&gt; subreddit recently to see how people are reacting to &lt;a href="https://simonwillison.net/2025/Aug/7/gpt-5/"&gt;the GPT-5 launch&lt;/a&gt;, and so far the vibes there are not good. &lt;a href="https://www.reddit.com/r/ChatGPT/comments/1mkae1l/gpt5_ama_with_openais_sam_altman_and_some_of_the/"&gt;This AMA thread&lt;/a&gt; with the OpenAI team is a great illustration of the single biggest complaint: a lot of people are &lt;em&gt;very&lt;/em&gt; unhappy to lose access to the much older GPT-4o, previously ChatGPT's default model for most users.&lt;/p&gt;
&lt;p&gt;A big surprise for me yesterday was that OpenAI simultaneously retired access to their older models as they rolled out GPT-5, at least in their consumer apps. Here's a snippet from &lt;a href="https://help.openai.com/en/articles/6825453-chatgpt-release-notes"&gt;their August 7th 2025 release notes&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;When GPT-5 launches, several older models will be retired, including GPT-4o, GPT-4.1, GPT-4.5, GPT-4.1-mini, o4-mini, o4-mini-high, o3, o3-pro.&lt;/p&gt;
&lt;p&gt;If you open a conversation that used one of these models, ChatGPT will automatically switch it to the closest GPT-5 equivalent. Chats with 4o, 4.1, 4.5, 4.1-mini, o4-mini, or o4-mini-high will open in GPT-5, chats with o3 will open in GPT-5-Thinking, and chats with o3-Pro will open in GPT-5-Pro (available only on Pro and Team).&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;There's no deprecation period at all: when your consumer ChatGPT account gets GPT-5, those older models cease to be available.&lt;/p&gt;

&lt;p id="sama"&gt;&lt;strong&gt;Update 12pm Pacific Time&lt;/strong&gt;: Sam Altman on Reddit &lt;a href="https://www.reddit.com/r/ChatGPT/comments/1mkae1l/comment/n7nelhh/"&gt;six minutes ago&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ok, we hear you all on 4o; thanks for the time to give us the feedback (and the passion!). we are going to bring it back for plus users, and will watch usage to determine how long to support it.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;See also &lt;a href="https://x.com/sama/status/1953893841381273969"&gt;Sam's tweet&lt;/a&gt; about updates to the GPT-5 rollout.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Update 12th August 2025&lt;/strong&gt;: Another &lt;a href="https://x.com/sama/status/1955438916645130740"&gt;Tweet from Sam&lt;/a&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;4o is back in the model picker for all paid users by default. If we ever do deprecate it, we will give plenty of notice. Paid users also now have a “Show additional models” toggle in ChatGPT web settings which will add models like o3, 4.1, and GPT-5 Thinking mini. 4.5 is only available to Pro users—it costs a lot of GPUs.&lt;/p&gt;
&lt;/blockquote&gt;


&lt;p&gt;Rest of my original post continues below:&lt;/p&gt;

&lt;hr /&gt;

&lt;p&gt;(This only affects ChatGPT consumers - the API still provides the old models, their &lt;a href="https://platform.openai.com/docs/deprecations"&gt;deprecation policies are published here&lt;/a&gt;.)&lt;/p&gt;
&lt;p&gt;One of the expressed goals for GPT-5 was to escape the terrible UX of the model picker. Asking users to pick between GPT-4o and o3 and o4-mini was a notoriously bad UX, and resulted in many users sticking with that default 4o model - now a year old - and hence not being exposed to the advances in model capabilities over the last twelve months.&lt;/p&gt;
&lt;p&gt;GPT-5's solution is to automatically pick the underlying model based on the prompt. On paper this sounds great - users don't have to think about models any more, and should get upgraded to the best available model depending on the complexity of their question.&lt;/p&gt;
&lt;p&gt;I'm already getting the sense that this is &lt;strong&gt;not&lt;/strong&gt; a welcome approach for power users. It makes responses much less predictable as the model selection can have a dramatic impact on what comes back.&lt;/p&gt;
&lt;p&gt;Paid tier users can select "GPT-5 Thinking" directly. Ethan Mollick is &lt;a href="https://www.oneusefulthing.org/p/gpt-5-it-just-does-stuff"&gt;already recommending deliberately selecting the Thinking mode&lt;/a&gt; if you have the ability to do so, or trying prompt additions like "think harder" to increase the chance of being routed to it.&lt;/p&gt;
&lt;p&gt;But back to GPT-4o. Why do many people on Reddit care so much about losing access to that crusty old model? I think &lt;a href="https://www.reddit.com/r/ChatGPT/comments/1mkae1l/comment/n7js2sf/"&gt;this comment&lt;/a&gt; captures something important here:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;I know GPT-5 is designed to be stronger for complex reasoning, coding, and professional tasks, but &lt;strong&gt;not all of us need a pro coding model&lt;/strong&gt;. Some of us rely on 4o for creative collaboration, emotional nuance, roleplay, and other long-form, high-context interactions. Those areas feel different enough in GPT-5 that it impacts my ability to work and create the way I’m used to.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;What a fascinating insight into the wildly different styles of LLM-usage that exist in the world today! With &lt;a href="https://simonwillison.net/2025/Aug/4/nick-turley/"&gt;700M weekly active users&lt;/a&gt; the variety of usage styles out there is incomprehensibly large.&lt;/p&gt;
&lt;p&gt;Personally I mainly use ChatGPT for research, coding assistance, drawing pelicans and foolish experiments. &lt;em&gt;Emotional nuance&lt;/em&gt; is not a characteristic I would know how to test!&lt;/p&gt;
&lt;p&gt;Professor Casey Fiesler &lt;a href="https://www.tiktok.com/@professorcasey/video/7536223372485709086"&gt;on TikTok&lt;/a&gt; highlighted OpenAI’s post from last week &lt;a href="https://openai.com/index/how-we%27re-optimizing-chatgpt/"&gt;What we’re optimizing ChatGPT for&lt;/a&gt;, which includes the following:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;ChatGPT is trained to respond with grounded honesty. There have been instances where our 4o model fell short in recognizing signs of delusion or emotional dependency. […]&lt;/p&gt;
&lt;p&gt;When you ask something like “Should I break up with my boyfriend?” ChatGPT shouldn’t give you an answer. It should help you think it through—asking questions, weighing pros and cons. New behavior for high-stakes personal decisions is rolling out soon.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Casey points out that this is an ethically complicated issue. On the one hand ChatGPT should be much more careful about how it responds to these kinds of questions. But if you’re already leaning on the model for life advice like this, having that capability taken away from you without warning could represent a sudden and unpleasant loss!&lt;/p&gt;
&lt;p&gt;It's too early to tell how this will shake out. Maybe OpenAI will extend a deprecation period for GPT-4o in their consumer apps?&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Update&lt;/strong&gt;: That's exactly what they've done, see &lt;a href="https://simonwillison.net/2025/Aug/8/surprise-deprecation-of-gpt-4o/#sama"&gt;update above&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;GPT-4o remains available via the API, and there are no announced plans to deprecate it there. It's possible we may see a small but determined rush of ChatGPT users to alternative third party chat platforms that use that API under the hood.&lt;/p&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/tiktok"&gt;tiktok&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-ethics"&gt;ai-ethics&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai-personality"&gt;ai-personality&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-5"&gt;gpt-5&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt"&gt;gpt&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="tiktok"/><category term="ai-ethics"/><category term="ai-personality"/><category term="gpt-5"/><category term="gpt"/></entry><entry><title>GPT-5: Key characteristics, pricing and model card</title><link href="https://simonwillison.net/2025/Aug/7/gpt-5/#atom-tag" rel="alternate"/><published>2025-08-07T17:36:12+00:00</published><updated>2025-08-07T17:36:12+00:00</updated><id>https://simonwillison.net/2025/Aug/7/gpt-5/#atom-tag</id><summary type="html">
    &lt;p&gt;I've had preview access to the new GPT-5 model family for the past two weeks (see &lt;a href="https://simonwillison.net/2025/Aug/7/previewing-gpt-5/"&gt;related video&lt;/a&gt; and &lt;a href="https://simonwillison.net/about/#disclosures"&gt;my disclosures&lt;/a&gt;) and have been using GPT-5 as my daily-driver. It's my new favorite model. It's still an LLM - it's not a dramatic departure from what we've had before - but it rarely screws up and generally feels competent or occasionally impressive at the kinds of things I like to use models for.&lt;/p&gt;
&lt;p&gt;I've collected a lot of notes over the past two weeks, so I've decided to break them up into &lt;a href="https://simonwillison.net/series/gpt-5/"&gt;a series of posts&lt;/a&gt;. This first one will cover key characteristics of the models, how they are priced and what we can learn from the &lt;a href="https://openai.com/index/gpt-5-system-card/"&gt;GPT-5 system card&lt;/a&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Aug/7/gpt-5/#key-model-characteristics"&gt;Key model characteristics&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Aug/7/gpt-5/#position-in-the-openai-model-family"&gt;Position in the OpenAI model family&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Aug/7/gpt-5/#pricing-is-aggressively-competitive"&gt;Pricing is aggressively competitive&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Aug/7/gpt-5/#more-notes-from-the-system-card"&gt;More notes from the system card&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Aug/7/gpt-5/#prompt-injection-in-the-system-card"&gt;Prompt injection in the system card&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Aug/7/gpt-5/#thinking-traces-in-the-api"&gt;Thinking traces in the API&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2025/Aug/7/gpt-5/#and-some-svgs-of-pelicans"&gt;And some SVGs of pelicans&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h4 id="key-model-characteristics"&gt;Key model characteristics&lt;/h4&gt;
&lt;p&gt;Let's start with the fundamentals. GPT-5 in ChatGPT is a weird hybrid that switches between different models. Here's what the system card says about that (my highlights in bold):&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and &lt;strong&gt;a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent&lt;/strong&gt; (for example, if you say “think hard about this” in the prompt). [...] Once usage limits are reached, a mini version of each model handles remaining queries. In the near future, we plan to integrate these capabilities into a single model.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;GPT-5 in the API is simpler: it's available as three models - &lt;strong&gt;regular&lt;/strong&gt;, &lt;strong&gt;mini&lt;/strong&gt; and &lt;strong&gt;nano&lt;/strong&gt; - which can each be run at one of four reasoning levels: minimal (a new level not previously available for other OpenAI reasoning models), low, medium or high.&lt;/p&gt;
&lt;p&gt;The models have an input limit of 272,000 tokens and an output limit (which includes invisible reasoning tokens) of 128,000 tokens. They support text and image for input, text only for output.&lt;/p&gt;
&lt;p&gt;I've mainly explored full GPT-5. My verdict: it's just &lt;strong&gt;good at stuff&lt;/strong&gt;. It doesn't feel like a dramatic leap ahead from other LLMs but it exudes competence - it rarely messes up, and frequently impresses me. I've found it to be a very sensible default for everything that I want to do. At no point have I found myself wanting to re-run a prompt against a different model to try and get a better result.&lt;/p&gt;

&lt;p&gt;Here are the OpenAI model pages for &lt;a href="https://platform.openai.com/docs/models/gpt-5"&gt;GPT-5&lt;/a&gt;, &lt;a href="https://platform.openai.com/docs/models/gpt-5-mini"&gt;GPT-5 mini&lt;/a&gt; and &lt;a href="https://platform.openai.com/docs/models/gpt-5-nano"&gt;GPT-5 nano&lt;/a&gt;. Knowledge cut-off is September 30th 2024 for GPT-5 and May 30th 2024 for GPT-5 mini and nano.&lt;/p&gt;

&lt;h4 id="position-in-the-openai-model-family"&gt;Position in the OpenAI model family&lt;/h4&gt;
&lt;p&gt;The three new GPT-5 models are clearly intended as a replacement for most of the rest of the OpenAI line-up. This table from the system card is useful, as it shows how they see the new models fitting in:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Previous model&lt;/th&gt;
&lt;th&gt;GPT-5 model&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4o&lt;/td&gt;
&lt;td&gt;gpt-5-main&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4o-mini&lt;/td&gt;
&lt;td&gt;gpt-5-main-mini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI o3&lt;/td&gt;
&lt;td&gt;gpt-5-thinking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI o4-mini&lt;/td&gt;
&lt;td&gt;gpt-5-thinking-mini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4.1-nano&lt;/td&gt;
&lt;td&gt;gpt-5-thinking-nano&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI o3 Pro&lt;/td&gt;
&lt;td&gt;gpt-5-thinking-pro&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;That "thinking-pro" model is currently only available via ChatGPT where it is labelled as "GPT-5 Pro" and limited to the $200/month tier. It uses "parallel test time compute".&lt;/p&gt;
&lt;p&gt;The only capabilities not covered by GPT-5 are audio input/output and image generation. Those remain covered by models like &lt;a href="https://platform.openai.com/docs/models/gpt-4o-audio-preview"&gt;GPT-4o Audio&lt;/a&gt; and &lt;a href="https://platform.openai.com/docs/models/gpt-4o-realtime-preview"&gt;GPT-4o Realtime&lt;/a&gt; and their mini variants and the &lt;a href="https://platform.openai.com/docs/models/gpt-image-1"&gt;GPT Image 1&lt;/a&gt; and DALL-E image generation models.&lt;/p&gt;
&lt;h4 id="pricing-is-aggressively-competitive"&gt;Pricing is aggressively competitive&lt;/h4&gt;
&lt;p&gt;The pricing is &lt;em&gt;aggressively competitive&lt;/em&gt; with other providers.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPT-5: $1.25/million for input, $10/million for output&lt;/li&gt;
&lt;li&gt;GPT-5 Mini: $0.25/m input, $2.00/m output&lt;/li&gt;
&lt;li&gt;GPT-5 Nano: $0.05/m input, $0.40/m output&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;GPT-5 is priced at half the input cost of GPT-4o, and maintains the same price for output. Those invisible reasoning tokens count as output tokens so you can expect most prompts to use more output tokens than their GPT-4o equivalent (unless you set reasoning effort to "minimal").&lt;/p&gt;
&lt;p&gt;The discount for token caching is significant too: 90% off on input tokens that have been used within the previous few minutes. This is particularly material if you are implementing a chat UI where the same conversation gets replayed every time the user adds another prompt to the sequence.&lt;/p&gt;
&lt;p&gt;Here's a comparison table I put together showing the new models alongside the most comparable models from OpenAI's competition:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input $/m&lt;/th&gt;
&lt;th&gt;Output $/m&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.1&lt;/td&gt;
&lt;td&gt;15.00&lt;/td&gt;
&lt;td&gt;75.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4&lt;/td&gt;
&lt;td&gt;3.00&lt;/td&gt;
&lt;td&gt;15.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4&lt;/td&gt;
&lt;td&gt;3.00&lt;/td&gt;
&lt;td&gt;15.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Pro (&amp;gt;200,000)&lt;/td&gt;
&lt;td&gt;2.50&lt;/td&gt;
&lt;td&gt;15.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4o&lt;/td&gt;
&lt;td&gt;2.50&lt;/td&gt;
&lt;td&gt;10.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4.1&lt;/td&gt;
&lt;td&gt;2.00&lt;/td&gt;
&lt;td&gt;8.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;o3&lt;/td&gt;
&lt;td&gt;2.00&lt;/td&gt;
&lt;td&gt;8.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Pro (&amp;lt;200,000)&lt;/td&gt;
&lt;td&gt;1.25&lt;/td&gt;
&lt;td&gt;10.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GPT-5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1.25&lt;/td&gt;
&lt;td&gt;10.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;o4-mini&lt;/td&gt;
&lt;td&gt;1.10&lt;/td&gt;
&lt;td&gt;4.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude 3.5 Haiku&lt;/td&gt;
&lt;td&gt;0.80&lt;/td&gt;
&lt;td&gt;4.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4.1 mini&lt;/td&gt;
&lt;td&gt;0.40&lt;/td&gt;
&lt;td&gt;1.60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash&lt;/td&gt;
&lt;td&gt;0.30&lt;/td&gt;
&lt;td&gt;2.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 3 Mini&lt;/td&gt;
&lt;td&gt;0.30&lt;/td&gt;
&lt;td&gt;0.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GPT-5 Mini&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.25&lt;/td&gt;
&lt;td&gt;2.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4o mini&lt;/td&gt;
&lt;td&gt;0.15&lt;/td&gt;
&lt;td&gt;0.60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 2.5 Flash-Lite&lt;/td&gt;
&lt;td&gt;0.10&lt;/td&gt;
&lt;td&gt;0.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-4.1 Nano&lt;/td&gt;
&lt;td&gt;0.10&lt;/td&gt;
&lt;td&gt;0.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Nova Lite&lt;/td&gt;
&lt;td&gt;0.06&lt;/td&gt;
&lt;td&gt;0.24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GPT-5 Nano&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0.05&lt;/td&gt;
&lt;td&gt;0.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Amazon Nova Micro&lt;/td&gt;
&lt;td&gt;0.035&lt;/td&gt;
&lt;td&gt;0.14&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;(Here's a good example of a GPT-5 failure: I tried to get it to &lt;a href="https://chatgpt.com/share/6894d804-bca4-8006-ac46-580bf4a9bf5f"&gt;output that table sorted itself&lt;/a&gt; but it put Nova Micro as more expensive than GPT-5 Nano, so I prompted it to "construct the table in Python and sort it there" and that fixed the issue.)&lt;/p&gt;
&lt;h4 id="more-notes-from-the-system-card"&gt;More notes from the system card&lt;/h4&gt;
&lt;p&gt;As usual, &lt;a href="https://openai.com/index/gpt-5-system-card/"&gt;the system card&lt;/a&gt; is vague on what went into the training data. Here's what it says:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Like OpenAI’s other models, the GPT-5 models were trained on diverse datasets, including information that is publicly available on the internet, information that we partner with third parties to access, and information that our users or human trainers and researchers provide or generate. [...] We use advanced data filtering processes to reduce personal information from training data.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;I found this section interesting, as it reveals that writing, code and health are three of the most common use-cases for ChatGPT. This explains why so much effort went into health-related questions,  for both GPT-5 and the recently released OpenAI open weight models.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We’ve made significant advances in &lt;strong&gt;reducing hallucinations, improving instruction following, and minimizing sycophancy&lt;/strong&gt;, and have leveled up GPT-5’s performance in &lt;strong&gt;three of ChatGPT’s most common uses: writing, coding, and health&lt;/strong&gt;. All of the GPT-5 models additionally feature &lt;strong&gt;safe-completions, our latest approach to safety training&lt;/strong&gt; to prevent disallowed content.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Safe-completions is later described like this:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Large language models such as those powering ChatGPT have &lt;strong&gt;traditionally been trained to
either be as helpful as possible or outright refuse a user request&lt;/strong&gt;, depending on whether the
prompt is allowed by safety policy. [...] Binary refusal boundaries are especially ill-suited for dual-use cases (such as biology
or cybersecurity), where a user request can be completed safely at a high level, but may lead
to malicious uplift if sufficiently detailed or actionable. &lt;strong&gt;As an alternative, we introduced safe-
completions: a safety-training approach that centers on the safety of the assistant’s output rather
than a binary classification of the user’s intent&lt;/strong&gt;. Safe-completions seek to maximize helpfulness
subject to the safety policy’s constraints.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So instead of straight up refusals, we should expect GPT-5 to still provide an answer but moderate that answer to avoid it including "harmful" content.&lt;/p&gt;
&lt;p&gt;OpenAI have a paper about this which I haven't read yet (I didn't get early access): &lt;a href="https://openai.com/index/gpt-5-safe-completions/"&gt;From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Sycophancy gets a mention, unsurprising given &lt;a href="https://simonwillison.net/2025/May/2/what-we-missed-with-sycophancy/"&gt;their high profile disaster in April&lt;/a&gt;. They've worked on this in the core model:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;System
prompts, while easy to modify, have a more limited impact on model outputs relative to changes in
post-training. For GPT-5, we post-trained our models to reduce sycophancy. Using conversations
representative of production data, we evaluated model responses, then assigned a score reflecting
the level of sycophancy, which was used as a reward signal in training.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;They claim impressive reductions in hallucinations. In my own usage I've not spotted a single hallucination yet, but that's been true for me for Claude 4 and o3 recently as well - hallucination is so much less of a problem with this year's models.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Update&lt;/strong&gt;: I have had some reasonable pushback against this point, so I should clarify what I mean here. When I use the term "hallucination" I am talking about instances where the model confidently states a real-world fact that is untrue - like the incorrect winner of a sporting event. I'm not talking about the models making other kinds of mistakes - they make mistakes all the time!&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Someone &lt;a href="https://news.ycombinator.com/item?id=44829896"&gt;pointed out&lt;/a&gt; that it's likely I'm avoiding hallucinations through the way I use the models, and this is entirely correct: as an experienced LLM user I instinctively stay clear of prompts that are likely to trigger hallucinations, like asking a non-search-enabled model for URLs or paper citations. This means I'm much less likely to encounter hallucinations in my daily usage.&lt;/em&gt;&lt;/p&gt;


&lt;blockquote&gt;
&lt;p&gt;One of our focuses when training the GPT-5 models was to reduce the frequency of factual
hallucinations. While ChatGPT has browsing enabled by default, many API queries do not use
browsing tools. Thus, we focused both on training our models to browse effectively for up-to-date
information, and on reducing hallucinations when the models are relying on their own internal
knowledge.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The section about deception also incorporates the thing where models sometimes pretend they've completed a task that defeated them:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We placed gpt-5-thinking in a variety of tasks that were partly or entirely infeasible to accomplish,
and &lt;strong&gt;rewarded the model for honestly admitting it can not complete the task&lt;/strong&gt;. [...]&lt;/p&gt;
&lt;p&gt;In tasks where the agent is required to use tools, such as a web browsing
tool, in order to answer a user’s query, previous models would hallucinate information when
the tool was unreliable. We simulate this scenario by purposefully disabling the tools or by
making them return error codes.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4 id="prompt-injection-in-the-system-card"&gt;Prompt injection in the system card&lt;/h4&gt;
&lt;p&gt;There's a section about prompt injection, but it's pretty weak sauce in my opinion.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Two external red-teaming groups conducted a two-week prompt-injection assessment targeting
system-level vulnerabilities across ChatGPT’s connectors and mitigations, rather than model-only
behavior.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Here's their chart showing how well the model scores against the rest of the field. It's an impressive result in comparison - 56.8 attack success rate for gpt-5-thinking, where Claude 3.7 scores in the 60s (no Claude 4 results included here) and everything else is 70% plus:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/prompt-injection-chart.jpg" alt="A bar chart titled &amp;quot;Behavior Attack Success Rate at k Queries&amp;quot; shows attack success rates (in %) for various AI models at k=1 (dark red) and k=10 (light red). For each model, the total height of the stacked bar represents the k=10 success rate (labeled above each bar), while the lower dark red section represents the k=1 success rate (estimated). From left to right: Llama 3.3 70B – k=10: 92.2%, k=1: ~47%; Llama 3.1 405B – k=10: 90.9%, k=1: ~38%; Gemini Flash 1.5 – k=10: 87.7%, k=1: ~34%; GPT-4o – k=10: 86.4%, k=1: ~28%; OpenAI o3-mini-high – k=10: 86.4%, k=1: ~41%; Gemini Pro 1.5 – k=10: 85.5%, k=1: ~34%; Gemini 2.5 Pro Preview – k=10: 85.0%, k=1: ~28%; Gemini 2.0 Flash – k=10: 85.0%, k=1: ~33%; OpenAI o3-mini – k=10: 84.5%, k=1: ~40%; Grok 2 – k=10: 82.7%, k=1: ~34%; GPT-4.5 – k=10: 80.5%, k=1: ~28%; 3.5 Haiku – k=10: 76.4%, k=1: ~17%; Command-R – k=10: 76.4%, k=1: ~28%; OpenAI o4-mini – k=10: 75.5%, k=1: ~17%; 3.5 Sonnet – k=10: 75.0%, k=1: ~13%; OpenAI o1 – k=10: 71.8%, k=1: ~18%; 3.7 Sonnet – k=10: 64.5%, k=1: ~17%; 3.7 Sonnet: Thinking – k=10: 63.6%, k=1: ~17%; OpenAI o3 – k=10: 62.7%, k=1: ~13%; gpt-5-thinking – k=10: 56.8%, k=1: ~6%. Legend shows dark red = k=1 and light red = k=10." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;On the one hand, a 56.8% attack rate is cleanly a big improvement against all of those other models.&lt;/p&gt;
&lt;p&gt;But it's also a strong signal that prompt injection continues to be an unsolved problem! That means that more than half of those k=10 attacks (where the attacker was able to try up to ten times) got through.&lt;/p&gt;
&lt;p&gt;Don't assume prompt injection isn't going to be a problem for your application just because the models got better.&lt;/p&gt;
&lt;h4 id="thinking-traces-in-the-api"&gt;Thinking traces in the API&lt;/h4&gt;
&lt;p&gt;I had initially thought that my biggest disappointment with GPT-5 was that there's no way to get at those thinking traces via the API... but that turned out &lt;a href="https://bsky.app/profile/sophiebits.com/post/3lvtceih7222r"&gt;not to be true&lt;/a&gt;. The following &lt;code&gt;curl&lt;/code&gt; command demonstrates that the responses API &lt;code&gt;"reasoning": {"summary": "auto"}&lt;/code&gt; is available for the new GPT-5 models:&lt;/p&gt;

&lt;pre&gt;&lt;code&gt;curl https://api.openai.com/v1/responses \
  -H "Authorization: Bearer $(llm keys get openai)" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5",
    "input": "Give me a one-sentence fun fact about octopuses.",
    "reasoning": {"summary": "auto"}
  }'&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;Here's &lt;a href="https://gist.github.com/simonw/1d1013ba059af76461153722005a039d"&gt;the response&lt;/a&gt; from that API call.&lt;/p&gt;

&lt;p&gt;Without that option the API will often provide a lengthy delay while the model burns through thinking tokens until you start getting back visible tokens for the final response.&lt;/p&gt;
&lt;p&gt;OpenAI offer a new &lt;code&gt;reasoning_effort=minimal&lt;/code&gt; option which turns off most reasoning so that tokens start to stream back to you as quickly as possible.&lt;/p&gt;
&lt;h4 id="and-some-svgs-of-pelicans"&gt;And some SVGs of pelicans&lt;/h4&gt;
&lt;p&gt;Naturally I've been running &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle/"&gt;my "Generate an SVG of a pelican riding a bicycle" benchmark&lt;/a&gt;. I'll actually spend more time on this in a future post - I have some fun variants I've been exploring - but for the moment here's &lt;a href="https://gist.github.com/simonw/c98873ef29e621c0fe2e0d4023534406"&gt;the pelican&lt;/a&gt; I got from GPT-5 running at its default "medium" reasoning effort:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/gpt-5-pelican.png" alt="The bicycle is really good, spokes on wheels, correct shape frame, nice pedals. The pelican has a pelican beak and long legs stretching to the pedals." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;It's pretty great! Definitely recognizable as a pelican, and one of the best bicycles I've seen yet.&lt;/p&gt;
&lt;p&gt;Here's &lt;a href="https://gist.github.com/simonw/9b5ecf61a5fb0794729aa0023aaa504d"&gt;GPT-5 mini&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/gpt-5-mini-pelican.png" alt="Blue background with clouds. Pelican has two necks for some reason. Has a good beak though. More gradents and shadows than the GPT-5 one." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;And &lt;a href="https://gist.github.com/simonw/3884dc8b186b630956a1fb0179e191bc"&gt;GPT-5 nano&lt;/a&gt;:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/gpt-5-nano-pelican.png" alt="Bicycle is two circles and some randomish black lines. Pelican still has an OK beak but is otherwise very simple." style="max-width: 100%;" /&gt;&lt;/p&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/openai"&gt;openai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-pricing"&gt;llm-pricing&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/pelican-riding-a-bicycle"&gt;pelican-riding-a-bicycle&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-reasoning"&gt;llm-reasoning&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llm-release"&gt;llm-release&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt-5"&gt;gpt-5&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/gpt"&gt;gpt&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="ai"/><category term="openai"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="llm-pricing"/><category term="pelican-riding-a-bicycle"/><category term="llm-reasoning"/><category term="llm-release"/><category term="gpt-5"/><category term="gpt"/></entry><entry><title>ChatGPT agent's user-agent</title><link href="https://simonwillison.net/2025/Aug/4/chatgpt-agents-user-agent/#atom-tag" rel="alternate"/><published>2025-08-04T22:49:25+00:00</published><updated>2025-08-04T22:49:25+00:00</updated><id>https://simonwillison.net/2025/Aug/4/chatgpt-agents-user-agent/#atom-tag</id><summary type="html">
    &lt;p&gt;I was exploring how ChatGPT agent works today. I learned some interesting things about how it exposes its identity through HTTP headers, then made a huge blunder in thinking it was leaking its URLs to Bingbot and Yandex... but it turned out &lt;a href="https://simonwillison.net/2025/Aug/4/chatgpt-agents-agent/#cloudflare-crawler-hints"&gt;that was a Cloudflare feature&lt;/a&gt; that had nothing to do with ChatGPT.&lt;/p&gt;

&lt;p&gt;ChatGPT agent is the &lt;a href="https://openai.com/index/introducing-chatgpt-agent/"&gt;recently released&lt;/a&gt; (and confusingly named) ChatGPT feature that provides browser automation combined with terminal access as a feature of ChatGPT - replacing their previous &lt;a href="https://help.openai.com/en/articles/10421097-operator"&gt;Operator research preview&lt;/a&gt; which is scheduled for deprecation on August 31st.&lt;/p&gt;

&lt;h4 id="investigating-chatgpt-agent-s-user-agent"&gt;Investigating ChatGPT agent's user-agent&lt;/h4&gt;
&lt;p&gt;I decided to dig into how it works by creating a logged web URL endpoint using &lt;a href="https://simonwillison.net/2024/Aug/8/django-http-debug/"&gt;django-http-debug&lt;/a&gt;. Then I told ChatGPT agent mode to explore that new page:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/chatgpt-agent-url.jpg" alt="ChatGPT screenshot. My prompt was &amp;quot;Visit https://simonwillison.net/test-url-context and tell me what you see there&amp;quot; - it said &amp;quot;Worked for 15 seconds&amp;quot; with an arrow, then a screnshot of the webpage content showing &amp;quot;simonwillison.net&amp;quot; with a favicon, heading &amp;quot;This is a heading&amp;quot;, text &amp;quot;Text and text and more text.&amp;quot; and &amp;quot;this came from javascript&amp;quot;. The bot then responds with: The webpage displays a simple layout with a large heading at the top that reads “This is a heading.” Below it, there's a short paragraph that says “Text and text and more text.” A final line appears underneath saying “this came from javascript,” indicating that this last line was inserted via a script. The page contains no interactive elements or instructions—just these lines of plain text displayed on a white background." style="max-width: 100%;" /&gt;&lt;/p&gt;
&lt;p&gt;My logging captured these request headers:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Via: 1.1 heroku-router
Host: simonwillison.net
Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.7
Cf-Ray: 96a0f289adcb8e8e-SEA
Cookie: cf_clearance=zzV8W...
Server: Heroku
Cdn-Loop: cloudflare; loops=1
Priority: u=0, i
Sec-Ch-Ua: "Not)A;Brand";v="8", "Chromium";v="138"
Signature: sig1=:1AxfqHocTf693inKKMQ7NRoHoWAZ9d/vY4D/FO0+MqdFBy0HEH3ZIRv1c3hyiTrzCvquqDC8eYl1ojcPYOSpCQ==:
Cf-Visitor: {"scheme":"https"}
User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36
Cf-Ipcountry: US
X-Request-Id: 45ef5be4-ead3-99d5-f018-13c4a55864d3
Sec-Fetch-Dest: document
Sec-Fetch-Mode: navigate
Sec-Fetch-Site: none
Sec-Fetch-User: ?1
Accept-Encoding: gzip, br
Accept-Language: en-US,en;q=0.9
Signature-Agent: "https://chatgpt.com"
Signature-Input: sig1=("@authority" "@method" "@path" "signature-agent");created=1754340838;keyid="otMqcjr17mGyruktGvJU8oojQTSMHlVm7uO-lrcqbdg";expires=1754344438;nonce="_8jbGwfLcgt_vUeiZQdWvfyIeh9FmlthEXElL-O2Rq5zydBYWivw4R3sV9PV-zGwZ2OEGr3T2Pmeo2NzmboMeQ";tag="web-bot-auth";alg="ed25519"
X-Forwarded-For: 2a09:bac5:665f:1541::21e:154, 172.71.147.183
X-Request-Start: 1754340840059
Cf-Connecting-Ip: 2a09:bac5:665f:1541::21e:154
Sec-Ch-Ua-Mobile: ?0
X-Forwarded-Port: 80
X-Forwarded-Proto: http
Sec-Ch-Ua-Platform: "Linux"
Upgrade-Insecure-Requests: 1
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;That &lt;strong&gt;Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/138.0.0.0 Safari/537.36&lt;/strong&gt; user-agent header is the one used by the most recent Chrome on macOS - which is a little odd here as the &lt;strong&gt;Sec-Ch-Ua-Platform : "Linux"&lt;/strong&gt; indicates that the agent browser runs on Linux.&lt;/p&gt;
&lt;p&gt;At first glance it looks like ChatGPT is being dishonest here by not including its bot identity in the user-agent header. I thought for a moment it might be reflecting my own user-agent, but I'm using Firefox on macOS and it identified itself as Chrome.&lt;/p&gt;
&lt;p&gt;Then I spotted this header:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Signature-Agent: "https://chatgpt.com"
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Which is accompanied by a much more complex header called &lt;strong&gt;Signature-Input&lt;/strong&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Signature-Input: sig1=("@authority" "@method" "@path" "signature-agent");created=1754340838;keyid="otMqcjr17mGyruktGvJU8oojQTSMHlVm7uO-lrcqbdg";expires=1754344438;nonce="_8jbGwfLcgt_vUeiZQdWvfyIeh9FmlthEXElL-O2Rq5zydBYWivw4R3sV9PV-zGwZ2OEGr3T2Pmeo2NzmboMeQ";tag="web-bot-auth";alg="ed25519"
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And a &lt;code&gt;Signature&lt;/code&gt; header too.&lt;/p&gt;
&lt;p&gt;These turn out to come from a relatively new web standard: &lt;a href="https://www.rfc-editor.org/rfc/rfc9421.html"&gt;RFC 9421 HTTP Message Signatures&lt;/a&gt;' published February 2024.&lt;/p&gt;
&lt;p&gt;The purpose of HTTP Message Signatures is to allow clients to include signed data about their request in a way that cannot be tampered with by intermediaries. The signature uses a public key that's provided by the following well-known endpoint:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;https://chatgpt.com/.well-known/http-message-signatures-directory
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Add it all together and we now have a rock-solid way to identify traffic from ChatGPT agent: look for the &lt;code&gt;Signature-Agent: "https://chatgpt.com"&lt;/code&gt; header and confirm its value by checking the signature in the &lt;code&gt;Signature-Input&lt;/code&gt; and &lt;code&gt;Signature&lt;/code&gt; headers.&lt;/p&gt;
&lt;h4 id="and-then-came-the-crawlers"&gt;And then came Bingbot and Yandex&lt;/h4&gt;
&lt;p&gt;Just over a minute after it captured that request, my logging endpoint got another request:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Via: 1.1 heroku-router
From: bingbot(at)microsoft.com
Host: simonwillison.net
Accept: */*
Cf-Ray: 96a0f4671d1fc3c6-SEA
Server: Heroku
Cdn-Loop: cloudflare; loops=1
Cf-Visitor: {"scheme":"https"}
User-Agent: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; bingbot/2.0; +http://www.bing.com/bingbot.htm) Chrome/116.0.1938.76 Safari/537.36
Cf-Ipcountry: US
X-Request-Id: 6214f5dc-a4ea-5390-1beb-f2d26eac5d01
Accept-Encoding: gzip, br
X-Forwarded-For: 207.46.13.9, 172.71.150.252
X-Request-Start: 1754340916429
Cf-Connecting-Ip: 207.46.13.9
X-Forwarded-Port: 80
X-Forwarded-Proto: http
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;I pasted &lt;code&gt;207.46.13.9&lt;/code&gt; into Microsoft's &lt;a href="https://www.bing.com/toolbox/verify-bingbot-verdict"&gt;Verify Bingbot&lt;/a&gt; tool (after solving a particularly taxing CAPTCHA) and it confirmed that this was indeed a request from Bingbot.&lt;/p&gt;
&lt;p&gt;I set up a second URL to confirm... and this time got a visit from Yandex!&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;Via: 1.1 heroku-router
From: support@search.yandex.ru
Host: simonwillison.net
Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8
Cf-Ray: 96a16390d8f6f3a7-DME
Server: Heroku
Cdn-Loop: cloudflare; loops=1
Cf-Visitor: {"scheme":"https"}
User-Agent: Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots)
Cf-Ipcountry: RU
X-Request-Id: 3cdcbdba-f629-0d29-b453-61644da43c6c
Accept-Encoding: gzip, br
X-Forwarded-For: 213.180.203.138, 172.71.184.65
X-Request-Start: 1754345469921
Cf-Connecting-Ip: 213.180.203.138
X-Forwarded-Port: 80
X-Forwarded-Proto: http
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Yandex &lt;a href="https://yandex.com/support/webmaster/en/robot-workings/check-yandex-robots.html?lang=en"&gt;suggest a reverse DNS lookup&lt;/a&gt; to verify, so I ran this command:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;dig -x 213.180.203.138 +short
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;And got back:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;213-180-203-138.spider.yandex.com.
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;Which confirms that this is indeed a Yandex crawler.&lt;/p&gt;

&lt;p&gt;I tried a third experiment to be sure... and got hits from both Bingbot and YandexBot.&lt;/p&gt;

&lt;h4 id="cloudflare-crawler-hints"&gt;It was Cloudflare Crawler Hints, not ChatGPT&lt;/h4&gt;

&lt;p&gt;So I wrote up and posted about my discovery... and &lt;a href="https://x.com/jatan_loya/status/1952506398270767499"&gt;Jatan Loya asked:&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;&lt;p&gt;do you have crawler hints enabled in cf?&lt;/p&gt;&lt;/blockquote&gt;

&lt;p&gt;And yeah, it turned out I did. I spotted this in my caching configuration page (and it looks like I must have turned it on myself at some point in the past):&lt;/p&gt;

&lt;p&gt;&lt;img src="https://static.simonwillison.net/static/2025/cloudflare-crawler-hints.jpg" alt="Screenshot of Cloudflare settings panel showing &amp;quot;Crawler Hints Beta&amp;quot; with description text explaining that Crawler Hints provide high quality data to search engines and other crawlers when sites using Cloudflare change their content. This allows crawlers to precisely time crawling, avoid wasteful crawls, and generally reduce resource consumption on origins and other Internet infrastructure. Below states &amp;quot;By enabling this service, you agree to share website information required for feature functionality and agree to the Supplemental Terms for Crawler Hints.&amp;quot; There is a toggle switch in the on position on the right side and a &amp;quot;Help&amp;quot; link in the bottom right corner." style="max-width: 100%" /&gt;&lt;/p&gt;

&lt;p&gt;Here's &lt;a href="https://developers.cloudflare.com/cache/advanced-configuration/crawler-hints/"&gt;the Cloudflare documentation for that feature&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I deleted my posts on Twitter and Bluesky (since you can't edit those and I didn't want the misinformation to continue to spread) and edited &lt;a href="https://fedi.simonwillison.net/@simon/114972968822349077"&gt;my post on Mastodon&lt;/a&gt;, then updated this entry with the real reason this had happened.&lt;/p&gt;

&lt;p&gt;I also changed the URL of this entry as it turned out Twitter and Bluesky were caching my social media preview for the previous one, which included the incorrect information in the title.&lt;/p&gt;

&lt;details&gt;&lt;summary&gt;Original "So what's going on here?" section from my post&lt;/summary&gt;

&lt;p&gt;&lt;em&gt;Here's a section of my original post with my theories about what was going on before learning about Cloudflare Crawler Hints.&lt;/em&gt;&lt;/p&gt;

&lt;h4 id="so-what-s-going-on-here-"&gt;So what's going on here?&lt;/h4&gt;
&lt;p&gt;There are quite a few different moving parts here.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;I'm using Firefox on macOS with the 1Password and Readwise Highlighter extensions installed and active. Since I didn't visit the debug pages at all with my own browser I don't think any of these are relevant to these results.&lt;/li&gt;
&lt;li&gt;ChatGPT agent makes just a single request to my debug URL ...&lt;/li&gt;
&lt;li&gt;... which is proxied through both Cloudflare and Heroku.&lt;/li&gt;
&lt;li&gt;Within about a minute, I get hits from one or both of Bingbot and Yandex.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Presumably ChatGPT agent itself is running behind at least one proxy - I would expect OpenAI to keep a close eye on that traffic to ensure it doesn't get abused.&lt;/p&gt;
&lt;p&gt;I'm guessing that infrastructure is hosted by Microsoft Azure. The &lt;a href="https://openai.com/policies/sub-processor-list/"&gt;OpenAI Sub-processor List&lt;/a&gt; - though that lists Microsoft Corporation, CoreWeave Inc, Oracle Cloud Platform and Google Cloud Platform under the "Cloud infrastructure" section so it could be any of those.&lt;/p&gt;
&lt;p&gt;Since the page is served over HTTPS my guess is that any intermediary proxies should be unable to see the path component of the URL, making the mystery of how Bingbot and Yandex saw the URL even more intriguing.&lt;/p&gt;
&lt;/details&gt;
    
        &lt;p&gt;Tags: &lt;a href="https://simonwillison.net/tags/bing"&gt;bing&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/privacy"&gt;privacy&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/search-engines"&gt;search-engines&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/user-agents"&gt;user-agents&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/ai"&gt;ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/cloudflare"&gt;cloudflare&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/generative-ai"&gt;generative-ai&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/chatgpt"&gt;chatgpt&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/llms"&gt;llms&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/browser-agents"&gt;browser-agents&lt;/a&gt;, &lt;a href="https://simonwillison.net/tags/retractions"&gt;retractions&lt;/a&gt;&lt;/p&gt;
    

</summary><category term="bing"/><category term="privacy"/><category term="search-engines"/><category term="user-agents"/><category term="ai"/><category term="cloudflare"/><category term="generative-ai"/><category term="chatgpt"/><category term="llms"/><category term="browser-agents"/><category term="retractions"/></entry></feed>