LLM digest: May 2026
Sent
I published 53 posts on my blog in May. Here's your sponsors-only summary of the most important trends and highlights from the past month.
As always, this issue and previous issues are archived in my simonw-private/monthly GitHub repository.
AI got expensive, and Anthropic had a really good month
The big theme of April and May is that AI got expensive. In April we saw GPT-5.5 at 2x the price of GPT-5.4 and Opus 4.7 at sort-of 1.4x the price of Opus 4.6.
It turns out there was a more important detail that I'd missed: by the end of April both OpenAI and Anthropic had fully rolled out new enterprise pricing which makes larger companies pay full API pricing for their token usage. Those $100 and $200/month subscription plans that offer a significant discount for Claude Code and Codex users? No longer an option for those larger customers.
Anthropic look set to have their first profitable quarter. They just raised a $65B Series H at a $965B valuation, meaning they're currently valued higher than OpenAI (who raised $122B for an $852B valuation on March 31st). In Anthropic's Series H announcement they claimed $47B in annualized "run-rate" revenue. I plotted that against their previously shared figures - ramping up from just $9B at the end of 2025 - and they may have the fastest-scaling revenue of any company in history.
But what does "annualized run-rate" actually mean? According to a leak to Karen Kwok at Reuters:
Anthropic defines “run-rate revenue” in two parts. Use the last 28 days of sales from customers charged on a consumption basis and multiply it by 13. Then, multiply the monthly subscription take by 12, and add the two together.
In I think Anthropic and OpenAI have found product-market fit I proposed that coding agents (and general purpose agents) for enterprise companies are the first proven business models for the big AI labs that have a chance of making a significant dent against their enormous expenses.
I've been reviewing my own spending (first using ccusage, then AgentsView) and I'm getting about $1,000/month in API-priced token value each for $100/month with both Anthropic and OpenAI (my spend is $200/month total). It's easy to see how paying full API price could rack up some serious bills.
One company (Amazon perhaps?) supposedly spent $500 million on Claude in a single month! Given Anthropic's habit of reporting "annualized run-rate" as 13x consumption revenue from the past 28 days, that one customer could be responsible for $6.5B of their reported $47B run-rate!
There was also a lot of online chatter about Uber blowing their annual AI budget in the first four months of 2026. I think this story was over-played - if you set a budget in 2025 without guessing how wildly popular coding agents like Claude Code and Codex would become following the November 2025 inflection point, then it's no surprise you'd low-ball 2026.
One thing is very clear though: the whole "Tokenmaxxing leaderboard" trend needs to be over already. It's gone from just obviously stupid to obviously stupid and ruinously expensive.
The model releases were a little disappointing
After April's flurry of new models, May was actually quite slow. The two big ones were Claude Opus 4.8 and Gemini 3.5 Flash.
Anthropic themselves described Claude Opus 4.8 as “a modest but tangible improvement”. It's refreshing to see a lab admit that their latest model is an incremental improvement over its predecessor, and that declaration also fitted the big theme of the announcement which was "honesty".
I published some notes and five pelicans for each of the five thinking levels, and I'm pleased to report that the "max" level did indeed draw the best pelican of the lot (at a hefty price of 43 cents).
Meanwhile, Google released Gemini 3.5 Flash during the keynote at this year's Google I/O. Most of the Gemini models spend a bunch of time in -preview mode but this one went straight to general availability. Google have already rolled it out to power the Gemini app, the new Google Antigravity agent framework and their "Gemini Enterprise Agent Platform".
Fitting the pricing trend from the other labs, Gemini 3.5 Flash is 3x the price of Gemini 3 Flash preview!
A better indication of comparative pricing comes from Artificial Analysis, who measure the cost to run their proprietary intelligence benchmark against each model - which takes both price and the number of tokens used into account. 3.5 Flash (high) cost $1,551.60, significantly more than Gemini 3.1 Pro (preview) at $892.28 and 5.6x the cost of Gemini 3 Flash Preview (reasoning) at $278.26.
3.5 Flash drew me a pelican that hedgehog on Hacker News said "looks like it’s in Miami for a crypto conference."
Opus 4.8 and Gemini 3.5 Flash are clearly the new frontier models for their respective labs, but they don't yet seem to me to be beating GPT-5.5. Google say Gemini 3.5 Pro is coming "next month" and Anthropic promise Claude Mythos-class models "in the coming weeks".
Conferences and podcasts
I attended Anthropic's Code w/ Claude developer event and live-blogged the keynote. At last year's event they shipped Claude 4, but this year they had very few new announcements at the event itself. The big one was a partnership with xAI to use their infamous Colossus data centers - the ones in Memphis with the terrible air pollution track record. We later learned from the SpaceX IPO S-1 that Anthropic "has agreed to pay us $1.25 billion per month through May 2029, with capacity ramping in May and June 2026 at a reduced fee."
My favorite new feature from the conference was Dreams, a perfectly named system where Claude can run a scheduled job that "reads an existing memory store alongside past session transcripts, then produces a new, reorganized memory store".
I gave a keynote at Heavybit's Write-Only Code Summit, which was preceded by this interview on their High Leverage podcast. I extracted a section of that transcript into a blog post Vibe coding and agentic engineering are getting closer than I’d like.
Then at PyCon US I helped chair the new AI track, and gave a lightning talk on The last six months in LLMs in five minutes - now available on my blog as an annotated slide deck, hopefully out on video some time soon.
I launched Datasette Agent and made a lot of progress on Datasette
I had a very busy month of coding. My big project was Datasette Agent, an AI assistant plugin for Datasette that's been in the works for a few months. Solving the problem of how to build a chat-based agent that justified its existence in a world full of agents was a fun challenge. It's very Datasette-ish - out of the box it can answer questions by running read-only queries against your SQLite databases, but it also extends Datasette's own plugin system to allow plugins to provide new tools for the agent. So far there are tool plugins for plotting charts and running code in a Fly.io Sprite sandbox, and I'm having a lot of fun spinning up new plugins for all sorts of other concerns within the Datasette ecosystem.
You can try a live demo at agent.datasette.io (sign in with GitHub). The demo is currently running on Gemini 3.1 Flash-Lite, which is 1/6th the price of Gemini 3.5 Flash and evidently a very capable model. Even the smaller models that run on my laptop are working great with Datasette Agent - SQLite SQL and basic tool calls are widely supported these days.
Work on Datasette Agent inspired two new Datasette alphas with significant new features. Datasette 1.0a30 adds an extensive "jump menu", where hitting / brings up a menu for jumping to tables, queries or custom content from plugins. Datasette 1.0a31 landed a long-overdue rethink of how stored SQL queries work in Datasette - you can now execute write SQL queries directly (provided you have the necessary permissions), and stored queries can be saved to Datasette's internal database either privately or shared with others. Both of those releases were featured on the new Datasette project blog.
What I'm using, May 2026 edition
This month I've spent more time in the OpenAI Codex desktop app than any other coding agent surface, generally using GPT-5.5 high or xhigh. I've tinkered a little bit with their fast mode too, but I'm too nervous to use it full time as I worry it will exhaust my allowance.
I really like this app. It seamlessly interoperates with the codex CLI tool, so any session I start with that shows up in the app automatically. It has a very solid built-in web preview and agent-powered browser tool: in many of my sessions I'll have it run a development server so I can preview changes as it makes them.
It also has a capable mobile integration, so I can leave my laptop on my desk at home and then continue directing sessions from my phone while I'm out walking the dog and photographing the local pelicans.
Codex has the feature I most want from a coding agent: an easy "export transcript as Markdown" option. I use that to paste transcripts into Gists like this one which I then link from my PRs or commit messages.
(The export feature disappeared in a recent release, but thankfully that was a mistake and it's coming back shortly.)
I continue to use Claude Code for web as my asynchronous, cloud-based agent of choice. Many of my projects start out in Claude Code for web and later switch to Codex on my laptop.
I'm currently paying Anthropic $100/month and OpenAI $100/month. This gives me access to GPT-5.5 Pro in ChatGPT which I'm finding to be an absurdly capable research assistant for complex questions - it's effectively the Deep Research pattern under a different name; it frequently churns away for 10 minutes before digging up exactly what I needed, often from dozens or hundreds of different sources.
I just discovered the AgentsView Python tool by Wes McKinney and I love it. Run the following command (no installation necessary) to get a localhost web app that provides search and aggregated statistics against your local transcripts for Claude Code, Codex, Pi, and other agents.
uvx agentsview serve
Or run this command to get a token and cost breakdown per-model directly in your terminal:
uvx agentsview usage daily --breakdown
Under the hood it creates a SQLite database, so having run it once you can then explore all of your data using Datasette like this:
uvx datasette ~/.agentsview/sessions.db
Miscellaneous extras
- The Pope's encyclical Magnifica Humanitas of His Holiness Pope Leo XIV on Safeguarding the Human Person in the Time of Artificial Intelligence is an extraordinary document - insightful, technically accurate and likely to be deeply influential. I published a few of my own notes here, and I particularly enjoyed Jacob Ward's optimistic TikTok coverage of the document. Since May was the month of Anthropic and they managed to have one of their co-founders, Christopher Olah, attend the event at the Vatican, I liked Corey Quinn's observation:
I cannot believe I'm saying this, but getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen.
- I figured out how to use my LLM CLI tool in the shebang line of a script, meaning you can write scripts in human languages:
#!/usr/bin/env -S llm -f Generate an SVG of a pelican riding a bicycle - For a while I've been wondering if I could build a better version of Datasette Lite - the Datasette Python app running entirely in the browser via Pyodide and WebAssembly - using Service Workers. I ran a coding agent research project the other day which convinced me this can work, and got this Datasette prototype up and running. More notes here.
That's it for May!
If this newsletter was useful, feel free to forward it to friends who might find it useful too, especially if they might be convinced to sign up to sponsor me for the next one!