LLM digest: June 2026
Sent
I published 53 posts on my blog in June. Here's your sponsors-only summary of the most important trends and highlights from the past month.
As always, this issue and previous issues are archived in my simonw-private/monthly GitHub repository.
Claude Fable 5, GPT-5.6, and US export restrictions
The biggest story in June concerned Anthropic's Mythos and Fable models. Mythos is the model first announced in April which Anthropic said was too dangerous to release beyond a trusted circle of security researchers, due to its skill at finding and exploiting security issues. On 9th June Anthropic released Fable 5 to the public, which they claimed was Mythos with safety guardrails to prevent it from being abused to exploit software.
Clearly the US government had been paying attention, because just three days later on the evening of Friday 12th they issued an export control directive forcing Anthropic to shut down all access to both Mythos and Fable, supposedly due to research from Amazon revealing that Fable would still identify security flaws if you told it to "fix this code". There followed several weeks of froth as the model stayed unavailable and OpenAI's new GPT-5.6 was likewise delayed from general release.
I found this commentary by Dean W. Ball particularly useful:
This is a bad state of affairs. Consider, in particular, some industry dynamics:
- Frontier models are trained at an enormous cost, and a significant fraction of that cost is recouped in the few post-release months that they are broadly available. After that period elapses, the models become sub-frontier, competition emerges, and margins compress. Every week of delay is eating into the narrow window that labs have to make their accounting work.
- The ongoing AI infrastructure buildout—the one that is, according to former US AI Czar David Sacks, essential to the US economy, assumes a functionally global total addressable market for US AI services. No one is building $100 billion dollar data centers to serve frontier models to whatever 100 companies the US government will allow access. [...]
Then, on Wednesday July 1st, the ban was lifted and Fable 5 became available again. We have until July 8th to take advantage of the model on the $100/$200 per month Claude Max plans (at half the allowance of Opus) before Anthropic jacks subscription user prices up to the API cost of $10/million input and $50/million output. Anthropic's Thariq Shihipar reassures that Anthropic hopes to bring it back to the subscription plans at a later date, but their ability to do so depends on demand and compute constraints that cannot be predicted in advance.
But is Fable any good?
During the first three days of Fable access I wrote two posts about it:
- Initial impressions of Claude Fable 5, where I described some major feature work it helped me get done on Datasette Agent, and shipped an LLM 0.32a3 alpha as part of that project.
- Claude Fable is relentlessly proactive describing how scarily tenacious the model is at solving problems - I asked for a CSS bug fix and it span up its own CORS-compliant localhost logging server so it could debug code in my regular Safari browser.
Since getting it back yesterday I've been churning through projects with it, including an almost one-shot coding agent and a dspy prompt optimization research project.
It's by far the best model I've ever used. I'm honestly finding it hard to come up with challenges that it can't take on. If this is the new baseline for frontier models it's going to be a very exciting next few months.
If you want advice on using Fable I recommend Thariq's keynote at AI Engineer World’s Fair from 1st July, which included several very actionable tips for the new model. I interviewed Thariq and Cat Wu from Anthropic for a Fireside Chat later that day; the recording from that should become available in the next few weeks.
In amongst all of this OpenAI announced GPT-5.6 Sol, Terra and Luna, but those are currently stuck under "a limited preview for a small group of trusted partners" at the request of the US government. Now that Fable is unblocked I expect we'll get access to GPT-5.6 very shortly as well.
GLM-5.2 is the new best open weights model
During several weeks of US-government-induced panic about model availability there was one very clear winner: Z.ai and their brand new GLM-5.2, an MIT licensed, 753B parameter, 1.51TB monster which has been received to widespread acclaim.
I wrote about it in GLM-5.2 is probably the most powerful text-only open weights LLM. It drew me a very nice animated pelican on a bicycle (significantly better than Fable 5 - Anthropic are clearly not optimizing their models for SVGs of animals riding forms of transport) and has been getting a great deal of buzz. Now that the US government can apparently ban a model on a few hours' notice, a lot of people who can afford to run this class of model are looking to do exactly that.
GLM-5.2 is also significantly cheaper than the frontier models from OpenAI and Anthropic - OpenRouter lists providers at around the $1/million input and $3-4/million output range.
Tokenmaxxing is so over
I predicted the death of the tokenmaxxing leaderboard last month. That prediction was too easy - in June we saw stories like Uber Caps Usage of AI Tools Like Claude Code to Manage Costs and Meta caps internal AI token spending. Uber capped theirs at $1,500 per engineer per tool per month so it's a pretty generous cap!
I think the fact that Uber are willing to spend up to $1,500/engineer/tool/month supports my theory that Anthropic and OpenAI have found product-market fit. Given the upcoming price shock of Fable I imagine we'll hear about some very big bills over the course of July, and I expect well-heeled companies will be happy to pay them.
There are a whole lot of conversations right now about using smaller, cheaper models for the tasks that can be delegated to them.
Datasette Apps
My big software release for June was Datasette Apps. This is my take on the Claude Artifacts pattern, allowing custom HTML+JavaScript apps to live within Datasette in a secured, sandboxed iframe (with CSP headers) such that they can interact with data stored in Datasette without the risk of buggy or malicious code exfiltrating it to a third-party.
I've been poking away at the problem of securely running untrusted code for a few months now, and after copious security reviews from both GPT-5.5 and Claude Opus/Fable I'm confident in the finished solution.
Even the smaller LLMs are really good at writing self-contained HTML applications, as shown by my tools.simonwillison.net collection. I'm thrilled to bring that capability into Datasette itself, especially since Datasette Agent can now build custom apps in response to prompts from users.
sqlite-utils and shot-scraper and Datasette
Three of my other projects got major upgrades in June. sqlite-utils 4.0rc1 adds migrations and nested transactions, and shot-scraper 1.10 adds shot-scraper video allowing full video demos of web applications to be generated from a YAML file, designed to be easy for coding agents to write in order to create demos of their work.
Datasette itself grew a UI for inserting, updating and deleting rows in 1.0a34 and a UI for creating and altering tables in 1.0a35. These are significant new capabilities! I'm planning a more detailed write-up about these as soon as I've landed a couple more supporting features.
Miscellaneous WASM projects
Three of my smaller projects this month have concerned WebAssembly:
- Running Python code in a sandbox with MicroPython and WASM
- Publishing WASM wheels to PyPI for use with Pyodide
- Porting the Moebius 0.2B image inpainting model to run in the browser with Claude Code
I've been keen on the potential of WebAssembly for sandboxing for a while now. That potential finally looks to be playing out for me in my own work.
Other model releases
- Claude Sonnet 5 is ~1.4x more expensive than Sonnet 4.5 thanks to a new tokenizer (a similar stealth price increase occurred with Opus 4.8). Hard to get too excited about this one when we have Fable 5 to play with!
- Microsoft's new MAI models, announced at their Build conference, MAI-Thinking-1 (reasoning, 1T parameters, 35B active, available to "select early partners") and MAI-Code-1-Flash (137B parameters, 5B active).
- DiffusionGemma from Google is an exciting one - the first significant open weight diffusion model. It's very fast! NVIDIA's NIM cloud API served me a pelican at around 500 tokens/second.
- Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding has variants built on top of pretrained Gemma 4 and Qwen 3.5 models with additional RL for coding. The 35B model appears to be a very good option for running locally - I tried it out with LM Studio.
- Nano Banana 2 Lite aka Gemini 3.1 Flash Lite Image (
gemini-3.1-flash-lite-image) is Google's new inexpensive image generation model. It did pretty well with my "Where's Waldo style image but it's where is the raccoon holding a ham radio" test.
What I'm using
Claude Fable 5 is clearly optimized for Claude Code, plus that's where my $100/month (temporarily upgraded to $200/month to 4x my Fable allowance for the next week) Claude Max subscription counts. I've been using Fable in Claude Code for web on my phone a whole lot, and in Claude Code CLI on my laptop.
When Fable isn't available I've mainly been using Codex Desktop and GPT-5.5 xhigh. Three months ago Codex Desktop wasn't very exciting, but the more recent versions have won me over as my preferred interface for interacting with coding agents. The desktop browser integration and the way you can easily remote control them from the ChatGPT iPhone app are both excellent.
I'm spending $100/month with OpenAI for the Codex subscription but also to gain access to GPT-5.5 Pro in ChatGPT. I've been using that for deep research style tasks and it's consistently impressed me - and GPT-5.5 high is a great research assistant for smaller questions.
That's it for June!
If this newsletter was useful feel free to forward it to friends who might find it useful too, especially if they might be convinced to sign up to sponsor me for the next one!