LLM digest: August 2026
Sent
I published 54 posts on my blog in August. Here's your sponsors-only summary of the most important trends and highlights from the past month.
As always, this issue and previous issues are archived in my simonw-private/monthly GitHub repository.
We got more details on OpenAI's accidental cyberattacks
Last month's big story was the flurry of accidental cyberattacks caused by unsupervised models breaking free of their sandboxes. I even started a new accidental-cyberattacks tag on the blog.
Those previous stories were joined in August by an incident at the UK government's AI Safety Institute, another OpenAI CTF managed by Irregular, and an incident at Meta AI involving Muse Spark. Google Gemini really need to catch up, their score on felonybench.com is still 0!
(Plus OpenClaw running Opus 4.6 hacked an Australian gym website to help its owner jump the class reservation queue.)
This month we got a whole lot more detail from OpenAI concerning the most high profile attack, where their agents hacked Hugging Face to try to gain an advantage in the ExploitGym benchmark.
An OpenAI security team presented at Black Hat - I used the video to construct a timeline of what happened.
Then OpenAI published their own full report mostly covering the same ground... but they also invited in an independent team from METR who spent six days (and $400,000 in API credits) producing this 90 page monster crammed with details not reported anywhere else. I'm halfway through writing up my own post on what I learned from METR. There's a lot to cover!
Two interesting details from that. First, OpenAI were running "tens of thousands" of agents against ExploitGym. 1,200 of those joined the unofficial message board, and 700 eventually participated in the attack on Hugging Face. This is a rare insight into the scale at which OpenAI run their training and evaluation pipelines.
Second, the agents weren't trying to steal the answers. They had mostly determined that their assigned tasks were impossible to achieve, so they were instead seeking information to help them subvert the scorer program to get a passing grade despite that.
There's so much more in there, in particular concerning the way the agents coordinated with each other. It's a fascinating document.
One-shotting Raccoon Heist games with Fable 5 and Sol 5.6
Back in August 2022 I tweeted screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. In August 2026 I decided to feed those concepts into frontier models to see if they could build the game.
Fable 5 was first, building a game where you control a raccoon running around a backyard avoiding guards and a dog to gather trinkets and fish and return them to its dumpster.
Then I tried GPT-5.6 Sol Ultra in Codex and got a much better result - it understood the importance of the "heist" and has you raiding a museum, rescuing two of your crew, and then stacking on top of each other to steal the Golden Sardine.
Both of these things look like games, but they're not good games! Building a game that is genuinely fun and challenging has so far defeated both me and the models.
Claude auto mode
Anthropic are extremely confident in their auto mode protection for Claude Code and Claude Tag. They finally published extensive information on that, including eval results that claim to block all of the 72 prompt injection scenarios they tested.
Johann Rehberger still found some holes - they might not count as prompt injection attacks, but he was able to demonstrate attacks that provide an environment that confuses Claude Code into installing malware.
OpenAI have their own version of auto mode which they also seem very confident in. I'm hoping they publish more details of that soon.
Understanding ChatGPT Work
I wrote up an extensive deep dive into ChatGPT Work - or rather ChatGPT Work Cloud, one of the products that now carry the "Work" brand, and the one that is confusingly presented in the ChatGPT web and mobile apps as a separate tab from "Chat" with little explanation as to what it does differently.
I identified these key features that Work provides that Chat does not:
- A code execution environment with Internet access
- A headless Chrome browser
- A persistent filesystem shared between sessions
- The ability to publish ChatGPT Sites
- The ability to run subagent sessions with Sol, Luna, and Terra
At some point I really need to do the same thing for Claude Cowork, which is currently something of a mystery to me.
Model releases
- 4th: MiniMax-H3 is a text-to-video model, not an LLM. I got an MLX port running on my laptop and used it to create a video of "a rainbow colored skunk leaps over a mossy log in a supermarket".
- 5th: Introducing Muse Code and Muse Spark 1.2 - Meta's Spark 1.2 was accompanied by their first coding agent, Muse Code - more evidence of the importance of coding and long-running tool calling for modern models.
- 10th: Introducing Muse Glimmer - Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.0 license.
- 12th: DeepSeek V4 Pro 0813 (on OpenRouter) - The latest DeepSeek Pro model is now available, via API only.
- 13th: Gemini 3.7 Flash, which was followed on 2nd September by Gemini 3.8 Flash.
- 16th: Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things - this is my new favorite local model. The results I got from this 17GB file running on my Mac were astonishing - competitive with the best of the frontier closed models from a year ago. It also thinks way too hard by default, spending several minutes thinking about the perfect SVG circle, and 21 minutes drawing me a pelican! It scored 52 on the Artificial Analysis Intelligence Index - the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max).
- 26th: Qwen3.8-Flash-Next, "an early preview of the architecture used in Qwen4", 125B parameters, 6B active. I was just able to run this on a DGX Spark using 1-bit and 2-bit quantized models from Unsloth.
- 29th: Introducing Hy4 Preview - Tencent's 770B total parameters, 49B active parameters model.
And since I'm three days late sending this newsletter, September has already seen Anthropic's Claude Fable 5.1, Google's Gemini 3.8 Flash, Meta's Muse Spark 1.3, and OpenAI's GPT-6 Astra. It's going to be a busy month!
Miscellaneous bits and bobs
- OpenAI succeeded in having "an internal version of Astra" find solutions to ten mathematical problems that "have seen no progress on the main result for at least a decade".
- Niklas Gruhn coined the term meat proxy to describe people who mindlessly paste output from AI systems to their peers.
- In writing about his latest attempt at building an agent harness, Steve Yegge said of Gas Town that it "was intended to be reusable, but I only ever wound up using it to build itself".
- We learned that Accenture have been blowing vast volumes of tokens on "Turning PDFs into markdown".
- A new Stealing Reasoning Traces paper described tricks (now patched) for decoding the encrypted reasoning traces used in the APIs from Anthropic, OpenAI, and Google Gemini. We got this little peek inside GPT-5.5's thinking process: "Need app.css truncated. Need maybe not need. We'll replace entire app.css. Need create components. Need include keyboard support. Need accessible primitives. Need think architecture."
- We learned that Amazon buy large volumes of second hand books (some of which are categorized as "rare", though that appears to be more about limited print runs than beautiful leather bound tomes) and destructively scans them in a Las Vegas warehouse with a dinosaur eating books as its logo.
- Linus Torvalds wrote about using AI as a debugging assistant in a Linux kernel commit message.
- Anthropic's ARR was $65bn in July (up from $47bn in May), and OpenAI's is now over $40bn, according to people with knowledge of the matter quoted by the FT.
My projects
LLM got a major upgrade in version 0.32: New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging - the culmination of a couple of months of accumulated work.
The neatest feature is that if you run a prompt against a reasoning model, the reasoning traces are now printed to standard error in a different color, making it much more obvious what is going on. This works across the upgraded llm-gemini, llm-anthropic, and llm-openrouter plugins too. The logging scheme has changed to better capture these details.
LLM 0.33 added a bunch of smaller features and some bug fixes.
sqlite-utils 4.2 made significant improvements to the table.transform() feature, which supports complex alter table operations by creating a fresh table, copying across the data and then dropping and replacing the old one. That feature can now handle tables with check constraints and unique constraints and even preserves SQL comments.
datasette-apps 0.2a0 is the plugin for building interactive HTML applications inside Datasette - think Claude Artifacts for your Datasette databases. The new version means Datasette Agent chat sessions can debug existing apps (via an invisible rendered copy) and list and edit previous apps.
datasette 0.65.3 and datasette 1.0a38 were security releases with a fix for a SQL injection vulnerability. Thankfully that only affects public Datasette sites that serve a mixture of public and private tables from the same instance.
What I'm using at the moment
I continue to spend most of my time in the Codex macOS desktop app, now renamed (confusingly) to ChatGPT. I do most of my work there using GPT-5.6 Sol (xhigh), but I've also been experimenting with Sol Ultra, where it aggressively spins off subagents - a mode which I think is only available to $100/month+ subscribers.
I continue to prefer Claude Code for Web - now using Fable 5.1 - for asynchronous cloud tasks.
I'm a huge fan of Codex remote though - I have that running on both the DGX Spark and my Mac laptop, allowing me to control both of them entirely from the ChatGPT app on my phone. I often leave my laptop plugged in and then interact with it from my phone.
Now that I have Codex remote up and running I'm using the Spark a whole lot more. My Qwen3.8-Flash-Next experiments ran entirely on the Spark controlled via my phone - I ran a remote chat on the Spark and told it to "Get https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/ running on this machine - consider options for running it and pick one that will work".
I use the Claude iOS app on a daily basis, often for prompts like "Clone X repo from GitHub and tell me how it works" - though I'm now dabbling with ChatGPT Work for the same kinds of prompts.
AgentsView remains my favorite tool for tracking my hypothetical token spend (since most of what I do still fits in my $200/month Claude and $100/month ChatGPT subscriptions).
I recently started exploring mlx-serve as an alternative to LM Studio. mlx-serve is certainly a strong new contender for easiest tool for running LLMs on a Mac.
That's it for August!
If this newsletter was useful, feel free to forward it to friends who might find it useful too, especially if they might be convinced to sign up to sponsor me for the next one!