LLM digest: March 2026
Sent
I published 68 posts on my blog in March - plus 20 releases, 10 tools, 10 research reports and 2 museums. Here's your sponsors-only summary of the most important trends and highlights from the past month.
As always, this issue and previous issues are archived in my simonw-private/monthly GitHub repository.
More agentic engineering patterns
I added 8 new chapters to my Agentic Engineering Patterns guide:
- 2nd: GIF optimization tool using WebAssembly and Gifsicle - the first in a new collection of annotated prompts.
- 4th: Anti-patterns: things to avoid - don't inflict giant unreviewed PRs on your collaborators!
- 6th: Agentic manual testing - having agents "manually" test code (using curl etc) in addition to writing automated tests can help find all sorts of interesting edge-case bugs.
- 10th: AI should help us produce better code - something of a personal manifesto. Shipping worse code with agents is a choice.
- 15th: What is agentic engineering? - since I'm writing a whole book about it I figured I should define it somewhere!
- 16th: How coding agents work - the fundamentals.
- 17th: Subagents - a handy feature of Claude Code that's now available in Codex too.
- 21st: Using Git with coding agents - coding agents mean we can use all of those features of Git that previously had a little too much friction.
I gave a talk about Agentic Engineering at the Pragmatic Summit in February, and they've now published the video - so I wrote up my highlights from that talk.
Streaming experts with MoE models on a Mac
Dan Woods kicked off a fascinating spurt of collaborative research when he used Andrej Karpathy's autoresearch pattern to implement and then optimize Apple's 2023 LLM in a flash paper against current hardware and the latest models.
LLM in a flash describes a pattern I call streaming experts. It notes that Mixture-of-Experts models only need a subset of their parameters in memory at a time, and if you have a fast enough SSD you can stream them into memory for every token.
Dan used this to get Qwen3.5-397B-A17B 4-bit - a 209GB model - running at 4.36 tokens/second on a 48GB Mac.
This inspired a flurry of activity which so far has seen Kimi K2.5 (1T parameters) running on a 96GB Mac and Qwen3.5-397B-A17B running (very slowly) on an iPhone!
Model releases in March
- 3rd: Gemini 3.1 Flash-Lite - Google's latest model is an update to their inexpensive Flash-Lite family, dirt-cheap at $0.25/million tokens of input and $1.5/million output - 1/8th the price of Gemini 3.1 Pro. Here are the pelicans.
- 5th: Introducing GPT‑5.4 - Two new API models: GPT-5.4 and GPT-5.4 Pro. These are the new OpenAI flagship models and they're very, very good. OK pelicans too.
- 16th: Introducing Mistral Small 4 - Apache 2 licensed 119B parameter (Mixture-of-Experts, 6B active) model. Not a great pelican.
- 17th: GPT-5.4 mini and GPT-5.4 nano, which can describe 76,000 photos for $52 - I wrote about these in a bit more detail, the nano model is even cheaper than Gemini 3.1 Flash-Lite.
- 30th: Mr. Chatterbox is a (weak) Victorian-era ethically trained model you can run on your own computer - a really fun tiny model pre-trained exclusively on British Library texts published between 1837 and 1899. I wrote a plugin to run this one in LLM.
Plus a bunch I haven't written about yet, most notably the new 1-bit Bonsai from PrismML, a tiny model family which claims "order-of-magnitude improvements in intelligence density".
The Qwen 3.5 model family came out in February and are rapidly gaining respect as people spend more time with them. Sadly the research team behind them lost some key figures at the start of March - I wrote about that in Something is afoot in the land of Qwen.
Vibe porting
Vibe porting is one of several emerging names for the practice of porting code from one language to another using LLMs. It's similar in spirit to clean-room implementations, where code is turned into a spec and that spec is used to write fresh code, potentially under a different license than the code it is replicating.
I wrote about a high profile instance of this in Can coding agents relicense open source through a “clean room” implementation of code?. Dan Blanchard, the long-time maintainer of the chardet library, released a brand new version that was a complete AI-assisted rewrite, and dropped the old LGPL license in favour of MIT.
chardet's original author Mark Pilgrim broke his 15 year internet silence to object to the relicensing.
The resulting GitHub issues thread eventually attracted this comment from LGPL v3 license co-author Richard Fontana who said that no one "has identified persistence of copyrightable expressive material from earlier versions in 7.0.0" and hence the LGPL v3 had likely not been violated in this case.
Meanwhile Cloudflare, fresh off their rewrite of Next.js, have set their sights on WordPress.
Supply chain attacks against PyPI and NPM
March was an awful month for supply chain security.
CI security scanning tool Trivy got compromised, and since their software runs in CI in many other projects this started a wave of further attacks.
The most prominent of these was an attack on LiteLLM, where a credential stealer was published to PyPI for 46 minutes (before being quarantined) but still got downloaded 46,996 times.
Then yesterday another stolen credential attack hit Axios, an NPM package with 101 million weekly downloads, again adding a credential stealer to that package.
Now is a great time to familiarize yourself with the various dependency cooldown options for modern package managers - configuration options for saying "only upgrade to package releases that are at least X days old".
In related packaging news, OpenAI are acquiring Astral, the company behind Python package manager uv. I have thoughts!
Stuff I shipped
I spent March mainly focused on adding file upload support to Datasette and integrating Datasette more deeply with LLM.
- 7th: datasette-table-diagram 0.1a0 - Show Entity Relationship diagrams of tables in Datasette
- 7th: dclient 0.5a3 - A client CLI utility for Datasette instances
- 9th: llm-tools-edit 0.1a0 - LLM plugin providing tools for editing files
- 17th: llm 0.29 - Access large language models from the command-line
- 18th: datasette 1.0a26 - An open source multi-tool for exploring and publishing data
- 23rd: datasette-files 0.1a2 - Upload files to Datasette
- 25th: datasette-llm 0.1a1 - LLM integration plugin for other plugins to depend on
- 25th: datasette-files-s3 0.1a1 - datasette-files S3 backend
- 25th: datasette-files-s3 0.1a2
- 26th: datasette-llm 0.1a2
- 27th: datasette-showboat 0.1a2 - Datasette plugin for SHOWBOAT_REMOTE_URL
- 30th: llm-mrchatterbox 0.1 - Chat with Mr Chatterbox, trained on a corpus of over 28,000 Victorian-era British texts published between 1837 and 1899
- 30th: llm-mrchatterbox 0.1.1
- 30th: datasette-llm 0.1a3
- 30th: datasette-files 0.1a3
- 31st: llm-echo 0.3 - Debug plugin for LLM providing an echo model
- 31st: llm-echo 0.4
- 31st: llm 0.30
- 31st: llm-all-models-async 0.1 - Register async versions of models from LLM plugins that only provide a sync version
- 31st: datasette-llm 0.1a4
What I'm using, March 2026 edition
I'm still defaulting to Claude Code and Opus 4.6, but I've been very impressed by GPT-5.4 and have been running that through Codex. I've found myself having GPT-5.4 review code written by Claude a fair bit - it's particularly good at security reviews.
I upgraded to a new 128GB M5 Max MacBook Pro, which has very much rekindled my interest in local models. I can spare a solid 100GB now and still have space to run other applications.
I'm mostly using LM Studio for this, but I'm also exploring Ollama's launch command for conveniently launching Claude Code or Codex against a local model.
gpt-oss 120B may be 250+ days old but it's still a very impressive local model. I'm having fun with the larger members of the Qwen 3.5 family too.
I've been using Claude skills a little more, as described in Experimenting with Starlette 1.0 with Claude skills. I've also been having a lot of fun Vibe coding SwiftUI apps.
And a couple of museums
I spent 3.5 days in New York City this month, which was enough time to tick two tiny museums off my list:
The New York Earth Room is an apartment filled with 280,000 pounds of soil, most of it still the original soil that was installed there in 1977.
The John M. Mossman Lock Collection at the General Society of Mechanics and Tradesmen of the City of New York is a fabulous appointment-only collection of locks, mostly from bank vaults and almost all of which were collected prior to 1928.
That's it for March!
If this newsletter was useful feel free to forward it to friends who might find it useful too, especially if they might be convinced to sign up to sponsor me for the next one!