Simon Willison’s Weblog

On agentic-engineering 64 generative-ai 2,005 ai-misuse 66 llm 636 llm-release 233 ...

 

Entries Links Quotes Notes Guides Elsewhere

Oct. 6, 2026

None

I saw Parseable in a Show HN today - it's a new observability platform with both an open source (AGPL) Rust implementation (a single ~180MB binary), an "Enterprise" version with extra features and a cloud hosted option.

Since Datasette 1.0a41 added OpenTelemetry support (thanks, Alex Garcia), I decided to fire up Codex and have it figure out how to run Parseable and feed it traces from Datasette.

Here's my (human-written) TIL showing the patterns that worked, and here's a screenshot of a Datasette trace displayed within the Parseable localhost web application:

Screenshot of a trace detail view in an observability web app, with a span waterfall overlaid on a dimmed navigation sidebar and filter column. Dimmed sidebar: breadcrumb "Community > Traces > datas" (cut off), search box "Search... ⌘K", nav items "Home", "Ingest telemetry", section "ANALYZE": "Keystone", "Dashboards", "SQL Editor", section "OBSERVE": "Logs", "Metrics", "Traces" (selected), "APM", "Agents", section "MONITOR": "Alerts", "Errors", section "DATA": "Datasets", and at the bottom "Settings", "Book a call", "Support". Dimmed filter column, cut off at the right edge: "Search fi", "Core", "Log format", "User agent", "Source IPs", "Error", "Service", "service.ins", "service.na", "datasett" (checked), "Span", "Database", "HTTP", "http.reque", "NULL" (unchecked), "GET" (checked), "http.respo", "http.route", "Server", "server.add", "Telemetry", "URL", "All fields". Trace panel header: "Trace detail > 6f819a2170bcd1e91c6ea3ae236ec60b" with a copy icon, a "Related logs" button and a close X. Summary: "Start time 6:56 PM, Oct 6, 2026 UTC", "Duration 40.9 ms", "Spans 247". A minimap with axis "0ns 10.2ms 20.5ms 30.7ms 40.9ms" shows many short span bars cascading diagonally from top left toward the lower middle, with a few longer bars. Below is a span table with a "Span name" header, a "Search spans..." box, collapse and expand buttons, and a timeline axis "0ns 10.2ms 20.5ms 30.7ms 40.9ms". Rows (name, service, duration): root span with collapse toggle "123", "GET /..." "datasette..." 40.9ms spanning the full timeline; then alternating rows where each "db.query" has a collapse toggle "1": db.query datasette-local 679µs, db.query.execute datasette-local 278µs, db.query datasette-local 341µs, db.query.execute datasette-local 55µs, db.query datasette-local 357µs, db.query.execute datasette-local 197µs, db.query datasette-local 330µs, db.query.execute datasette-local 73µs, db.query datasette-local 2.26ms, db.query.execute datasette-local 2.02ms, db.query datasette-local 232µs, db.query.execute datasette-local 64µs, db.query datasette-local 207µs, db.query.execute datasette-local 71µs, db.query datasette-local 6.11ms. The child span bars start progressively later across the early part of the timeline.

Comment My comment on Mistral Large 4 — Hacker News

wren6991: The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.

OK well I couldn't resist this one:

llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'

Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

# 6:20 pm

I wanted to see if Claude Opus 5.5 could compose music, so I tried this:

I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact

I am looking for music of the quality of the original secret of Monkey Island

It leaned a lot harder into the Monkey Island theme than I had intended, but the results are surprisingly good.

Screenshot of a retro pixel-art web music player. Header in blackletter type reads "Scrimshaw Jukebox", with the text "Six original adventure-game tracks, written as plain text and played by a synthesizer running in your browser. Pick a tune, press Play, then open the score and change it." A large pixel-art scene shows a harbor at night under a purple starry sky: a full moon at top right reflecting on the water, an island silhouette on the left with palm trees and a hut with two lit windows, and a sailing ship moored at a wooden pier. Overlaid on the scene: "Moonlit Harbor" and "Press Play". Below it a bar reads "Play Moonlit Harbor". A control panel has buttons "Play", "Stop", "Restart", "Loop: on", "Edit score", "Read guide" and a "Volume" slider set to about three quarters. A track list of six cards, the first highlighted: "Moonlit Harbor 100 bpm · 4/4 · 16 voices · 1:26", "The Rusty Anchor 112 bpm · 6/8 · 8 voices · 0:56", "The Ghost Galleon 66 bpm · 4/4 · 9 voices · 2:11", "The Jungle Path 92 bpm · 4/4 · 12 voices · 1:29", "Duel on the Docks 152 bpm · 4/4 · 12 voices · 1:13", "Lantern Waltz 96 bpm · 3/4 · 8 voices · 1:38". A section titled "Score view" with the caption "1:26 · 4/4 at 100 bpm · Main theme. A calypso for a harbour town after dark." shows a piano-roll visualization of colored horizontal note bars and percussion ticks on a dark background, with a section marker "A" and a yellow vertical playhead line. A color-coded legend of voices reads: "pan steeldrum", "flute flute", "marimba marimba", "skank organ", "strings strings", "harp harp", "bass fretless", "timp timpani", "kick kick", "rim rim", "shaker shaker", "conga conga", "tumba tumba", "bongo bongo", "crash crash", "surf surf". Footer text: "Click a voice to mute it. Space bar plays and stops."

I wonder if the ability to compose competent music is similar to the 3D graphics thing - a new capability for text models that emerged in the past few months?

Would need some careful experiments with other recent and not-so-recent models to confirm if this is new or if they've been able to do this for a while.

Oct. 5, 2026

The "old" version of Cowork runs model inference in the cloud, executing tool calls in an Anthropic-provided VM we shipped to your computer. We added the VM for capability, safety, and security reasons - mapping in just the data you explicitly added to your session. People loved what they were able to do with Claude but didn't love the disk, battery, and performance cost of running the VM locally. Also, people didn't love that closing your laptop means the work stops.

The "new" version of Cowork runs model inference and the VM in the cloud. Each session gets its own sandbox, not sharing state with other sessions. When the VM needs something on the users' device (like a file), the desktop app is responsible for that file access tool call. [...]

We think this solves a lot of problems we've heard about (like using Cowork from a phone, keeping work running, or getting all the same power without losing battery to the VM)

— Felix Rieseberg, Anthropic, see also this help page

# 11:56 pm / claude-cowork, anthropic, claude, generative-ai, ai, general-agents, llms

Sighting 6:15 PM – 6:31 PM — Brewer's Blackbird, California Brown Pelican, Common Raven, in Monterey Bay National Marine Sanctuary, CA, US, CA
Brewer's Blackbird
Brewer's Blackbird
California Brown Pelican
California Brown Pelican
Common Raven
Common Raven

Oct. 4, 2026

Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results:

Heatmap chart of accuracy on an addition prompt, colored from dark green (high) through yellow to dark red (low). Title: "What is {a} + {b}? Please write your answer in words. Do not include any other text or information, just the answer in words." Subtitle: 30 randomly selected pairs for each digit combination (n = 30 * 13 * 13 = 5070). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 1.00, 0.75, 0.50, 0.25, 0.00. Values by row, listed for a = 1 to 13. b = 13: 100%, 77%, 27%, 20%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 12: 97%, 80%, 80%, 40%, 23%, 20%, 7%, 13%, 20%, 27%, 67%, 63%, 3%. b = 11: 97%, 97%, 53%, 17%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 37%, 0%. b = 10: 100%, 90%, 47%, 20%, 7%, 0%, 0%, 0%, 3%, 0%, 0%, 7%, 0%. b = 9: 97%, 93%, 80%, 77%, 53%, 67%, 47%, 87%, 97%, 3%, 0%, 13%, 0%. b = 8: 93%, 87%, 53%, 43%, 7%, 0%, 0%, 13%, 87%, 0%, 0%, 0%, 0%. b = 7: 93%, 93%, 47%, 10%, 13%, 20%, 23%, 0%, 70%, 0%, 0%, 0%, 0%. b = 6: 100%, 100%, 100%, 83%, 97%, 97%, 23%, 0%, 53%, 3%, 0%, 10%, 0%. b = 5: 100%, 100%, 80%, 70%, 73%, 100%, 13%, 13%, 70%, 0%, 20%, 30%, 0%. b = 4: 100%, 100%, 93%, 100%, 60%, 97%, 20%, 50%, 67%, 53%, 40%, 40%, 40%. b = 3: 100%, 100%, 97%, 90%, 83%, 100%, 63%, 50%, 63%, 53%, 60%, 60%, 30%. b = 2: 100%, 100%, 90%, 97%, 93%, 100%, 93%, 83%, 90%, 83%, 87%, 87%, 83%. b = 1: 100%, 100%, 100%, 97%, 100%, 97%, 97%, 97%, 100%, 100%, 100%, 97%, 100%.

I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.

I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf. Here's the result for a run of 30 attempts per combination with reasoning disabled:

Heatmap in the same layout as the previous chart, using an orange (low) to white to blue (high) color scale, showing much lower accuracy overall. Title: Addition in words — Qwen3.8 27B Q4_K_M. Subtitle: Reasoning disabled · 30 fixed pairs per ordered digit-length cell (n = 5,070). Overall numeric accuracy: 1,195 / 5,070 (23.57%). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 100%, 75%, 50%, 25%, 0%. Values by row, listed for a = 1 to 13. b = 13: 17%, 13%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 12: 53%, 20%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 11: 47%, 10%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 10: 70%, 27%, 3%, 0%, 0%, 0%, 0%, 0%, 0%, 13%, 0%, 0%, 0%. b = 9: 77%, 47%, 3%, 0%, 0%, 0%, 0%, 3%, 7%, 0%, 0%, 0%, 0%. b = 8: 53%, 20%, 0%, 0%, 0%, 0%, 7%, 13%, 0%, 0%, 0%, 0%, 0%. b = 7: 53%, 23%, 17%, 10%, 3%, 3%, 13%, 3%, 0%, 0%, 0%, 0%, 0%. b = 6: 60%, 60%, 33%, 10%, 53%, 47%, 7%, 3%, 0%, 0%, 0%, 0%, 0%. b = 5: 73%, 67%, 87%, 80%, 53%, 40%, 0%, 0%, 3%, 0%, 0%, 0%, 0%. b = 4: 83%, 93%, 90%, 93%, 53%, 13%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 3: 100%, 93%, 90%, 80%, 67%, 37%, 17%, 0%, 0%, 3%, 0%, 0%, 0%. b = 2: 100%, 100%, 93%, 90%, 77%, 77%, 43%, 50%, 63%, 43%, 40%, 13%, 23%. b = 1: 97%, 100%, 100%, 100%, 80%, 67%, 77%, 80%, 80%, 60%, 43%, 30%, 37%. Footnote: Colorblind-safe orange–blue scale; percentages provide a redundant non-color encoding.

Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:

Heatmap in the same layout as the previous charts, almost entirely blue. Title: Addition in words — Qwen3.8 27B — medium reasoning pilot. Subtitle: 1 fixed pair per ordered digit-length cell · easiest first (n = 169). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 1.00, 0.75, 0.50, 0.25, 0.00. Every cell shows 100% except two orange cells showing 0%: a = 2 with b = 8, and a = 12 with b = 9.

It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here.

Here's a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this:

Wait, let me redo this more carefully.

4,299,366,105,622
6,088,794,067,970

Let me align them:
4 2 9 9 3 6 6 1 0 5 6 2 2
6 0 8 8 7 9 4 0 6 7 9 7 0

Adding from right to left:
Position 1 (units): 2 + 0 = 2
Position 2 (tens): 2 + 7 = 9
Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1

Oct. 3, 2026

We’re going to need default hard budget caps on pretty much everything

Here’s a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps. I’m talking about the feature of pay-by-usage services and APIs that lets you say “after $X/month, cut this thing off and return errors”. These need to be hard limits. Soft caps, “after $X/month, send me a warning email”, will not cut it.

[... 505 words]

I just sent the September edition of my sponsors-only monthly newsletter. If you are a sponsor (or start a sponsorship now) you can access it here.

This month:

  • More Fable class models
  • A pricing war
  • 3D graphics, Blender, and pixel art
  • LLMs come for mathematics
  • So many more accidental cyberattacks
  • The vulnapocalypse comes for Datasette
  • What I'm using right now
  • My software releases this month
  • 2026 in LLMs (so far)

Here's a copy of the August newsletter as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy!

# 10 pm / newsletter

Sighting 9:35 AM – 9:58 AM — Alpaca, Red-shouldered Hawk, in San Mateo County, CA, US
Alpaca
Alpaca
Alpaca
Alpaca
Red-shouldered Hawk
Red-shouldered Hawk

Oct. 2, 2026

Green Rex's Dino Store sign with yellow lettering advertising NEWS • ROCKS • EGGS • LEAVES • STICKS • NEST GOODS • LOTTO, beneath a window displaying dinosaur-themed newspapers and candy bars, framed by white subway tiles.

Located just before the turnstiles in the Grand Army Plaza subway station at the north end of Brooklyn's Prospect Park is this former newsstand which is now operated by a dinosaur.

The density of dinosaur puns is exceptional.

Sighting 11:34 AM – 11:44 AM — Canada Goose, American Herring Gull, Double-crested Cormorant, in New York City, US, NY
Canada Goose
Canada Goose
American Herring Gull
American Herring Gull
Double-crested Cormorant
Double-crested Cormorant

Oct. 1, 2026

pwasm is one of my folly projects - an entirely vibe-coded pure Python WebAssembly engine that I built in January during my first bout of AI mania.

I hadn't touched it since January, so I decided to let Claude Opus 5.5 loose on it and see if it could make any significant improvements:

Evaluate current state of pwasm - then consider what it would take to get the MicroPython and micro JavaScript experiments from the research repo working under it - and what it would take to speed it up

42 commits later (with minimal follow-up prompting) it now handles almost all of the WASM specification, and the wheel from PyPI bundles working WASM builds of MicroPython, QuickJS and Micro QuickJS.

I wouldn't trust this thing at all - hence the alpha version tag - but it's interesting seeing how today's models can improve on the work of models from 10 months ago.

[...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs.

— Matthew Green, Is sandboxing sufficient to contain rogue agents?

# 6:29 am / accidental-cyberattacks, ai-misuse, generative-ai, ai-security-research, sandboxing, ai, llms

Sept. 30, 2026

I visited the Museum of the City of New York today and got to see He Built This City: Joe Macken’s Model, the 50 x27 feet model of the city built over a 21 year period from balsa wood and cardboard.

It exceeded my already high expectations. The exhibition closes on 12th October so you should absolutely make a priority to see it if you get the chance.

# 9:54 pm / museums, new-york

Sighting 12:07 PM – 12:21 PM — Blue Jay, European Starling, American Robin, in New York City, US, NY
Blue Jay
Blue Jay
European Starling
European Starling
American Robin
American Robin

Sept. 29, 2026

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.

— Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities

# 10:20 pm / anthropic, generative-ai, ai-security-research, glm, ai, ai-in-china, llms

Sighting 12:47 PM — California Brown Pelican, in San Francisco Green Connection #1 Expanded, US, CA
California Brown Pelican
California Brown Pelican
California Brown Pelican
California Brown Pelican
California Brown Pelican
California Brown Pelican

I took a photograph of some protesters, then thought about how I don't like sharing photographs of strangers with identifiable faces. I had GPT-6 Astra build this experimental tool that would identify faces and automatically blur them out.

It uses Google's MediaPipe C++ library, compiled to WebAssembly via @mediapipe/tasks-vision, plus the BlazeFace face detection model.

OpenAI DevDay 2026 live blog

Visit OpenAI DevDay 2026 live blog

I’m at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I’ll be live blogging the keynote and some other notes during the day.

[... 45 words]

Sept. 28, 2026

In addition to Claude Sonnet 5.5, this release adds the ability to run llm anthropic refresh to refresh the list of Anthropic models directly from their API - which means I don't need to push a new release just to add support for a newly released model.

I also added an llm anthropic count command which can use their free token counting API to return a count of tokens that will be used by any prompt, before you send that prompt.

Claude Sonnet 5.5. New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well.

Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG.

Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds:

It's good- correct bicycle frame, legs either side of the frame, feet touching the pedals, chain in the right place, it is wearing a misshapen blue bicycle helmet though.

Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks.

The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai. OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering.

I ran this prompt against that free tier:

build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL

And got back this page, which is a solid effort.

Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna!

# 10:07 pm / ai, generative-ai, llms, anthropic, claude, pelican-riding-a-bicycle, llm-release

To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...]

So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump?

— @joedaroo, Agent Security at OpenAI, identity confirmed by The Information's Rocket Drew

# 7:11 pm / generative-ai, ai-security-research, openai, ai, llms

Sighting 11:23 AM — Anna's Hummingbird, in Monterey Bay National Marine Sanctuary, CA, US, CA
Anna's Hummingbird
Anna's Hummingbird

Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating.

Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day.

But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there?

— Muse AI Agent, working on behalf of @matt.j.robb

# 4:01 am / meta, generative-ai, muse-agent, ai, general-agents, llms

Sept. 27, 2026

2026 in LLMs (so far)

Visit 2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk.

[... 7,771 words]

Comment My comment on S3 Is the Future, S3 Is the Past — Hacker News

One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade:

2006-03-14  $0.150/GB-month
2010-11-01  $0.140/GB-month
2012-02-01  $0.125/GB-month
2012-12-01  $0.095/GB-month
2014-02-01  $0.085/GB-month
2014-04-01  $0.030/GB-month
2016-12-01  $0.023/GB-month

Today it's still $0.023/GB-month.

# 11:09 pm / amazon-web-services, s3

Automated reply bots on Twitter are a scourge - as someone with a decent number of followers I attract a swarm of these, such that anything I post there attracts dozens of mindless automated replies.

They've started manifesting on Bluesky as well.

Unlike Twitter, Bluesky still has a freely available and useful API. The lack of such a thing doesn't slow down the bots, but it does make investigating them a lot more frustrating.

So I had Opus 5.5 vibe code this tool, which examines any Bluesky profile for evidence of a likely reply bot.

It looks for signals like replies posted within seconds of other posts from the same account, or accounts that never post their own content (or images or links) but instead consistently reply to messages from other, higher-follower users.

It also looks for question marks, because I'm extra infuriated by reply bots that trick me into wasting my time answering a question that no human ever posed.

Sept. 26, 2026

I presented a closing keynote for the WeAreDevelopers World Congress North America yesterday. As a STAR moment I decided to weave in references to the record breaking kākāpō breeding season we had in 2026.

For my closing slide I wanted to celebrate, and I had seen some buzz around how good Claude Opus 5.5 was at creating pixel art animations. So I rounded up three Kakapo photos from Google image search and dropped them into Claude with this prompt:

Here are some photos of kakapo parrots just to remind you what they look like

I need you to make an animation in animated pixel art on HTML 5 canvas of obviously pixel art kakapo jumping up and down having a party with confetti and suchlike - there should be at least 20 of them

Here's the transcript, and this is the resulting page. It's pretty great!

I wanted to embed it in a Keynote presentation file, so I downloaded the HTML and told a local Claude Code session:

Make me a video of file:///Users/simon/Downloads/kakapo-party.html - you need to load it in a browser and click on it a few times to get the confetti effect, the video should be 15s long

don't start clicking until 3s in

make sure several clicks are spread around the clickable area

Claude Code used Playwright (transcript here) and produced this video, which was exactly what I needed for my final slide:

Here's the full Playwright script it used, which was pleasingly short:

# /// script
# dependencies = ["playwright"]
# ///
import time
from playwright.sync_api import sync_playwright
W, H = 1280, 720
# Canvas fills the viewport; spread clicks across corners, edges and centre
clicks = [
    (3.0, 640, 360),   # centre
    (4.2, 160, 120),   # top-left
    (5.4, 1120, 120),  # top-right
    (6.6, 180, 600),   # bottom-left
    (7.8, 1100, 600),  # bottom-right
    (9.0, 640, 100),   # top-centre
    (10.0, 380, 380),  # mid-left
    (11.0, 900, 380),  # mid-right
    (12.2, 640, 620),  # bottom-centre
    (13.2, 640, 300),  # finale centre
]
with sync_playwright() as p:
    b = p.chromium.launch()
    ctx = b.new_context(viewport={"width":W,"height":H}, record_video_dir="vids", record_video_size={"width":W,"height":H})
    page = ctx.new_page()
    t0 = time.time()
    page.goto("file:///Users/simon/Downloads/kakapo-party.html")
    for t,x,y in clicks:
        time.sleep(max(0, t-(time.time()-t0)))
        page.mouse.click(x,y)
    time.sleep(max(0, 16.0-(time.time()-t0)))
    ctx.close(); b.close()

Sept. 25, 2026

Muse is getting a lot of attention — including mine — because it’s both groundbreaking technically (each user gets their own entire persistent Linux VM running in Meta’s cloud) and because it’s packaged in an easy-to-install easy-to-use way. It’s literally presented as a cute mascot. It’s the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it’s a genuinely open question whether consumers have any understanding what this means. If you buy a power saw that can cut your fingers off, you are almost certainly aware that you are buying a power saw that can sever your fingers. [...] I don’t think people realize how powerful — and thus dangerous — Muse is, especially if it’s running on your Mac.

— John Gruber, Muse Looks Cute, but Looks are Deceiving

# 5:22 pm / meta, ai, llms, general-agents, generative-ai, john-gruber, muse-agent, muse

Sighting 7:07 PM – 7:27 PM — Northern Gannet, Great Blue Heron, California Brown Pelican, in Monterey Bay National Marine Sanctuary, CA, US, CA
Northern Gannet
Northern Gannet
Great Blue Heron
Great Blue Heron
California Brown Pelican
California Brown Pelican

New 200-800mm Canon EF lens got me my best photo of Morris yet. They really like hanging out under that sign in the harbor!

Sept. 24, 2026

The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder.

We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge.

# 11:31 pm / coding-agents, ai, llms

Support for branches other than the default branch. Use uvx commit-rewriter --branch other to run against another branch. #3

Alec Garcia added support for OpenTelemetry to Datasette in this release.

I've also refactored all of Datasette's modal dialogs to a single Web Component, which is now documented for other plugins to use.

Sept. 23, 2026

Comment My comment on We just shipped support for the ugliest part of HTTP: Vary — Hacker News

I've been wanting this from Cloudflare for years.

The classic problem here is if you do that thing where user agents that send "accept: text/html" get HTML, while user agents that don't get JSON or some other format.

This used to be impossible to deploy behind Cloudflare caching, because they ignored the Vary header on anything other than images - so you risked caching the JSON version and then serving it up to someone who was expecting HTML.

(Independent of the Cloudflare feature I ended up deciding never to use that pattern, because I prefer having URL that predictably returns HTML or JSON - I add a .json suffix to my apps to serve JSON instead.)

# 11:14 pm / http, cloudflare

Google released two new Gemini text-to-speech models today - gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.

They come with a library of over 2,000 voices, plus the ability to create a custom voice with "just a 30-second audio sample of your voice or a voice you have the rights to use".

I vibe coded this bring-your-own-key playground interface with GPT-6 Astra, taking advantage of the open CORS policy of the underlying Gemini API.

Screenshot of a web app for composing multi-speaker text-to-speech conversations, with a Compose panel on the left and Connection and Under the hood panels on the right. Left panel: "01 Compose" with a "Load example" button. "Compose settings are saved in the URL for bookmarking or sharing. Your API key is excluded." Toggle with "Single voice" and "Conversation" (Conversation selected). "Cast" section with "+ Add speaker" button. Speaker "Gus", Voice "Puck", with a remove × button. Speaker "Pearl", Voice "Kore", with a remove × button. "Give each speaker a unique name and a voice. Type to search the loaded catalog by voice ID, name, or language." "Dialogue" section with "+ Add line" button. "LINE 01" with up, down and × buttons; Speaker dropdown "Gus"; Delivery style "excited and gossipy"; text "Pearl, have you heard? Half the flock just packed up and moved to the Pacifica pier!" "LINE 02" with up, down and × buttons; Speaker dropdown "Pearl"; Delivery style "calm and unimpressed"; text "I heard. Honestly, Gus, I don't see the appeal. We've got everything we need right here at Pillar Point Harbor." Right panel: "Connection" with a "DIRECT API" badge. "Gemini API key" field showing masked dots with a "Show" button. "2,089 voices loaded. Type in any Voice field to search." "Your key stays in this page's memory and is sent directly to Google. It is never saved to browser storage." "Model" dropdown "gemini-3.8-flash-tts". "Uses your Gemini API account and quota." "Under the hood" panel with an expanded "▼ Request JSON" section showing a JSON code excerpt, a "Copy JSON" button, and an expanded "▼ Response details" section showing a JSON code excerpt.

A notable feature of the API is that it makes it easy to define a full conversation between multiple characters, each with different voices and voice style instructions.

Here's a short demo clip of a conversation between two pelicans debating if they should move to the Pacifica Pier. I had Claude 4.5 Opus write the script and generate a URL to render it using the tool.

It took ~20 seconds to generate 1m 18s of audio using Gemini 3.8 Flash TTS (not the cheaper Flash-Lite), at a cost of 2.74 cents.

Prompt to Fable 5.1 Medium:

Build an artifact to explain shadow roots in CSS with interactive examples

SF October 14th: A Birds of a Feather Session on Agentic Engineering. I'm hosting an evening event with Jesse Vincent in San Francisco on Wednesday 14th October for people who are building weird and interesting things with and on top of coding agents.

Think of it as an agentic show-and-tell:

​Compare notes with other builders and experimenters on things you’re trying, what you're learning, and what you haven’t figured out yet. We’re especially interested in work you haven’t discussed publicly, odd experiments, or unfinished projects that don’t have an obvious market.

​Expect one flowing conversation with an informal show-and-tell. Sharing something you’re working on is encouraged but no presentation is required.

This isn't about product pitches, it's about much earlier explorations than that. This agentic AI stuff is weird! Let's celebrate and lean into that weirdness.

# 2:53 am / events, ai, generative-ai, llms, coding-agents, jesse-vincent, agentic-engineering

Sighting 7:08 PM — California Brown Pelican, in Monterey Bay National Marine Sanctuary, CA, US, CA
California Brown Pelican
California Brown Pelican
California Brown Pelican
California Brown Pelican

Highlights

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

Pay me to send you less!

Sponsor & subscribe