Fun with Unicode
13th September 2002
Hixie has submerged himself in Unicode. Stuart muses that the reason Unicode is so (potentially) huge is a legacy of the Y2K problem. I prefer the explanation given in XML in a Nutshell (my current reading matter of choice for three-and-a-half-hour-train-journeys-from-hell):
Unicode can potentially hold more than a million characters, but no one is willing to say in public where they think most of the remaining million characters will come from. *
* Footnote: Privately, some developers are willing to admit that they’re preparing for the day when we’re part of a Galactic Federation of thousands of intelligent species
More recent articles
- Reverse engineering Codex CLI to get GPT-5-Codex-Mini to draw me a pelican - 9th November 2025
- Video + notes on upgrading a Datasette plugin for the latest 1.0 alpha, with help from uv and OpenAI Codex CLI - 6th November 2025
- Code research projects with async coding agents like Claude Code and Codex - 6th November 2025