512 bytes to a graphical hello world, and the font bug that took 14 builds before breakfast

The UnoDOS repository opens on January 22 with a commit containing no code at all: "Initial project setup with CLAUDE.md," the instruction file for the AI agents that would write everything after it. The instructions existed before the software did, which is the whole project in miniature. Two commits later, the same day, the code starts: v0.2.0, "Implement boot loader with graphical Hello World." Within hours, the version number climbed to v0.2.4, each step a small fix: floppy write tooling, a QEMU boot compatibility issue, a graphics corruption bug. Then, later the same day, a rewrite: v3.0.0, "Welcome to UnoDOS 3!" A fresh start on the same idea.

What followed, across the next few hours, is one of the cleanest small case studies I have from the whole project in what an AI coding agent looks like when it is confidently, repeatedly wrong about something very small.

The bug

The task was ordinary: render a bitmap font, one character at a time, to a VGA framebuffer. By v3.1.0 the agent had implemented "Complete ASCII bitmap font set," and small fixes followed for a clock display and a demo of the character set. Then, at v3.1.5, the commit message is "Fix version text corruption from overlapping elements." Text on screen was landing wrong.

What happened next is a sequence I have not seen described honestly very often, because it does not flatter anyone: four raw debug commits (DEBUG: Simplified clock to test draw_ascii_4x6, DEBUG: Add static TEST at demo position Y=120, DEBUG: Move demo to Y=95, DEBUG: Add Y position markers), followed by a version bump to v3.2.0 that split the kernel into a three-stage boot architecture, and then this, in order, all on January 22 and into the 23rd:

Fourteen numbered attempts, each one a hypothesis about a font rendering bug, each one tested, and most of them wrong or incomplete, before a real fix landed as v3.2.1, "Fix font rendering by reorganizing memory layout." Even that needed a follow-up correction the next commit, and another the one after that, before the welcome message finally rendered correctly using what the commit calls "the working character range."

What was actually happening

Font rendering bugs in real-mode assembly on a VGA framebuffer are a genuine minefield: label addressing versus computed offsets, off-by-one errors in a glyph table, segment boundaries that silently change what a label resolves to depending on where the linker placed it. There is no exception thrown when you read the wrong byte of a font table. You just get the wrong glyph, or a garbled one, or nothing, and you have to reason backward from the pixels to the cause.

What the fourteen builds actually show is an agent doing something close to binary search by hand, in public, one commit at a time: hardcode a known-good character to rule out the font table itself (v3.2.0.2, v3.2.0.3), narrow where the addressing broke down by testing offset boundaries (v3.2.0.7, v3.2.0.8, v3.2.0.11), then isolate a single suspect pair of characters (v3.2.0.12, v3.2.0.13) once the range narrowed. That is not a bad process. It is close to correct debugging methodology. What makes it a story worth telling is how many of the individual steps in the middle were confidently framed as fixes ("Fix font rendering for all character offsets," v3.2.0.9) and were not, in fact, the fix, because two more numbered attempts followed it.

The lesson

The specific thing I took from this, and applied for the rest of the project, is: version the debug attempts, not just the features. It would have been easy, and is the default instinct in most projects, to squash all fourteen of those commits into one clean "fix font rendering" commit after the fact and move on. I did not, and I am glad, because that sequence is the actual record of an agent's reasoning under uncertainty, and it is more useful to me than any summary of it.

Two things follow from that. First: when an agent says "fixed," that claim has a shelf life of exactly until the next screen render. Treat it as a hypothesis, not a fact, until you have looked at the actual output. Second: a numbered sequence of small, falsifiable attempts, each one committed separately, is a better safety net than one large confident rewrite. If v3.2.0.9 had been the only commit, and it had been wrong (which it was), there would have been no record of what was tried before it, and the next attempt would have had to start from scratch instead of from "these five things we already ruled out."

This pattern, an agent confidently declaring victory and being wrong, then doing it again, is not unique to font rendering. It shows up again about three weeks later in the mouse driver, in a much stranger and more expensive form, because that one could not be diagnosed by staring at pixels. It needed a raw stack dump. That is the next post.

Source: https://github.com/hmofet/unodos, commits 4f7a0739 through d91c9f47, January 22-23.

The OS in this piece runs in your browser. No install, no sign-up: boot it in a tab, or download it for any of 22 machines.

Get the next one

New essays roughly every other week: the war stories, the method, and the receipts. No spam, unsubscribe in one click.

Prefer a reader? RSS.