FAT12 from scratch, and "the ACTUAL root cause": what confidently wrong looks like, five times over

Some bugs get fixed once. Some get fixed, un-fixed by the next thing that touches them, and fixed again in a different place, over and over, in a way that only makes sense once you can see the whole sequence at once. The FAT12 filesystem driver, on January 23, is the clearest example I have of the second kind, and it is worth walking through commit by commit, because the individual commit messages tell the story better than I can summarize it.

The setup

By v3.10.0, the OS could mount a FAT12 floppy image and read files from it, at least in principle. What followed was a single day, January 23, of the driver being fixed, again and again, by an agent that was confident each time and correct roughly once.

In order:

  1. 9010be37 "Fix FAT12 file search to skip VFAT long filename entries"
  2. a4dbb5e3 "Skip volume label entries in FAT12 directory search"
  3. 43831526 "Fix critical stack offset bug in FAT12 filename comparison"
  4. 7af69fc5 "Fix stack offset calculation - account for ALL pushes before cmpsb"
  5. 4057fb14 "Fix stack offset by pushing all regs BEFORE calculating DI"
  6. 8b52e1bd "Complete rewrite of FAT12 filename comparison logic"
  7. 09124699 "Fix CRITICAL segment register bug in FAT12 directory search"

Four separate attempts (3, 4, 5, and eventually 6, a full rewrite) at the same category of bug, a stack offset being computed wrong somewhere in the filename comparison routine, each one framed as the fix, each one followed by another attempt at the same class of problem. This is real-mode x86 assembly, where a stack offset is just an arithmetic constant you have to get exactly right by counting how many registers were pushed before it, and there is no compiler to catch you if you count wrong. An agent reasoning about "how many bytes have been pushed onto the stack at this point in the function" is doing arithmetic that is trivial to state and easy to get wrong once, and this sequence shows it being gotten wrong at least three times over, each time with full confidence, before commit 8 finally rewrote the routine from scratch rather than patching it again.

The commit that gives this post its title

Then, da2b13f5: "Fix CRITICAL bug: hardcoded sector 19 instead of using root_dir_start." A real bug, a good catch: the code was using a magic number where it should have used a computed value. And then, immediately after it, this:

437fdbbd Fix BROKEN CHS conversion - the ACTUAL root cause! CRITICAL BUG - COMPLETELY BROKEN LBA->CHS CONVERSION: The previous code was treating the LBA as if it were already in some packed CHS format... For LBA 19 (typical root directory start): Old: C=0, H=0, S=20 (INVALID, only 18 sectors/track). New: C=0, H=1, S=2 (CORRECT).

I want to be fair to this commit: the diagnosis it describes is real and correct. The logical-block-address to cylinder-head-sector conversion really was broken in exactly the way described, and fixing it really did matter. The commit message even explains, correctly, why the old code produced an invalid sector number that would fail an INT 13h disk read. This is good, specific, verifiable engineering reasoning, presented with total confidence, and the words "the ACTUAL root cause" are doing real work in that sentence: they are a claim that the search is over.

It was not over. The next day, January 24, five more fixes landed against the same filesystem code: a stack offset corrected from 12 bytes to 14 (111ca7d1), a stack misalignment causing an infinite loop (136be3de), a stack cleanup bug causing a different infinite loop (5a7057b3), a debug display routine that was itself corrupting the stack it was trying to inspect (304d350a), and a further stack cleanup fix in the LBA-to-CHS conversion routine itself, the same routine the "ACTUAL root cause" commit had just rewritten (de0d8faf, "debug09: Fix LBA to CHS conversion in fat12_read").

The CHS math was, in fact, fixed correctly by the commit that said so. It was also, provably, not the only thing wrong, and it was not even the last bug in the same function.

The lesson

I do not tell this story to make the agent look bad. Every individual diagnosis in this sequence, read on its own, is competent: specific, falsifiable, backed by a concrete before-and-after byte pattern. The failure is not in any single commit's reasoning. It is in the confidence of the framing, "the ACTUAL root cause," stated as if finding one true cause and fixing the rest of the bugs were the same event. They were not. Real-mode assembly bugs cluster. A stack offset error in one place is evidence that the calling convention around it is fragile everywhere, not evidence that you have found the one problem.

What I changed after this, and what shows up in the discipline the rest of this series describes, is: treat "root cause" as a claim to verify, not a status to accept. A commit message calling something the actual root cause should trigger more scrutiny, not less, because it is exactly the phrasing an agent reaches for when it has found a real, satisfying, plausible answer and stopped looking before checking whether anything else nearby depends on the same broken assumption. The concrete practice that grew out of this, and that the later port work formalized as conformance tests and instruction-level verification harnesses, is simple to state and was hard to earn: after a fix is declared, run the thing that was supposedly fixed, on the real target, and look at what actually happens, before believing the commit message.

That practice is what the rest of this series is actually about. The next several posts move past these early war stories into how "verify, don't trust" got built into the project's actual infrastructure: contracts, executable conformance tests, and harnesses that render a screen and diff the pixels instead of reading a claim and moving on.

Source: https://github.com/hmofet/unodos, commits 9010be37 through de0d8faf, January 23-24.

The OS in this piece runs in your browser. No install, no sign-up: boot it in a tab, or download it for any of 22 machines.

Get the next one

New essays roughly every other week: the war stories, the method, and the receipts. No spam, unsubscribe in one click.

Prefer a reader? RSS.