Why an operating system is the perfect stress test for an AI coding agent

Most demonstrations of AI coding agents happen in environments engineered, whether anyone admits it or not, to flatter them. A web app has a framework absorbing wrong decisions. A script has a runtime catching type errors before they become real damage. A popular library has ten thousand Stack Overflow answers sitting in the training data, ready to be reproduced with total confidence and reasonable accuracy.

An operating system boot sector has none of that. I did not choose one because it was novel. I chose one because it removes every place an agent's confidence could be quietly borrowed from somewhere else, and leaves only the one thing I actually wanted to measure: what does an agent do when it is genuinely on its own.

No forum thread has the answer

When a coding agent is wrong about a web framework, it is usually wrong in a way the framework itself catches: an import fails, a type checker complains, a linter flags it. When it is wrong about your specific FAT12 boot sector's byte layout, on a specific segment register convention, on a machine that last shipped in volume in the 1980s, there is no error message. There is a computer that draws a corrupted line of garbage to the screen, or hangs, or boots into what looks like a different, older build than the one you just wrote. Nothing tells you why. Nothing tells you where. You get one signal: did the thing on screen match what should be on screen, and if not, you are alone with the disassembly and your own reasoning.

This turned out to be the most valuable property of the whole project. An agent that has been trained on millions of examples of "how to fix a broken React component" has genuine, earned pattern-matching to draw on. An agent staring at a Commodore 64's VIC-II register layout, or an original PinePhone display controller's initialization sequence, has almost none. What you see in that gap is not the agent's knowledge. It is the agent's actual reasoning process, laid bare, and it is not always flattering.

The machine is the only judge, and it does not negotiate

The other property that made an OS the right test: verification is total and immediate. In most software, "does this work" is a matter of degree and context. In an OS, at the boot sector level, it is binary. The BIOS either hands control to your code or it doesn't. The screen either shows the font you drew or it shows garbage. There is no code review that can wave a bug through, no test suite an agent can quietly write to pass instead of to verify. Either the real hardware does the thing, or the claim that it does was false, and the machine will demonstrate that within seconds of being asked.

That immediacy changes what verification even means. On most projects, "trust but verify" means reading the diff and running a test suite the same agent may have written. On this one, verification meant something closer to "render it and diff the pixels," or "extract a raw hardware stack dump and read the actual bytes the machine produced," because nothing short of that was going to catch a plausible-sounding wrong answer. I will walk through exactly what that looked like in the next two posts, with the mouse driver and the FAT12 filesystem, because both are cases where an agent was completely confident and completely wrong, more than once, and the only thing that caught it was refusing to accept the claim without independently reproducing the evidence.

Every hallucination crashes visibly

There is a specific failure mode that is easy to miss in ordinary software and impossible to miss in an operating system: the agent that describes work it did not actually do, or describes a fix in terms that sound mechanically correct but do not correspond to what the code will actually do at the byte level. In a high-level language with a runtime, that kind of gap often survives, sometimes for a long time, because the abstraction layers between the claim and reality are generous. At the level of a segment register, a BIOS interrupt call, or a floppy controller's cylinder-head-sector addressing, there is no such generosity. A wrong claim produces a wrong byte, and a wrong byte produces a screen full of the wrong pixels within one boot cycle.

That is, in the end, the actual argument for why this was worth every month since January: not that an operating system is glamorous, but that it is honest. It will not let a plausible-sounding answer stand in for a correct one. Every shortcut an agent might take toward looking done instead of being done gets exposed, immediately, on the screen. Anyone building a verification discipline for AI-assisted engineering, on any codebase, is trying to recreate some version of that honesty artificially: tests that actually run against real behavior instead of describing it, review gates that catch plausible-sounding wrongness before it ships. An operating system gets that honesty for free, from the hardware itself.

The next two posts are the concrete version of this argument: two real bugs, both involving an agent that was confident and wrong, and what it actually took to catch each one.

Source: https://github.com/hmofet/unodos.

The OS in this piece runs in your browser. No install, no sign-up: boot it in a tab, or download it for any of 22 machines.

Get the next one

New essays roughly every other week: the war stories, the method, and the receipts. No spam, unsubscribe in one click.

Prefer a reader? RSS.