I ran AI coding agents like an engineering organization. They shipped an operating system you can boot in your browser.

Before the story, the artifact, because you do not have to take my word for any of this: UnoDOS is a graphical operating system written from scratch, and you can boot it in a browser tab right now, at unodos.arinbakht.com/try. It is the real OS under full CPU emulation, about a 17-megabyte download, on the desktop in six seconds on a fast machine and fifteen to twenty on an older laptop. It has its own drivers, its own TCP/IP stack with TLS, a web browser with two switchable JavaScript engines, a word processor and spreadsheet that read and write real Microsoft file formats (the old binary ones and .docx/.xlsx alike), a C compiler and IDE that build UnoDOS applications on UnoDOS, and UnoCode, a VS Code-class editor whose extensions are JavaScript running on the OS's own runtime. One of its applications is a Doom engine written in Python, running on the OS's own MicroPython runtime, and that is playable in a browser too, at unodos.arinbakht.com/duum.

No Linux underneath. No existing kernel. A boot sector up.

The experiment

In January I started an experiment. Could I run AI coding agents the way you would run a small engineering team, with standards, review, and institutional memory, on a project hard enough that shortcuts would be visible immediately?

I picked an operating system. Not because operating systems are a novel choice for a hobby project (they are one of the oldest), but because an OS has a property that most software doesn't: there is nowhere to hide. You cannot paste a Stack Overflow answer for your own kernel. There is no framework absorbing your mistakes. The machine either boots or it doesn't, and it does not care how the code got written.

Seventy-three active development days later, the modern build above is the flagship. And then there is the part I did not plan.

The middle act: the same OS, for 22 machines

Somewhere along the way, the experiment grew a second axis. The same source tree now builds and ships binaries for 22 platforms, verified against the repository's own claims ledger rather than from memory:

The full breakdown, including every gap and every platform I have not gotten to real hardware yet, lives in a file called PLATFORMS.md in the public repository. I built it as a discipline, not a brochure: it exists because I do not want to be the kind of person whose claims fall apart the first time someone checks them, and the only way to guarantee that is to make checking easy.

The part that needs saying plainly

I did not write most of this code by hand. I directed AI coding agents that did.

I want to say that once, clearly, and then spend the rest of this series being specific about what it actually meant, because both of the easy reactions to that sentence are wrong. It was not effortless, and it was not fake. What it actually was: a long argument with a very fast, very confident collaborator that is wrong constantly, and the discovery that the only way to make that collaborator useful is to stop trusting anything it says and start trusting only what the machine does when the code runs.

That distinction, claims versus verified behavior, turned out to be the entire game. An agent will tell you a bug is fixed. It will tell you with total confidence, citing a plausible root cause, using the right vocabulary. Sometimes it is right. Often, especially early on, it is wrong, and the telling difference between a project that works and one that quietly rots is whether anything downstream of that claim actually checks it against the real machine.

This series is going to walk through what that looked like in practice: the font bug that took fourteen numbered debug builds in one evening before I found the actual cause, the mouse driver that only started working after I extracted a raw stack dump and read it like a detective reads a crime scene, the afternoon a commit message confidently declared the "ACTUAL root cause" of a bug that then needed four more fixes. I am not going to clean these stories up. The failures are the point. A series about building verification discipline that hides its own failures would be lying about the thing it claims to teach.

The rules I held to

A few constraints made this a fair test rather than a demo:

What's next

The next post goes back to the actual reason an operating system was the right stress test in the first place: what happens to an AI coding agent when there is no manual for the machine it is targeting, no forum thread with the answer, and the only oracle available is whether the thing you built actually turns on.

Source: https://github.com/hmofet/unodos. The in-browser boot is at https://unodos.arinbakht.com/try/, Duum at https://unodos.arinbakht.com/duum/, and downloads for all 22 platforms and the platform ledger are both linked from the README.

The OS in this piece runs in your browser. No install, no sign-up: boot it in a tab, or download it for any of 22 machines.

Get the next one

New essays roughly every other week: the war stories, the method, and the receipts. No spam, unsubscribe in one click.

Prefer a reader? RSS.