← All writing

Mailroom writes back

Mailroom, the agent that reads my email every 15 minutes, has spent its whole career in a deliberately lopsided posture: allowed to read everything, allowed to say nothing. Two earlier posts covered its containment and how my corrections teach it; since the last one, its trust ledger filled in nicely. Over two weeks it filed 221 pieces of bulk mail out of my inbox on its own authority, and I didn’t have to rescue a single one. In this house privileges are earned one at a time, and the ledger said it was time for the next one: the mail that actually matters still waited on me, and I am a slower correspondent than anyone in my life deserves. So Mailroom now writes back. Sort of.

A wrong label is a small mistake; the message lands in the wrong folder, where search can still find it. An email sent in my name is a different animal entirely: it speaks to a human being on my behalf, and there is no undo. Those two mistakes don’t belong on the same ladder, so the new privilege shipped pre-shrunk. Mailroom drafts replies only for the rare messages it has already labeled as a real person wanting me specifically, with a reply expected. And the drafting path cannot send: it contains no send call; the only one in the codebase is still the permissions probe from the guardrail post, which ran once, at setup, to prove a point. One more invariant matters nearly as much to me: drafting never touches whether a message is read. Unread is how I know what I haven’t dealt with, and an agent that marks mail read while helping is erasing my to-do list. The draft path makes zero read-state calls, so that promise holds by construction rather than by good intentions.

The drafts come from the same local 27-billion-parameter model that classifies the mail, writing against a voice spec distilled from 366 replies I actually sent, segmented by relationship, because how I write to an old friend and how I write to a vendor barely share a language. The model produces text and nothing else; the recipient, the threading, and the headers are deterministic code, the same propose-don’t-execute split every agent here runs on. The finished draft sits in the Gmail thread looking exactly like one I’d started myself and wandered away from (a genre I am prolific in).

The part I’m proudest of is how it learns, because the feedback costs me nothing. Every draft opens a pair. When my real reply lands in the sent folder, an hourly sweep closes the pair, preserving the machine’s draft next to what I actually wrote, and deletes the unsent ghost. On Sunday, the weekly improvement run that already rewrites the classifier reads the week’s pairs and folds the differences into a capped “Learned from your edits” section of the voice spec: 12 bullets at most, fenced between markers only the improver may touch, with a guard that rejects any lesson quoting email text verbatim, so a stranger’s prose can’t write itself into my voice. The drafter reads the whole spec every time; a lesson learned Sunday shows up in Monday’s draft with no extra wiring. With labels I have to flag mistakes by hand. Here the correction is the reply I was going to write anyway. My sent mail is the red pen now, and I don’t even have to pick it up. I want to be precise about the boundary this widens: the Sunday run is a cloud call, and the pairs it studies include my own replies. My words are the one thing I’ve decided can leave the house, and even that has an expiry date; the distilling job is earmarked for a local model the day one is good enough to take it.

The first version of the learning loop broke before it ever saw a real pair. I had the new section stitched into the existing Sunday script by string surgery, and it mangled its own code on the first attempt (escape characters, the classic way). I reverted and rebuilt it as a standalone module, which the failure improved: every stage in this pipeline is now its own small, fail-soft piece, so a crashed voice lesson can’t take down the classifier rewrite riding the same Sunday schedule.

Qualifying messages are rare by design, so the firsts arrived on my correspondents’ schedule. The first draft was written on August 1. The first pair to close opened on the 4th, for a contractor’s scheduling thread; my real reply closed it the next morning, and that Sunday the improver read the two side by side and wrote its first five lessons into the spec. They’re about me, not him: I close short and warm and skip signing my name, I answer the proposal on the table instead of restating everything, I state my own constraint first and then offer the workable alternative. The week the diff comes back empty is the week the ghostwriter has my voice, and I’ve already decided what happens then: nothing changes. The send button was never on offer; it stays mine.