The switch I hope never fires
There’s a version of the single-point-of-failure problem that people who run their family’s infrastructure mostly don’t say out loud. Every password, photo, and document in my family’s life routes through systems I operate alone. I’m in good health (my robot coach is adamant about keeping it that way), but “what if I’m suddenly not here” is a question infrastructure has to answer, the same as “what if the disk dies.” So I built Lighthouse, a dead man’s switch: a system that stays silent while I keep answering a weekly check-in, and acts only if I stop. When it acts, it delivers to my wife the instructions she’d need: how to access and manage anything and everything, which of the blinking boxes in the rack actually matter, and some final words from yours truly. Bleak, but real.
What makes this thing interesting to engineer is that both directions of failure are catastrophic, and they’re opposites. If it fires falsely, it delivers a package that opens with the digital equivalent of “if you’re reading this…” to my wife while I’m fine and unreachable on a plane. If it fails to fire, it has failed at the only job it exists to do. Every design choice below traces back to which of those two directions it protects.
The first choice: it cannot live on the homelab. I’ve argued before that the fire marshal shouldn’t stand in the burning building; this is that principle at its limit, a system that must specifically outlive its operator and its operator’s hardware. It runs instead on Cloudflare Workers, code hosted on someone else’s always-on infrastructure with no server of mine anywhere in the path, and it fits inside the free tier, so the Worker itself has no bill to lapse (the text-message and email legs still ride outside accounts, which is its own chain to keep alive). The payload stores pointers, never the keys themselves: it directs my wife to the password manager’s family recovery and to an emergency kit she can hold in her hands (printing it is still on my list). Leak the whole payload and you get a sappy letter and a map, not a key.
The rhythm is one Discord message a week with a single button; press it, done. The switch has also learned to read proof of life I was already emitting: a push to any of a hand-picked list of my code repositories counts as a pulse, so ordinary weeks reset the ladder without my thinking about it. The list is deliberately explicit rather than “all my repositories,” because the two mistakes aren’t symmetric: a list that’s too short sends me a needless ping, and a list that’s too broad could one day let some automated commit keep a dead man looking alive. Ignore every signal, and a ladder climbs over roughly two weeks: daily nudges, then text messages to my phone, then a check-in to my wife (“checking on Jesse, reply PAUSE if he’s fine”), and only after all of that, delivery. Any response at any rung before delivery resets everything. I walked the failure cases on purpose: a two-week phone-free vacation would climb as far as her check-in rung, and I’ve decided that’s correct, because she’d be on the trip. (There’s now a snooze command for planned absences; the ladder’s behavior when I forget to use it is the part I had to decide.) The ladder itself is a deliberately boring state machine: pure functions, 104 tests, no cleverness anywhere I could avoid it. That ban includes AI. I built and validated every piece of this with a model as my peer programmer, and no model appears anywhere in the execution path.

One design question sounds like trivia and decides everything: when the system acts, does it send the message first or save its record first? The answer depends on which catastrophe a crash between those two steps would cause. An escalation sends first, because the dangerous version is a warning that got saved but never delivered: the ladder would keep climbing on warnings it only believes people received, with no human actually in the loop. An acknowledgment saves first, because there the dangerous loss runs the other way: losing my “I’m alive” is a step toward a false fire, while a button that looks unpressed after I pressed it is merely cosmetic. Duplicate messages are the accepted cost of both orderings. And the ladder climbs at most one rung per tick of its clock, so if the clock itself fails for a day, the system stalls in place rather than catching up in a burst. Stalled is the failure mode I chose everywhere it was an option. Stalling has its own failure, though: a switch that quietly stops ticking has failed its one job as surely as one that fires wrongly. So the switch reports every heartbeat to an outside watchdog on infrastructure I don’t own and can’t take down along with me, and if the daily tick ever goes silent, that watchdog is what notices.

My (im)patience won’t allow me to rehearse a two-week ladder at real speed, so staging runs 1,440 times faster: a day per minute. It’s a separate Worker with its own Discord staging app and its own keys, so the accelerated clock cannot reach the real channel (my panicked wife): a cross-wired deploy fails at signature verification instead of rehearsing on my actual family. The full walk from alive to fired ran in 17 minutes and 18 seconds with every transition on schedule, and the datastore’s eventual consistency showed itself at that speed: a button press can take up to a minute to become visible, so its effect lands a tick late. For most of the ladder that is harmless, the acknowledgment arriving delayed but never lost, the equivalent of a late “He’s fine!” text. There was exactly one rung where “delayed but never lost” quietly broke: the last one. A press landing in the minute before the final tick could be invisible to that tick, and because fired is terminal, the switch would never read the acknowledgment that surfaced a beat later. The fix is a grace step ahead of the end: exhausting the ladder now arms a pending state that sends nothing and waits one extra tick, so a late acknowledgment still has a window to fold everything back to alive before a single message goes out. The best test was an accident. I staged a rung, got pulled away, and came back to a switch that had fired in my absence; my unplanned funeral became the proof that FIRED is terminal, and the system never talks itself back down from it. I’ve since recovered.

It’s armed now. The text-message legs survived the carrier’s verification gauntlet (a saga of its own; the messaging industry’s anti-spam machinery is not built for a two-person use case) and passed a live drill: real texts, real replies, a stranger’s number silently ignored. The full fire drill ran too, rerouted to my own inbox instead of hers: the payload email, the check-your-email text, and the day-later repeat (a minute, at staging speed), all delivered, all readable. We walked the runbook together, and she texted her consent to the switch from her own phone. It has asked me every week since whether I’m still here. What remains is the part no amount of engineering covers: writing the real letter, and printing the emergency kit.