Healthy in every snapshot, sick for days
The drive lights on my storage array had started blinking in unison. Not the scattered flicker of eight disks doing unrelated work; a synchronized pulse, every bay at once, like the array was breathing. Around the same time, my alert pipeline kept producing storage health reports that concluded benign, which should have been reassuring and wasn’t. Nothing I could point to was failing. What I had was the homeowner’s version of a noise in the car the mechanic can’t reproduce: a conviction that something was off, and no metric to pin it on.
So I did the disciplined thing instead of rebooting on a hunch: I asked the AI in my terminal to prove me right.
It couldn’t, and its work was honest. The pool wasn’t mid-scrub (the periodic full-disk integrity check that lights every drive for hours); it was nearly idle. The synchronized blinking was the pool’s heartbeat: ZFS, my filesystem, batches writes and flushes them on a fixed cadence, and on a bored array that tick is the only activity left. Every disk blinking together is what healthy looks like. The zombie processes I’d spotted at login (programs that have exited but whose bookkeeping was never collected) were gone on recheck. One drive running at the top of its temperature limit earned a watch item, and that was the entire haul. Verdict after verdict, across days: healthy, hold, don’t reboot. My theory had been disk trouble, and the model dismantled it correctly. The LEDs were innocent. My unease was not.
Saturday afternoon the login banner broke the tie. A Linux server greets each login with a short status summary, and mine claimed swap was 92% full with 25 zombies loitering; live checks minutes later read nearly zero on both. Then the host stopped accepting new SSH connections while my existing session hummed along, a machine still alive but no longer answering the door. I overruled the standing hold, installed the 4 pending system updates, and rebooted. Clean shutdown, clean start, all 63 containers back within minutes. The medicine worked before anyone could name the disease.
The diagnosis was waiting in the metrics history, the one observer in the room that could see backward in time. Swap is the slow parking lot on disk where Linux moves memory pages it hopes not to need soon, and mine had been pinned at exactly 8.0 of 8.0 GB for four straight days. An overnight photo backup on Tuesday had squeezed memory hard enough to push the cold pages of running services out to that lot, and Linux never proactively empties swap; those pages come back only when something touches them. So the lot stayed full, because full is a state a working host can hold for days. When a real burst of writes arrived Saturday (about 26 GB, with my update run on top), the memory allocator finally had nowhere left to breathe, and the first casualty was new logins. At any given instant, nothing had been wrong. That was precisely the problem.
There’s a name for what my checks missed. Brendan Gregg’s USE method, a standard checklist for exactly this kind of hunt, asks three questions of every resource: utilization, saturation, and errors. For days the model and I interrogated utilization and errors and never asked the third, and I want to be precise about it, because the easy version of the answer is wrong. Swap pinned at 100% is utilization, a capacity reading a working host can sit at for days. Saturation is the backlog underneath: the time processes actually spend stalled waiting on memory, which Linux exposes as pressure stall information, not as swap occupancy. That pressure number stayed near zero right up until Saturday’s write burst, which is exactly why every instantaneous check came back clean. The host wasn’t saturated for four days; it was running on zero headroom for four days, and one burst turned no-headroom into a stall. The model was taking snapshots; the story was in the trend.
Score it honestly and both of us got half. I was right for the wrong reason: the host really was one write burst from falling over, and the blinking lights I blamed were innocent bystanders. The model was wrong by the right method: every probe it ran was true, and a host running on no headroom looks idle in every instantaneous one. The alert I added watches headroom (swap above 90% for two hours pages me); a second one now watches the saturation signal it can’t see, sustained memory pressure itself, so the trend gets to page me and not just the snapshot. The interesting question is what made my half of the partnership worth anything. It wasn’t intuition falling from the sky. It was pattern recognition built from years of living with these systems, and it stayed sharp because of how I use the tools, not despite them.
A small preprint out of MIT’s Media Lab (54 people writing essays, EEGs on) found the AI-assisted group showed the weakest brain connectivity of any cohort, and in the first session 15 of its 18 members couldn’t quote a single sentence from the essay they had finished minutes earlier. A small study, on a narrow task, and it still names the failure mode I care about. The researchers called the pattern cognitive debt: every answer you accept without processing is a small loan against your own ability, and the balance compounds quietly. My fleet automates my email, my alerts, my training plan. The study is about essays, not about debugging a server, but it names the direction I don’t want to drift: the tools doing so much of the thinking that I stop understanding the systems I’m responsible for.
So my agents operate under a standing order, written into the instruction file every one of them reads before we work (the employee handbook post tours that file): explanations are required, not optional. If you give me a command, tell me what it does. If you write code, walk me through it. I rarely want something to just work; I want to understand it well enough to tweak it, debug it, and explain it to someone else later. That single instruction converts every fix into a small lesson, and the tutor charges nothing for repetition. A human mentor gets tired of the third “why”; the model answers the 2am follow-up question with the same patience as the first, at whatever depth I ask. The swap saga is what that compounding looks like in practice: I could overrule the model’s hold because months of walked-through fixes had taught me what a starved memory allocator feels like, and the model could hand me the USE method in return. Neither of us finds that bug alone. That’s the argument for expertise in the AI era as I’ve come to see it: the ceiling isn’t the best model, it’s the best model paired with someone who understands the system it’s pointed at.
The habits that keep it that way are small. Never run a command you couldn’t explain back. Say what you expect before you look, so being wrong costs you an update instead of nothing. And when a fix lands, ask what the general lesson was, then write it down; my decision log exists because future me is the person I’m most often teaching. I recently wrote about deleting scaffolding that props up the model, and this essay is that one’s mirror: in the fleet, the workarounds melt and the verification appreciates; in the operator, the answers evaporate and the understanding appreciates. AI, used this way, is not a crutch for the things I don’t know. It’s the fastest path I’ve found to knowing them.