← All writing

Once a quarter, I lie to my monitoring

On purpose, and on a schedule.

My homelab’s alert pipeline had a quiet month, and that worried me more than a loud one would have. Over a few weeks I had excluded a dozen-plus classes of known-benign noise from it, and every exclusion is a bet that the pattern really is benign forever. A quieter channel and a deafer channel look identical from the outside; the dashboard is green either way. There’s a second bet underneath that one: I move between AI models constantly, and every capability jump the new ones bring widens everyone’s attack surface, mine included. Probing my own network before somebody else does struck me as bare-minimum due diligence.

The industry’s answer is chaos engineering, deliberately injecting failures to prove your systems handle them, and there’s a homelab-sized version of it: fake seven faults, each one labeled as synthetic so nothing downstream mistakes it for real trouble, all fired at a sacrificial victim container rather than anything the family touches. The pre-drill ritual is due diligence rather than damage control, confirming more than once that the drills are drills and not live rounds, because explaining to my wife or kid that the network is down because daddy was trying to be a security boy scout is unacceptable, whatever years of uptime precede it. Then the drill asks its one question. Does the pipeline notice, every time?

The first layer said yes, and I’ll take the win: all 7 probes (a failing disk, a power event, an unfamiliar login, a burst of application errors, and friends) fired their alerts at the detection layer. A month of aggressive noise-tuning had cost exactly zero recall there. If the story ended here, this post would be a screenshot and a pat on my own back.

Deeper in the pipeline is where it got interesting. When an alert fires, an AI investigator SSHes in, works the problem, and posts a verdict to Discord. Seven probes had become six alerts by this floor, and I went back to the logs to make sure that arithmetic was innocent: it was. The two disk-fault probes tripped the same detection rule, which folded them into a single alert, an aggregation doing its job, not a message going missing. Of the six alerts that flowed on, four came back as verdicts, each correctly recognized as a labeled synthetic and closed as benign. Two produced nothing. No verdict, no error, no record that the alerts had ever existed. And my first thought wasn’t a pipeline bug; it was the worse branch. I had built the injector with an AI, confirmed multiple times that its probes were benign and scoped to the victim container, and put real faith in those confirmations. But models hallucinate, and some have been known to act without asking, so for a bad few minutes the working theory was that I was attacking myself, and that whatever it was might still be executing. The probes re-checked as inert, which left only the boring explanation: the alerts had existed, and something of mine had eaten them. That post about the investigator laid down the rule I run this pipeline by: “An alert is allowed to arrive degraded. It is not allowed to vanish.” Two alerts had vanished, and without a drill built from known inputs, I could not have known; you cannot notice the absence of a message about an event you don’t know happened.

The culprit was a safeguard. The triage layer has a loop-prevention guard: an investigation’s own SSH commands generate log lines, and without protection the pipeline would investigate its own investigations forever. But the guard was written one notch too broad; while any investigation of a machine was running, it skipped every other alert from that same machine, unrelated ones included, and a skipped alert produced no message at all. Each investigation runs 20-30 seconds, so the window sounds harmless, until you describe the real-world shape of it: a genuine multi-system incident is exactly when several distinct alarms fire from one machine at once.

The fix went out the same day, and it was a narrowing, not a removal: the guard now skips only a true duplicate (the same fault type from the same machine while its investigation is in flight), and distinct faults investigate in parallel. I re-ran the failure directly, two different synthetic faults from one machine, and watched them complete side by side. Then the re-validation earned its keep twice, because it surfaced a third layer of deduplication nobody had mapped: the alerting platform itself mutes each rule for a full hour after it notifies. Fine for an ongoing-condition alarm, where the mute is really a reminder rate; badly wrong for discrete faults, where a second disk event arriving 40 minutes after the first would have been silenced by a timer that never read either message. So I split the difference by kind: the ongoing-condition alarms keep the hour on purpose, and the discrete-fault classes now re-arm in 10 minutes. That top-layer mute still exists; it just no longer makes hour-long, content-blind silencing decisions for the faults where a second event is its own emergency. Deduplication belongs downstream, where something with context can decide.

The scoring rule I hold this drill to is that a pass is a screen, not a green light. This run passed at the layer I was worried about and still lost two alerts one layer down, in a safeguard I wrote myself and trusted for months. The injector also can’t fake the smart-home system’s alert class yet, so that class runs on faith. That’s the next probe I’m building. For now the drills are quarterly; the next few runs will tell me whether that’s often enough.