Six seconds of microphone memory
The assistant woke up in the middle of my wife’s work call, again, and started answering her. She crossed the room to shut it up by hand while it kept opening the microphone. There are three of these little satellites in my house, running firmware I compile myself and a wake word I trained, answering from models on my own GPU with no trip to anybody’s cloud. That morning it fired three times in 11 minutes, already turned down to its least sensitive setting. However good the engineering is, my homelab exists to serve my family, and an experiment that isn’t serving them gets refined or retired. My options were rip and replace or fix, and for something this close to working, the decision to iterate was obvious.
A week earlier I had declined this exact fix. The wake word is a small model that runs on the device itself, and training it had already taken hours across thousands of simulated variations, so I didn’t believe the word was the problem. The false triggers didn’t sound anything like it, and I had never captured a single one. We happened to work from home a couple of extra days before a vacation, the false wakes stacked up in that window, and my wife had had enough of it and told me to do something about it. Time to get to work.
The trouble was that I had no evidence, and the pipeline was built so I never could. I worked on the Google Assistant team for a stretch, so I knew the shape of this already: hotword detection runs locally on the device, listening the whole time, and the device only starts streaming a query once the word has fired. Every record in my voice logs is therefore a transcript of what came after the trigger, which is exactly the audio that isn’t interesting for troubleshooting. The sound that fools the model is the second or so before it, and that audio exists nowhere at all: it sits in the device’s memory for a moment and then it’s gone.
So the fix had to live inside the firmware. I wrote a small ESPHome component called wake_capture that keeps a six-second ring of microphone audio in the device’s spare memory, continuously overwriting itself, converted to the same mono format and gain the wake-word model hears, so a clip approximates what the model heard rather than what a human would have. When a wake event fires, the component arms a 500-millisecond post-roll, snapshots the ring, and posts it as a WAV file from a low-priority background task so the assistant’s own response never stutters. Nothing is ever written to the device’s flash, and a software mute stops the capture along with everything else.

The clips land on a purpose-built collector I developed: a LAN-only service, standard library, small enough to read in one sitting, writing each WAV into the folder that feeds round two of the training. Retention is capped in the service itself at 500 files or 100 megabytes, whichever comes first, oldest pruned. I am the only person with access to anywhere those clips land, they never leave my network or get committed to a repo, and they are deleted the moment a training round consumes them. She knows the clips exist and what they are for. Six seconds of audio from a work-from-home day is rarely sensitive enough to be worth keeping, and it is the only material that can fix the thing that recorded it. When this same pipeline captured work-call fragments back in July, I read the whole corpus, purged the work content, and kept the family commands.
The office unit got the firmware first, on its own, and the other two followed two days later once it had proved it wouldn’t wedge anything. (Getting the flash onto the first unit cost me an evening of Mac privacy-permission debugging that had nothing to do with audio.)
The negatives are piling up now, a few a day, which is the first time this problem has produced evidence instead of complaints, and round two of the training waits on that pile. The harder question is the one I haven’t solved: anyone who knows these units exist and knows the word can give them a command, and my toddler shouldn’t have the same system-level controls or receive the same verbose responses I do. Telling speakers apart by voice alone, or by voice plus a presence sensor, or even voice plus a camera I already run, is still undecided, and the clock on it is my toddler’s vocabulary.