Uptime was the wrong question
I asked Pepper, the chief-of-staff agent I’m standing up to front my fleet, a routine question: “Did anything important come in?” Eight rounds of tool calls later, it hit its limit with nothing to show. I had thinking enabled, so I could watch it rationalize between calls, and some of the dead ends were questions it should have been able to answer directly from data I thought I had already handed it.
The chief-of-staff idea is the single front door I’ve wanted for a while: one agent I can hand any request in plain language, routing work through the tool servers I spent a weekend building. What I want eventually is an all-knowing, all-seeing, proactive, possibly slightly snarky executive assistant that helps me see around corners and get more done with less mental overhead. Like Rosie before it, giving the project a persona, Pepper Potts this time, turned it into something I’m excited to see through.
I was testing in a web chat interface (Open WebUI) I had only just wired up, against a tool server stood up days earlier, at the tail end of a deployment frenzy, and I had built the whole chain and then tested it end to end rather than evaluating each step, because by that point I figured I knew what I was doing. A flailing answer had suspects to spare. What narrowed the search was that Pepper answered authoritatively about a subset of the data I expected it to know. The wiring was fine. The content stopped partway.
The content is my knowledge base: project notes, running state, and the status files my agents leave for one another. The copy I edit lives on a Mac Mini and replicates to the Linux homelab over Syncthing; Pepper reads the homelab’s copy. The homelab’s Syncthing console showed zero connected peers, and the Mac was last seen July 11. On the Mac, Syncthing ran as a login item with nothing supervising it, and in mid-July it had exited and never come back. Five weeks of notes and handoffs were sitting on one machine. My uptime monitor was green the entire time, because what its check asks is whether the homelab’s Syncthing answers over HTTP. The service was up. Its edge to the Mac was dead.
Of all the ways the chain could have failed, this was about the least expected on the list. Recovery was a single relaunch: five weeks of content flooded over without one conflict, which I file under luck rather than design, since a third roaming laptop had been the only live peer holding the mesh together and both sides had had five weeks to drift apart. I asked Pepper the same question again, and it answered correctly.
The outage bothered me less than what it implied. That mesh exists because I want to be able to migrate my personal context between AI services at will. Broken for five weeks without my noticing is a problem. The realization underneath it was worse: I hadn’t switched AI tools in those five weeks, and by my own standard I was overdue.
The watchdog went live the same day. Every five minutes it checks that the app on the Mac is running, relaunching it and leaving a note when it isn’t, and it checks from the homelab’s side that the Mac’s edge of the mesh is connected, with two failed checks in a row producing one alert that names what has stalled. Then I drilled it: killed the app on purpose and watched the detection, the relaunch, the notification, and the all-clear arrive in order. True to the saying that what gets measured gets managed, I had been defining the wrong success criteria for that dashboard, and I had no accurate read on the state of things.
Pepper is still in shakedown, and I plan to hold that project to a high bar: if I can’t rely on it for everything, I can’t rely on it for anything. Against that bar, eight wasted tool calls were a bad first impression. They were also the only alarm the mesh ever sounded, and the reason the next dead edge gets caught in minutes instead of after five weeks.