The audit my AI coach never ranessay

An AI agent executes a bad plan as faithfully as a good one. It’s the failure mode the demos never show, and part of why Gartner forecasts that over 40% of agentic AI projects will be canceled by the end of 2027, over escalating costs, unclear value, and inadequate risk controls. My version of the problem was smaller and closer to home: nobody was checking whether the model was aimed at the right problem to begin with. I ran into it in the least likely place, my running coach, a program I wrote and had come to trust.

The last post in the Coach saga ended on a promise: adherence, the number every commercial training app guesses at, was about to get measured. Consider it measured. My adherence was excellent. And it was the wrong number to worry about, because for a month, every session came back green while the fitness metric I care most about drifted the wrong way.

That metric is VO2 max, the watch’s estimate of aerobic capacity, roughly, how big your engine is. Mine had climbed steadily through spring, which is most of why I trust the training system these posts keep referring to, and it impacts Garmin’s calculation of your “fitness age,” exactly the sort of vanity metric that motivates me. Unfortunately, that figure started sliding, gently and consistently, through weeks in which Coach and I executed the program essentially perfectly. Nothing in the daily loop flagged a problem, and that was the interesting part: nothing in the daily loop could. The coach checks each session against the program, but nobody was checking the program.

So I ran the audit my coach never runs: 30 days of what actually happened, side by side with what the program assumed. It read like a slow-motion own goal. The weekly “intense cardio” session, the one nominally responsible for pushing the engine, had delivered zero minutes above 90% of max heart rate in a month; its prescribed heart-rate band sat well below the intensity that actually builds aerobic capacity. The session my regimen file labeled “the VO2 max development session” was, physiologically, the easy one; the labels were inverted, and the plan believed its own labels (and I had learned to blindly trust the plan). Cardio had no progression lever at all (durations sat at the bottom of their ranges indefinitely), and my recovery-aware guardrails, the feature I was proudest of, had been trimming intensity on rough-recovery days with no floor beneath them. A month of well-meaning caps had shaved the stimulus to nothing. Even lifting had a version of the disease: “maintain” targets, once hit, hardened into ceilings.

None of these was a bug in the software. Every one was a bug in the program the software faithfully served.

Two loops at different speeds: the daily brief gathers, prescribes, and verifies inside the program, while a slower six-week review audits the program itself against outcomes.

The rebuild was content, not code. The intense day became a classic 4x4 interval session with a hard floor: reach at least 90% of max heart rate, from the first week, no easing in. Textbook stuff. The easy day went back to being honestly easy, heart rate unconstrained. The regimen gained an explicit progression policy (add reps first, then load; a stall flag when progress flatlines) and the recovery guardrails got the rule they always needed: reduce the dose, never zero it.

The fix that matters most, though, was a meta-fix. The daily brief coaches inside the program by design; asking it to also question the program is asking one loop to run at two speeds. So program review is now its own slower loop: every 6 weeks, on a calendar, an outcomes-versus-assumptions audit of the whole regimen, the same separation my work life would call session coaching versus periodization. Had that review existed, this month would have been caught in week two. “Every session was on plan” and “the plan is working” are different claims; I had been testing only the first, and the results (unfortunately) spoke for themselves.

The first rebuilt week is done, and the intervals are exactly as unpleasant as they’re supposed to be, which I’m choosing to read as a good sign. The next review is already on the calendar for six weeks out, and it has a standing question: what else does the daily loop believe that nobody’s checked? The coach still grades every session. Now something other than me gets to grade the coach.