← All writing

Second opinion, same blind spot

The third round of an external AI code review of my repos came back with 12 findings. All 12 were real, and all 12 were fixed the same day. Twelve for twelve, and this after the two previous rounds found just as many gaps. The bug that mattered most that day was not on the list, and neither model in the loop caught it.

Most of what runs in my homelab is drafted and built by a flagship frontier model, with me directing. A review from the model that wrote the code is like a student grading her own exam, so I wanted a second opinion that owed the first one nothing: a second flagship model, from a different company, handed the code cold along with my decision log, to see whether it reached the same conclusions. I have a standing rule that anything I catch myself repeating gets automated. Here that meant an MCP connection between the two tools and a written procedure for the loop, so the models interact in a structured, predictable way and tailor their output to each other instead of to me running markdown files between the two. I stay in the loop as the human who approves the final decision.

All three rounds ran across two days in August. The first round produced 14 findings and not one invented file path, which anyone who has asked a language model to review a codebase knows is worth saying out loud. It also inflated 7 of 11 severity ratings, grading a one-operator homelab like an enterprise. That calibration went into the next round’s prompt and held. The reviewer remarked more than once on how impressed it was by the state of what it inherited, then kept finding real bugs in it anyway; the two models worked alongside one another better than I could have hoped for.

The bug came out of one of round two’s fixes. Columbo, the investigator that chases my alerts, gets a shell through a restricted account, and a deny filter is there to keep that shell read-only. The review found the filter too permissive, and the fix tightened it, including a rule against shell redirects, the > that overwrites files. Written as a pattern, that rule also matches 2>&1, the harmless tail that folds error output into the transcript, and some form of it sits on the end of a large share of the diagnostic commands a Linux investigator runs. The morning the tightened filter went live, it refused two real investigations.

Round three had been told, explicitly, to hunt for regressions in the earlier rounds’ fixes. It listed ten candidates, all of which checked out; this was not among them. For an adjacent finding the reviewer had even dry-run the exact expression in question, and the handful of synthetic cases passed. What caught it was history. The investigator records every command and outcome in a database, so the builder replayed the new filter against 30 days of that record: 429 of 753 real commands would have been refused. The whole check is a 20-line script, and it settled in seconds what two frontier models had both signed off on.

The filter was rewritten that afternoon to parse commands instead of pattern-matching them, denying by command word and allowing error output only two places to go: folded into the transcript, or dropped entirely. Replayed against the recorded history again, now 763 commands, the rewrite blocks 3, each one on purpose.

The review loop: I hand the code cold to two frontier models from two companies, the builder and the reviewer, connected over MCP; three rounds go twelve for twelve and round three dry-runs the redirect rule on synthetic cases and passes it; the replay runs that rule against the investigator’s own 30-day command record and finds 429 of 753 real commands would have been refused; the rewrite faces the same record and blocks 3 of 763, each on purpose, then ships under the new rule that no filter change deploys until it has faced the agent’s history.

I still wanted the injection fixes shipped promptly. The severities were real, but narrow, most of them requiring access to my home network before they mattered, and moving at the speed of AI meant the whole list could close in a single sitting. What the round lacked was a better check before deployment. In 1986, Knight and Leveson tested the assumption that independently written programs fail independently and found the failures correlate: different authors, same blind spots. Two models from two companies are still two readers looking at the same handful of synthetic test cases. The check neither of them had exercised was the recorded past. So the rule I promoted out of the round is mechanical: no filter change on an agent’s tool path ships until it has been replayed against that agent’s own command history.

I see the memes about how being a technologist now just means hitting “allow” whenever an AI asks. I can’t really relate. I do run with most permission prompts off, “dangerously” so in the tooling’s own vocabulary, and the trade I hold myself to is understanding what is happening and why, well enough to catch the mistakes and steer. The scorecard was never mine to claim either way: twelve for twelve was two frontier models negotiating with each other. What I can take credit for is thinking to connect them, standing up the plumbing between them, and the structure that keeps them working to each other’s standards while I approve the final call.

The loop is on the calendar quarterly now, with the replay written into its procedure next to the calibration notes from round one. The rewritten filter shipped the same day as everything else, at the same speed. The difference is that before it deployed, it had to answer to 763 commands of the investigator’s own history, which is 763 more than the twelve-for-twelve version ever faced.