Series · 7 essays · 58 min read
Ghost in the Runtime
How AI agents go wrong in real work, from agreeing with everything to passing tests on broken code, and the gates that catch it.
Part 1
Why Your AI Agent Agrees With Everything: 10 Production Failure Modes
My review agent approved a regex, then reversed itself when I asked the opposite question. I map the 10 biases behind that and the runtime check for each one.
Part 2
Code Over Context: Why Written Agent Skills Break in Production
10 markdown skill files cost my agent 22,000 tokens per turn and broke on smaller models at turn 4. I distill written skills into deterministic code tools.
Part 3
Your #1 Arena Model Fails in Real Repositories: The Leaderboard Mirage
The #1 leaderboard model failed 34% of my edge cases. Leaderboards rank single turns and my agents run 30 to 60, so I put compiler gates in the commit loop.
Part 4
The Crutch vs. the Operating System: Why I Deleted 4,000 Lines of Agent Prompts
Coding agents at 72.8% on SWE-Bench drop to 25% on multi-file repos. I replaced 4,000 lines of markdown prompts with a 45-line AST gate in the commit hook.
Part 5
When the AST Hook Goes Green and the Test Still Fails: The Behavior Gate
An AST hook can be gamed. This second pre-commit hook locks each baseline test by hash, runs it against staged code, and rejects the commit on a failure.
Part 6
Your AI Says the Bug Is Fixed, but the Test Never Failed: Two Checks Before It Ships
My agent reshaped two public functions to suit one caller; five others sat outside the diff. Two git checks block that and a test that passes before the fix.
Part 7
Your AI Fixed the Bug and Every Test Passed, but the Tests Skipped the Fix
My agent fixed a bug with a new raise statement. Its 14 tests passed in one second and skipped it. I added a pre-commit hook that lists every skipped line.