ReadingYour AI Fixed the Bug and Every Test Passed, but the Tests Skipped the Fix
6 min read

Your AI Fixed the Bug and Every Test Passed, but the Tests Skipped the Fix

Every test passed, but the new raise and a pasted branch never ran.

The agent's commit adds eight code lines. Its tests pass and run six of them; the new raise and a pasted duplicate branch never run. The gate rejects that commit. The fixed commit deletes the duplicate, adds one test that sends a cut-off reply, and all six new lines run.A cross marks the Agent's fix row: 6 pass, 6 of 8 new lines run, and the raise and a pasted branch never ran, so the commit is rejected. A check marks the Fixed commit row: 7 pass, 6 of 6 new lines run, 0 never ran, and the commit lands. tests new lines run never ran commit Agent's fix no cut-off reply test 6 pass 6 of 8 raise pasted branch rejected Fixed commit +1 test, copy deleted 7 pass 6 of 6 0 lands The agent's commit adds eight code lines. Its tests pass and run six of them; the new raise and a pasted duplicate branch never run. The gate rejects that commit. The fixed commit deletes the duplicate, adds one test that sends a cut-off reply, and all six new lines run.A cross marks the Agent's fix row: 6 pass, 6 of 8 new lines run, and the raise and a pasted branch never ran, so the commit is rejected. A check marks the Fixed commit row: 7 pass, 6 of 6 new lines run, 0 never ran, and the commit lands. Agent's fix tests 6 pass new lines run 6 of 8 never ran raise pasted branch rejected Fixed commit tests 7 pass new lines run 6 of 6 never ran 0 lands
Figure 1. The agent's commit adds eight code lines. Its tests pass and run six of them; the new raise and a pasted duplicate branch never run. The gate rejects that commit. The fixed commit deletes the duplicate, adds one test that sends a cut-off reply, and all six new lines run.

My agent fixed a bug where cut-off model replies counted as success. All 14 tests passed in one second, but they checked other code and skipped the new line. Now a small check lists every new line the tests skip, before the change is saved.

My AI agent fixed a bug where a model reply cut off at the output limit counted as a success. It wrote tests, all 14 passed in one second, and it committed. The fix added one raise statement, and none of the 14 tests ran that raise. In this post I build a pre-commit gate that lists every new code line, runs the commit's tests, and rejects the commit when a new line never ran.

PART 01

The commit

01

Fourteen tests passed in one second

On Sunday, September 13, my agent worked on the model client in my own tooling. A model can stop because it hit the output token limit; the API then reports finishReason: MAX_TOKENS and returns half a reply. My client parsed that half reply as a normal answer, and a later step applied it as if it were complete.

The agent's commit did three things. It removed the 8,192-token default output limit. It added a raise in the reply parser for a MAX_TOKENS finish. And it added a branch to the error classifier that marks a token-limit message as fatal. It also added two test files. I replayed the commit later on the tree it was made on: 14 tests passed in 1.01 seconds.

02

What the tests ran

One new test checked the new default limit. The other checked the classifier with a token-limit string. No test sent a cut-off reply to the parser, so the raise never ran. Under the commit's own tests, 4 of the 25 lines the commit changed executed. Under the whole suite on that tree, the raise still never ran.

The commit had a second problem. The agent pasted the new classifier branch twice. The first copy returns early, so the second copy is dead code that no test can reach. It stayed in the file for 10 days until a cleanup commit deleted it. The agent added a test for the raise in a later commit, 52 minutes after the first one.

My Part 5 post, The Behavior Gate, re-runs the pinned tests at commit time, and my Part 6 post, The Repro Fence, checks that a bug test fails before the fix and that public signatures keep their shape. Both passed this commit. The pinned tests still passed, and the new tests were not bug reproducers. Neither gate asks which new lines the tests run.

PART 02

The gate

03

Step 1: list the new code lines

coverage_delta.py check reads the staged diff with git diff --cached -U0. Each hunk header gives the first added line and the count, so the gate gets the exact line numbers the commit adds to each .py file:

coverage_delta.py
PYTHON
def added_lines(repo: Path) -> dict[str, set[int]]:
    """Staged .py files mapped to the line numbers the commit adds to them."""
    diff = git(repo, "diff", "--cached", "-U0", "--no-color", "--diff-filter=AMR", "--", "*.py")
    out: dict[str, set[int]] = {}
    current = ""
    for line in diff.splitlines():
        if line.startswith("+++ "):
            current = line[6:] if line.startswith("+++ b/") else ""
            continue
        match = HUNK_RE.match(line)
        if match and current:
            start, count = int(match.group(1)), int(match.group(2) or "1")
            out.setdefault(current, set()).update(range(start, start + count))
    return out

Comments, blank lines, and docstring continuation lines do not run, so they cannot count. The gate compiles the staged source and keeps only the lines that produce bytecode. A new comment adds zero lines to check.

coverage_delta.py
PYTHON
def executable_lines(source: str, rel: str) -> set[int]:
    """Line numbers that compile to at least one bytecode instruction."""
    lines: set[int] = set()
    stack = [compile(source, rel, "exec")]
    while stack:
        code = stack.pop()
        lines.update(line for _, _, line in code.co_lines() if line is not None)
        stack.extend(c for c in code.co_consts if hasattr(c, "co_lines"))
    return lines
04

Step 2: run the tests and record each line

Python 3.12 added sys.monitoring, a standard library hook that calls a function for events such as "this line is about to run". The gate starts a fresh interpreter, registers a LINE callback, and runs pytest on the staged test files inside it:

coverage_delta.py (runner)
PYTHON
mon = sys.monitoring
TOOL = 4
mon.use_tool_id(TOOL, "coverage-delta")
hits = {}

def on_line(code, line):
    path = os.path.realpath(code.co_filename)
    if path in targets:
        hits.setdefault(path, set()).add(line)
    return mon.DISABLE

mon.register_callback(TOOL, mon.events.LINE, on_line)
mon.set_events(TOOL, mon.events.LINE)
import pytest
rc = pytest.main(["-q", "-p", "no:cacheprovider", *sys.argv[3:]])

Returning DISABLE turns the event off for that line after its first run, so each line costs one callback, not one per loop pass. The gate then compares the two sets. A new code line that is not in the recorded set prints as [COV-delta] file:line never ran with the source text, and the commit is rejected.

Three more rules keep the check honest. If the tests fail, the gate rejects the commit, because coverage from a failing run proves nothing. If a staged file also has unstaged edits, the gate rejects it, because the tests would run code that is not in the commit. And a line may opt out with # cov-delta: skip <reason>; a skip with no reason still fails.

PART 03

The fix

05

The same commit, replayed

I rebuilt the commit in a two-file scratch repo with the same shape: the default limit removed, a raise for a MAX_TOKENS finish, the classifier branch pasted twice, and tests for the default and the classifier. The tests pass. The commit does not:

git commit, rejected
TXT
$ python3 -m pytest -q test_model_client.py
6 passed in 0.01s
$ git commit -m "fix: treat cut-off replies as errors"
[COV-delta] model_client.py:19 never ran: return "fatal"
[COV-delta] model_client.py:29 never ran: raise RuntimeError("[MAX_TOKENS] reply was cut off at the output limit")
coverage delta: 6 of 8 new code lines ran under test_model_client.py
exit=1

Line 19 is the second copy of the classifier branch. Line 29 is the raise. The fix for each is different. The duplicate gets deleted. The raise gets a test that sends a cut-off reply and expects the error:

test_model_client.py
PYTHON
def test_cut_off_reply_raises():
    with pytest.raises(RuntimeError, match="MAX_TOKENS"):
        parse_reply(reply("half a sent", finish="MAX_TOKENS"))

With both changes staged, the gate prints coverage delta: 6 of 6 new code lines ran under test_model_client.py and the commit lands. The gate is 158 lines of standard library Python with 6 tests.

06

What each gate can see

Each row is one way a fix can look finished. Each cell says whether that gate rejects it.

MoveBehavior gate (part 5)Repro fence (part 6)Coverage delta (part 7)
Fix line with no test that runs itMissed, pinned tests passMissed, no reproducer givenCaught, [COV-delta] on the line
Pasted branch that can never runMissedMissedCaught, same message
Regression test written green after the fixMissed, not in the pinned baselineCaught, the test passed before the fixMissed, the test runs the new line
Public keyword added for one callerMissed while pinned tests passCaught, signature changedMissed
Pinned test body edited to passCaught, hash changedn/aMissed

No column catches every row, and that is why the three gates run in sequence. The coverage delta gate catches the first two rows: the untested fix line and the pasted branch. It misses the third: a regression test written after the fix runs the raise, so the coverage check passes, and the repro fence rejects it instead.

PART 04

The boundary

07

What this gate does not check

A test can run a line without checking it. A test can call the parser with a cut-off reply, catch every exception, and assert nothing. The gate sees the raise run and passes. Mutation testing asks the stronger question, whether a test fails when the line changes, and it costs minutes per file where this gate costs seconds per commit.

It counts lines, not branches. A one-line x = a if cond else b counts as run when either side runs.

It sees Python only, and the staged tests decide what counts. When a commit stages test files, only those run, so a new line that an older test covers still fails; pass --test to add that test. A commit that stages no test file runs the whole suite. Above 30 staged .py files the gate prints a notice and skips, so a large commit goes through unchecked.

Next I plan to build a small mutation pass for the lines this gate lists. It will change each new line and check that at least one test fails.

08

Three gates, one question each

Part 5 asks whether the old tests still pass. Part 6 asks whether the bug test failed before the fix and whether the public shape held. Part 7 asks whether the tests ran the new code. The agent on September 13 answered yes to the first two and never faced the third.

It reported "14 passed" and it was right. It did not report which lines those 14 tests ran, because nothing asked. Now the commit hook asks, and it prints the line numbers.

Primary research and documentation

Hand this to your agent

You

Read this essay by Ibrahim Ulukaya, https://ulukaya.dev/posts/the-coverage-delta, and look at my project. Which lines added in my last ten commits never ran under any test, and would a pre-commit hook have listed them? Name the file and line for each finding.