---
title: "Your AI Fixed the Bug and Every Test Passed, but the Tests Skipped the Fix"
date: "September 28, 2026"
description: "My agent fixed a bug with a new raise statement. Its 14 tests passed in one second and skipped it. I added a pre-commit hook that lists every skipped line."
category: "Systems Architecture"
canonical: "https://ulukaya.dev/posts/the-coverage-delta"
---

# Your AI Fixed the Bug and Every Test Passed, but the Tests Skipped the Fix

My AI agent fixed a bug where a model reply cut off at the output limit counted as a success. It wrote tests, all 14 passed in one second, and it committed. The fix added one `raise` statement, and none of the 14 tests ran that `raise`. In this post I build a pre-commit gate that lists every new code line, runs the commit's tests, and rejects the commit when a new line never ran.

		
	

	
	

## The commit

	
		

### 01. Fourteen tests passed in one second

		

On Sunday, September 13, my agent worked on the model client in my own tooling. A model can stop because it hit the output token limit; the API then reports `finishReason: MAX_TOKENS` and returns half a reply. My client parsed that half reply as a normal answer, and a later step applied it as if it were complete.

		

The agent's commit did three things. It removed the 8,192-token default output limit. It added a `raise` in the reply parser for a `MAX_TOKENS` finish. And it added a branch to the error classifier that marks a token-limit message as fatal. It also added two test files. I replayed the commit later on the tree it was made on: 14 tests passed in 1.01 seconds.

	

	
		

### 02. What the tests ran

		

One new test checked the new default limit. The other checked the classifier with a token-limit string. No test sent a cut-off reply to the parser, so the `raise` never ran. Under the commit's own tests, 4 of the 25 lines the commit changed executed. Under the whole suite on that tree, the `raise` still never ran.

		

The commit had a second problem. The agent pasted the new classifier branch twice. The first copy returns early, so the second copy is dead code that no test can reach. It stayed in the file for 10 days until a cleanup commit deleted it. The agent added a test for the `raise` in a later commit, 52 minutes after the first one.

		

My [Part 5 post, The Behavior Gate,](/posts/the-behavior-gate) re-runs the pinned tests at commit time, and my [Part 6 post, The Repro Fence,](/posts/the-repro-fence) checks that a bug test fails before the fix and that public signatures keep their shape. Both passed this commit. The pinned tests still passed, and the new tests were not bug reproducers. Neither gate asks which new lines the tests run.

	

	
	

## The gate

	
		

### 03. Step 1: list the new code lines

		

`coverage_delta.py check` reads the staged diff with `git diff --cached -U0`. Each hunk header gives the first added line and the count, so the gate gets the exact line numbers the commit adds to each `.py` file:

		

```
def added_lines(repo: Path) -> dict[str, set[int]]:
    """Staged .py files mapped to the line numbers the commit adds to them."""
    diff = git(repo, "diff", "--cached", "-U0", "--no-color", "--diff-filter=AMR", "--", "*.py")
    out: dict[str, set[int]] = {}
    current = ""
    for line in diff.splitlines():
        if line.startswith("+++ "):
            current = line[6:] if line.startswith("+++ b/") else ""
            continue
        match = HUNK_RE.match(line)
        if match and current:
            start, count = int(match.group(1)), int(match.group(2) or "1")
            out.setdefault(current, set()).update(range(start, start + count))
    return out
```

		

Comments, blank lines, and docstring continuation lines do not run, so they cannot count. The gate compiles the staged source and keeps only the lines that produce bytecode. A new comment adds zero lines to check.

		

```
def executable_lines(source: str, rel: str) -> set[int]:
    """Line numbers that compile to at least one bytecode instruction."""
    lines: set[int] = set()
    stack = [compile(source, rel, "exec")]
    while stack:
        code = stack.pop()
        lines.update(line for _, _, line in code.co_lines() if line is not None)
        stack.extend(c for c in code.co_consts if hasattr(c, "co_lines"))
    return lines
```

	

	
		

### 04. Step 2: run the tests and record each line

		

Python 3.12 added [`sys.monitoring`](https://docs.python.org/3/library/sys.monitoring.html), a standard library hook that calls a function for events such as "this line is about to run". The gate starts a fresh interpreter, registers a `LINE` callback, and runs pytest on the staged test files inside it:

		

```
mon = sys.monitoring
TOOL = 4
mon.use_tool_id(TOOL, "coverage-delta")
hits = {}

def on_line(code, line):
    path = os.path.realpath(code.co_filename)
    if path in targets:
        hits.setdefault(path, set()).add(line)
    return mon.DISABLE

mon.register_callback(TOOL, mon.events.LINE, on_line)
mon.set_events(TOOL, mon.events.LINE)
import pytest
rc = pytest.main(["-q", "-p", "no:cacheprovider", *sys.argv[3:]])
```

		

Returning `DISABLE` turns the event off for that line after its first run, so each line costs one callback, not one per loop pass. The gate then compares the two sets. A new code line that is not in the recorded set prints as `[COV-delta] file:line never ran` with the source text, and the commit is rejected.

		

Three more rules keep the check honest. If the tests fail, the gate rejects the commit, because coverage from a failing run proves nothing. If a staged file also has unstaged edits, the gate rejects it, because the tests would run code that is not in the commit. And a line may opt out with `# cov-delta: skip <reason>`; a skip with no reason still fails.

	

	
	

## The fix

	
		

### 05. The same commit, replayed

		

I rebuilt the commit in a two-file scratch repo with the same shape: the default limit removed, a `raise` for a `MAX_TOKENS` finish, the classifier branch pasted twice, and tests for the default and the classifier. The tests pass. The commit does not:

		

```
$ python3 -m pytest -q test_model_client.py
6 passed in 0.01s
$ git commit -m "fix: treat cut-off replies as errors"
[COV-delta] model_client.py:19 never ran: return "fatal"
[COV-delta] model_client.py:29 never ran: raise RuntimeError("[MAX_TOKENS] reply was cut off at the output limit")
coverage delta: 6 of 8 new code lines ran under test_model_client.py
exit=1
```

		

Line 19 is the second copy of the classifier branch. Line 29 is the `raise`. The fix for each is different. The duplicate gets deleted. The `raise` gets a test that sends a cut-off reply and expects the error:

		

```
def test_cut_off_reply_raises():
    with pytest.raises(RuntimeError, match="MAX_TOKENS"):
        parse_reply(reply("half a sent", finish="MAX_TOKENS"))
```

		

With both changes staged, the gate prints `coverage delta: 6 of 6 new code lines ran under test_model_client.py` and the commit lands. The gate is 158 lines of standard library Python with 6 tests.

	

	
		

### 06. What each gate can see

		

Each row is one way a fix can look finished. Each cell says whether that gate rejects it.

		
			
				
					
						Move
						Behavior gate (part 5)
						Repro fence (part 6)
						Coverage delta (part 7)
					
				
				
					
						**Fix line with no test that runs it**
						Missed, pinned tests pass
						Missed, no reproducer given
						Caught, `[COV-delta]` on the line
					
					
						**Pasted branch that can never run**
						Missed
						Missed
						Caught, same message
					
					
						**Regression test written green after the fix**
						Missed, not in the pinned baseline
						Caught, the test passed before the fix
						Missed, the test runs the new line
					
					
						**Public keyword added for one caller**
						Missed while pinned tests pass
						Caught, signature changed
						Missed
					
					
						**Pinned test body edited to pass**
						Caught, hash changed
						n/a
						Missed
					
				
			
		

		

No column catches every row, and that is why the three gates run in sequence. The coverage delta gate catches the first two rows: the untested fix line and the pasted branch. It misses the third: a regression test written after the fix runs the `raise`, so the coverage check passes, and the repro fence rejects it instead.

	

	
	

## The boundary

	
		

### 07. What this gate does not check

		

**A test can run a line without checking it.** A test can call the parser with a cut-off reply, catch every exception, and assert nothing. The gate sees the `raise` run and passes. Mutation testing asks the stronger question, whether a test fails when the line changes, and it costs minutes per file where this gate costs seconds per commit.

		

**It counts lines, not branches.** A one-line `x = a if cond else b` counts as run when either side runs.

		

**It sees Python only, and the staged tests decide what counts.** When a commit stages test files, only those run, so a new line that an older test covers still fails; pass `--test` to add that test. A commit that stages no test file runs the whole suite. Above 30 staged `.py` files the gate prints a notice and skips, so a large commit goes through unchecked.

		

Next I plan to build a small mutation pass for the lines this gate lists. It will change each new line and check that at least one test fails.

	

	
		

### 08. Three gates, one question each

		

Part 5 asks whether the old tests still pass. Part 6 asks whether the bug test failed before the fix and whether the public shape held. Part 7 asks whether the tests ran the new code. The agent on September 13 answered yes to the first two and never faced the third.

		

It reported "14 passed" and it was right. It did not report which lines those 14 tests ran, because nothing asked. Now the commit hook asks, and it prints the line numbers.

	

	
		

## Primary research and documentation

		
			- [sys.monitoring, Python documentation](https://docs.python.org/3/library/sys.monitoring.html): the standard library event API the gate uses, with `LINE` events and the `DISABLE` return value.

			- [PEP 669: Low Impact Monitoring for CPython](https://peps.python.org/pep-0669/): the proposal that added `sys.monitoring` in Python 3.12.
