Capital & Compute

Claude Code TDD: How Failing Tests Steer the Agent

How failing tests steer Claude Code in a startup monorepo: tests you watch fail first, a policy on which to write, and checks that catch fake passes.

· ai· coding-agents· tools· claude code· testing· By Capital & Compute
Five-step red-green testing loop with a side branch that catches tests passing before any code exists.

I am a full-stack developer at a startup, and Claude Code writes most of the code I ship. There is far too much of it for me to read line by line, so for a while a green test run was my signal that a change was safe. Then this happened.

The real case was messier, so here is the same trap in its simplest form.

A list page shows 10 items per page, and a small function works out how many pages there are:

export function pageCount(total: number, perPage: number): number {
  return Math.floor(total / perPage);
}

With 100 items that gives 10 pages, which is right. With 101 items it also gives 10 pages, so the last item never appears. There is no error and no warning. It is simply missing.

This is the test that sat next to that function, green the whole time:

it('counts the pages', () => {
  expect(pageCount(100, 10)).toBe(10);
  expect(pageCount(101, 10)).toBe(10); // copied from what the function returned
});

Read the second line slowly. It was written by recording what the code did, not what the code should do. So it never caught the bug. It protected it.

The fix is one word: round up instead of down. The more important change is the test, which now demands the right answer:

export function pageCount(total: number, perPage: number): number {
  return Math.ceil(total / perPage);
}

it('gives leftover items their own page', () => {
  expect(pageCount(101, 10)).toBe(11); // failed before the fix: got 10
  expect(pageCount(100, 10)).toBe(10); // an exact fit gets no empty page
  expect(pageCount(0, 10)).toBe(0);
});

Run the new test against the old function and it fails. That failure is the proof that the test checks the right thing.

That is the habit this whole post is about. I no longer try to read every line Claude writes. I read the tests instead, starting with the one that fails. A test that fails before the change and passes after it is the smallest thing a person can check that still tells you whether the feature works. A test that never failed tells you nothing.

This post is the testing half of the harness described in the Claude Code harness example. All code below is a generic reconstruction of the pattern with placeholder names, not the production files.

Why I do not just tell Claude to “do TDD”

A test like that one is the reason the repo has no “always write tests first” rule. If you tell Claude Code to do TDD and let it write both the test and the code, you mostly get tests that agree with the code, whatever the code does. That test was exactly this: a test and a bug that agreed with each other.

I am not the only one who has seen this. Birgitta Böckeler tested it properly in an August 2026 article on Martin Fowler’s site, TDD inside the agent loop - theater or actual value?. She compared agent runs with and without a TDD workflow, had a model judge the results, and found “no clearly discernable difference” in quality. She also caught tests that “checked the implementation’s output against itself,” and she has stopped asking agents to write tests first.

Where I landed is a little different. Writing the test first still pays off, but only when two things are true: I decide what the test checks, and I see it fail before the fix. Claude can type the test. It should not be the one deciding what “correct” means, and it should not grade its own work.

Anthropic’s best-practices documentation backs up that second point. Its example bug-fix prompt says to “write a failing test that reproduces the issue, then fix it,” and it suggests you “have one Claude write tests, then another write code to pass them.”

The loop: red first, and watch it fail

The core loop is short. The figure shows it, including the branch that catches most bad agent tests.

Red-green loop with a fake-pass checkFive steps: write the failing test, run it and require red, let the agent implement, run it and require green, then label the test and add a check if the bug can happen again. If the test is green before any code is written, it is a fake pass and goes back to step one.Next bugWrite the failing testYou write it, or approve exactly what it assertsRun it: must be redRecord how it fails, in a commentAgent implementsThe smallest change that makes it passRun it: must be greenThe workspace tests plus the guard it touchesLabel it, then guard itSay what broke; add a check if it can happen againGreen already?Fake passIt checks nothing. Rewrite it.
Red-green loop with a fake-pass check
StepWhat happensExpected result
1. Write the failing testYou write it, or approve exactly what it assertsNot a test run
2. Run it: must be redRecord how it fails, in a commentFails
3. Agent implementsThe smallest change that makes it passNot a test run
4. Run it: must be greenThe workspace tests plus the guard it touchesPasses
5. Label it, then guard itSay what broke; add a check if it can happen againNot a test run
Branch: Fake passGreen already? It checks nothing. Rewrite it.Return to step 1
The test-first loop used for bug fixes and behavior changes. The branch on the right catches most bad agent tests: if a new test passes before any code is written, it checks nothing.Source: Author's own repository practice (first-party), September 2026

The best example in the repo came from a performance bug. A settings screen re-rendered its entire grid on every checkbox click. The fix was a memoization change, and the regression test counts renders. The explanation doc for that fix states the rule better than I can: the test was written against the broken screen and watched fail, “so the green run certifies the fix rather than the test’s own assumptions.” The failing counts from the red run are recorded in the test’s comment, so a reviewer can see what red looked like without re-running history.

A generic version of the shape:

// apps/web/src/settings/grid.rerender.test.tsx
// Regression: toggling one cell re-rendered every row (120 renders for 1 click).
// Written red-first against the broken grid: before the fix this failed with
// "expected 1 render per row, got 120". Removing the memo in either file turns it red again.
import { render, screen } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { expect, it } from 'vitest';

import { countRenders } from '../test/render-probe';
import { SettingsGrid } from './settings-grid';

it('re-renders only the row whose cell changed', async () => {
  const renders = countRenders(SettingsGrid.Row);
  render(<SettingsGrid rows={makeRows(40)} />);
  renders.reset();

  await userEvent.click(screen.getAllByRole('checkbox')[3]);

  expect(renders.total()).toBe(1);
});

Three details carry the weight. The comment names what broke, which is what stops a later cleanup from deleting the test as redundant. The comment records the red result. And the last line of the comment says how to turn it red again, which is a ten-second way for anyone to prove the test really works.

The prompt that produces this is deliberately split into two turns:

Turn 1: Write a failing test that reproduces <the bug>. Run it. Show me the failure
        output and stop. Do not change any non-test file.
Turn 2: (after I have read the test and the red output) Make it pass with the smallest
        change. Run the workspace tests. Do not edit the test.

The pause between the turns is the entire point. Reading one test and one failure message takes a minute. Reading the 300-line diff that follows is optional once the test is right.

The failing test is the review gate

The same idea scales to security-sensitive code. The web app’s Content Security Policy is built by a function, and each branch of it is locked by a test that asserts the exact policy string. The docs for that area put it plainly: the failing test is the review gate, because the pull request shows the policy change twice, once in prose (the test) and once in code (the builder).

it('allows the analytics origin on marketing pages only', () => {
  expect(buildCsp({ route: 'marketing' })['connect-src']).toBe(
    "'self' https://analytics.example.com",
  );
  expect(buildCsp({ route: 'app' })['connect-src']).toBe("'self'");
});

When the agent changes the policy, this test goes red, and updating it is a visible, reviewable decision rather than a side effect buried in a helper. A reviewer who reads nothing else reads the test diff.

Decide which tests to write, in writing

The opposite failure is just as common. A recent r/ClaudeCode thread on keeping up with agent-written code included a 40-line change that came back with four new test files and 40 tests. That problem is real, and it is not solved by telling the agent “fewer tests.” It is solved by a written policy the agent can read, kept in the docs site and pointed to from the app’s CLAUDE.md (“Does this need a test? Read the testing strategy first”). A condensed, generic version:

## Tier A: always

- A regression test for every bug that actually escaped. Annotate it with what broke.
- Behavior a user or an API client depends on, asserted through observable outcomes.

## Tier B: only with a stated reason (write the reason in the test)

- Render-count or timing probes, written red-first against the bug they guard.

## Tier C: do not write

- A test asserting what the linter or the type checker already rejects.
- Tests of vendored or framework code.
- Snapshots with no semantic content.
- "A mocked function was called" when an observable outcome already proves it.

## Tier D: cost rules

- One timeout in config. A per-test timeout override is a smell, not a fix.
- A test must not measure the machine: stop the clock, do not widen the margin.

Tier D exists because of a specific audit. One app had accumulated 80 per-test timeout overrides across nine files, while the other apps had one and zero. None of those tests were testing the wrong thing; they were just paying ten to a hundred times the necessary runtime. The trim that followed cut suite runtime by about two thirds while removing only 4.6% of the tests. The shared config now carries a single line and a comment:

// vitest.config.ts
export default defineConfig({
  test: {
    testTimeout: 10_000, // If a test needs more than this, the test is the bug.
  },
});

The same audit produced a finding worth repeating: deleting tests to make CI faster attacks the smaller half of the cost. Before concluding a file is expensive, run it alone.

Fake passes: the agent’s favourite mistake

The most dangerous test is one that passes for the wrong reason, because it looks exactly like one that passes for the right reason. Agents write them constantly, not out of malice but because a green run is what they were asked for. Four patterns in the repo each exist because one of these slipped through.

A test that checks an empty list always passes. Many tests in the repo first find the files they check (every controller, every page, every component) and then check each one. If a refactor moves those files, the search finds nothing, and every check passes because there was nothing to check. So each of these tests starts by making sure it found something:

const controllers = globSync('src/**/*.controller.ts');

it('discovers the controllers it is about to check', () => {
  // Without this, a moved folder makes the scan return nothing and every
  // check below passes without checking anything.
  expect(controllers.length).toBeGreaterThanOrEqual(8);
});

it.each(controllers)('%s records an audit entry on every mutation', (file) => {
  const source = stripComments(readFileSync(file, 'utf8'));
  expect(source).toMatch(/this\.audit\.record\(/);
});

An empty function still gets called. Checking that a function is called proves nothing if someone has emptied out its body. The audit test also checks that the recorder’s body still performs the insert. The stripComments call is there for a very specific reason: the most likely fake pass is a debugging leftover like // await this.audit.record(...), which a plain text match would happily accept.

The check looks at the wrong thing. A pagination control marks its end state with aria-disabled, not the disabled attribute. toBeEnabled() ignores aria-disabled, so it passes on a dead-end button. The app’s CLAUDE.md now says to assert toHaveAttribute('aria-disabled', 'true') for that component, never toBeEnabled().

The test runs in the wrong place. Under a Node test environment, code guarded by typeof window !== 'undefined' silently takes the server branch, so every assertion about a signed-in browser session passes without ever running the browser path. That is why the test environment is chosen by file extension, not by habit (more on that below).

Test the guards, not just the code

A harness built on checks has a blind spot: a check that stops checking looks exactly like a codebase with nothing wrong. A lint rule that is enabled but unreachable (wrong glob, wrong plugin load order) exits 0, and exit 0 is indistinguishable from a clean run. In the repo, one React rule sat switched on but silently doing nothing for three weeks.

The fix is a probe script. It writes a small file that must trigger each blocking rule, lints it, and fails unless every rule fires at error severity:

// scripts/check-lint-rules-fire.mjs (generic version)
import { spawnSync } from 'node:child_process';
import { rmSync, writeFileSync } from 'node:fs';

const PROBE = 'apps/web/src/.lint-probe.tsx'; // gitignored; not a *.test.tsx, which is exempt
const MUST_FIRE = ['no-array-index-key', 'rules-of-hooks']; // oxlint reports codes as plugin(rule)

writeFileSync(PROBE, probeSourceThatBreaksEveryRule());
try {
  // oxlint exits non-zero when it finds errors, which is the expected case here.
  const { stdout } = spawnSync('pnpm', ['exec', 'oxlint', '--no-ignore', '--format=json', PROBE], {
    encoding: 'utf8',
  });
  const found = JSON.parse(stdout).diagnostics;
  const inert = MUST_FIRE.filter(
    (rule) => !found.some((d) => d.code.includes(`(${rule})`) && d.severity === 'error'),
  );
  if (inert.length) {
    console.error(`INERT under apps/web/src/**: ${inert.join(', ')}`);
    process.exit(1);
  }
} finally {
  rmSync(PROBE, { force: true });
}

Severity is checked, not just presence, because a rule demoted to a warning also enforces nothing in CI. The probe lives outside any exempt path, because a probe inside a test file would prove nothing.

The same principle applies to bespoke guards. A permissions drift checker in the repo runs a built-in self-test on every invocation: it feeds the checker fixtures that break each rule and fails if any rule stays quiet. Its header says the acceptance bar is that breaking each rule is shown to fail it, which “is what stops a refactor from quietly turning a rule into return [].” The hook self-tests described in the harness example follow the same rule, with near-miss fixtures that must not fire.

One more small guard is worth copying. The rule “every UI component has a story” started with a list of existing components that had no story yet. That list is shrink-only: the check fails if a listed component has since gained a story (so the entry must be removed) and never accepts a new name. Adding a name is never the fix for a new component; writing the story is.

Steer how the agent runs tests

Most of the steering is a handful of lines in CLAUDE.md and in the vitest configs.

Verify what you touched, not everything. The root CLAUDE.md tells the agent never to run the full verify suite locally, because that is the CI gate. It runs the workspace test and typecheck for what it changed, plus the one guard the change is about. The PR skills repeat the same instruction, so the agent does not burn ten minutes of runner time on every small edit.

The file extension decides the environment. Each frontend’s vitest config defines two projects, and the include globs enforce the split:

// apps/web/vitest.config.ts (generic)
export default defineConfig({
  test: {
    projects: [
      { test: { name: 'node', environment: 'node', include: ['src/**/*.test.ts'] } },
      {
        test: {
          name: 'dom',
          environment: 'jsdom',
          include: ['src/**/*.test.tsx'],
          setupFiles: ['./src/test/setup.ts'],
        },
      },
    ],
  },
});

A new test’s extension decides where it runs, so the agent cannot accidentally test a browser path in Node. Vitest documents this as test projects, and the config comment in the repo says it best: the file extension is the honest signal, enforced by the globs rather than by convention.

Mock at the boundary, never the thing under test. API end-to-end tests build the real application module in-process and swap services through dependency injection overrides, never vi.mock, “so the real guard chain runs.” Authentication gets one shared stub builder, because the shape of the stub is what kept drifting when every test hand-rolled its own.

A Stop hook can make the check deterministic. The hooks reference documents that a Stop hook can block the turn from ending until a check passes, and that Claude Code applies an eight-consecutive-continuation cap, reporting stop_hook_active: true on the input when it is already continuing because of a stop hook. A minimal version runs only the tests for workspaces the session changed:

#!/usr/bin/env bash
# .claude/hooks/test-before-stop.sh, registered on Stop
input=$(cat)
[ "$(printf '%s' "$input" | jq -r '.stop_hook_active')" = "true" ] && exit 0

changed=$(git diff --name-only HEAD | grep -oE '^(apps|packages)/[^/]+' | sort -u)
[ -z "$changed" ] && exit 0

filters=$(printf -- '--filter=./%s ' $changed)
if ! out=$(pnpm turbo run test $filters 2>&1); then
  printf 'Tests fail in the workspaces you changed. Fix them before finishing:\n%s\n' \
    "$(printf '%s' "$out" | tail -40)" >&2
  exit 2
fi

The broader verification-loop setup (prompt-level checks, /goal, a Stop hook) is covered in the agentic workflows playbook. The point here is narrower: scope the hook to what changed, or the agent learns to wait out a slow suite instead of fixing the test that matters.

Make CI unable to go silently green

The last layer is CI, and the theme is the same: every configuration that could produce a green run without running the assertions is closed off.

  • Tests hash their full inputs. A guard script fails the build if any test task in Turborepo narrows its cache inputs. The failure it prevents is subtle: exclude one data folder from the hash, and a change to that folder restores a cached pass while the specs that read it never execute. Its comment is the rule: a test task takes the default input set, whole, because a smaller cache key is “not worth a silent green.”
  • Shards are defined by exclusion. Tests run as a matrix of shards, and the last shard is written as negative filters (“everything except the three big apps”), so a newly added package lands in it automatically instead of never running in CI.
  • A red PR reports every failure. The matrix sets fail-fast: false, so one broken shard does not hide failures in the others, and an aggregator job reports a single required test check.
# .github/workflows/ci.yml (the shape of it)
test-shard:
  strategy:
    fail-fast: false
    matrix:
      shard:
        - { name: web, filter: '--filter=./apps/web' }
        - { name: admin, filter: '--filter=./apps/admin' }
        - { name: rest, filter: '--filter=!./apps/web --filter=!./apps/admin' }
  steps:
    - run: pnpm turbo run test ${{ matrix.shard.filter }}

What I would still add

Three gaps remain, and each is a known way for an agent to produce a green run that means nothing:

  • A lint rule for focused and skipped tests. Zero .only and .skip is a convention today. A single rule would make it a guarantee.
  • Mutation testing on the core packages (a tool that deliberately breaks the code to check that the tests notice). The manual version exists (the “remove the memo and this turns red” comments, the gutted-body assertion), but a tool would check it for every test, not just the ones someone annotated.
  • A “must fail first” check in the PR skill. The review skill could check out the test commit alone, run it against the old code, and require red. That would turn the most important step in the loop from a habit into a gate.

None of this makes the agent’s code easier to read. It makes most of it unnecessary to read. The cost argument is the one in who pays for AI slop in code: review is where the cost of agent output lands, and a correct failing test is the cheapest review there is. For how tests fit the wider set of techniques, see harness engineering in 2026.

Claude Code TDD FAQ

Does TDD work with Claude Code?
It works when a human decides what the failing test asserts and someone watches it fail before the fix. Telling the agent to do TDD and letting it write both the test and the code tends to produce tests that confirm whatever the code does, which is the failure mode Birgitta Böckeler describes on Martin Fowler's site.
How do I stop Claude Code from writing too many tests?
Give it a written testing policy it can read, pointed to from CLAUDE.md. List what always gets a test (escaped-bug regressions, user-facing behavior), what needs a stated reason, and what never gets one (things the linter or type checker already reject, vendored code, snapshots with no meaning, mock-was-called assertions).
What is a fake pass in testing?
A test that passes without checking anything: it asserts over an empty list, checks that a no-op function was called, uses a matcher that ignores the relevant attribute, or runs in an environment that skips the code path. If a new test is green before any code exists, treat it as a fake pass and rewrite it.
How do I make Claude Code run tests before it finishes?
Register a Stop hook that runs the tests for the workspaces the session changed and exits 2 with the failure output when they fail. Check stop_hook_active in the hook input to avoid blocking forever; Claude Code also caps consecutive stop-hook continuations at eight.
Should the agent be allowed to edit a failing test?
Not in the same turn that makes it pass. Split the work: one turn writes the failing test and stops, a human reads it, and a second turn implements without touching the test. Changing a test is fine as a separate, visible decision.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Coding agents