Test-Driven Development With Claude Code
Tests first is the single best way to make an AI coding agent reliable: it gives Claude a pass/fail signal to iterate against. A red-green-refactor workflow for Claude Code, prompts that stop it cheating the tests, enforcing it with hooks and /goal, and where TDD with agents falls short.
The most important thing you can give a coding agent is a way to check its own work. Without one, Claude stops when the code looks done and you become the test suite. (Claude Code best practices)
Test-driven development — write the test first, watch it fail, make it pass, clean up — is the most direct way to provide that check. It turns out to suit agents even better than it suits humans.
Why TDD works so well with agents
- A precise target. A failing test is an unambiguous spec: input, expected output, edge cases.
- A loop the agent can close alone. Claude runs the tests, reads the failures, edits, reruns — without waiting for you.
- Protection against regressions from the agent's next change, or the one after.
- Reviewable intent. Reviewing ten lines of tests is easier than reviewing two hundred lines of implementation. (Review an AI-written PR)
The workflow
1. Red: write the tests first, and only the tests
We're adding
calculateInvoiceTotal(lines, { taxRate, discount }). Write tests only — no implementation. Cover: empty invoice, a single line, multiple lines, percentage discount, discount larger than the subtotal (total never below zero), rounding to 2 decimals, invalid tax rate throws. Use Vitest. Don't create the function yet.
Then:
Run the tests and confirm they fail for the right reason.
"Fails because the function doesn't exist" is fine. A test that passes before any implementation is testing nothing.
2. Review the tests yourself
This is your main job. Are the cases right? Is the rounding rule what the business wants? Is anything missing? Fix the tests now — they're the spec. (Unit vs integration vs E2E tests)
Commit them. (Claude Code and Git)
3. Green: implement until they pass
Implement calculateInvoiceTotal so all tests pass. Do not modify the test file. Run the full test suite when done and show me the output.
4. Refactor
Tests are green. Refactor for clarity without changing behaviour; rerun tests after each change.
Stop it gaming the tests
Agents under pressure to make tests pass sometimes take shortcuts: special-casing test inputs, weakening assertions, skipping tests, or mocking away the thing under test. Guard against it:
- Say it explicitly: "Don't edit, skip or delete tests. If a test seems wrong, stop and tell me why."
- Commit the tests first, so any change to them shows up in the diff.
- Block edits with a hook — a
PreToolUsehook can refuse writes to*.test.tsduring implementation. (Claude Code hooks) - Prefer real dependencies (a test database) over mocks where practical. (How to test an API)
- Get a second opinion:
/code-reviewin a fresh context, or a subagent asked to check the implementation isn't special-casing.
Make the loop automatic
- Put the test command in CLAUDE.md — "Run
pnpm testbefore declaring any task done." (How to write a CLAUDE.md) - A Stop hook can run the suite and block Claude from finishing while it's red.
/goalsets a completion condition ("all tests ininvoice.test.tspass") that an evaluator checks after every turn, so Claude keeps going until it's met.
A skill for the team
Package the workflow as a skill so anyone can run it:
---
name: tdd
description: Implement a feature test-first
---
1. Write failing tests for the requested behaviour, covering edge cases. Do not implement.
2. Run them; confirm they fail for the right reason. Stop and show the tests for review.
3. After approval, implement until green without changing tests.
4. Refactor; rerun the full suite; show the output.
Where it falls short
- UI and visual work — use screenshot comparison or Playwright checks instead. (Playwright MCP, Playwright tutorial)
- Exploratory work where you don't know the behaviour yet — spike first, then pin it down with tests.
- Tests only check what they check. A green suite with thin coverage is false comfort. (Getting AI to write tests that catch bugs)
EasySpawn gives Claude Code a persistent server with your test database and services already running, so the red-green loop works the same at 2 a.m. as when you're watching. See how it works or join the waitlist.
Related: Getting AI to Write Tests That Catch Bugs · Spec-Driven Development · Claude Code Best Practices · Vitest vs Jest
Keep reading
Spec-Driven Development With Claude Code: Write the Spec, Then Let the Agent Build
Spec-driven development means agreeing on a written specification before an AI agent writes code. What a good spec contains, a practical workflow with Claude Code (spec → plan → tasks → implement → verify), templates, and when it's overkill.
How to Resume a Claude Code Session (and What Resuming Can't Bring Back)
claude --continue and claude --resume reopen yesterday's conversation in seconds. But a resumed session restores the conversation, not the world it was working in. The commands, the habits that make resuming reliable, and the gap between conversation state and environment state.