Blog
4 min read

Test-Driven Development With Claude Code

Tests first is the single best way to make an AI coding agent reliable: it gives Claude a pass/fail signal to iterate against. A red-green-refactor workflow for Claude Code, prompts that stop it cheating the tests, enforcing it with hooks and /goal, and where TDD with agents falls short.

The most important thing you can give a coding agent is a way to check its own work. Without one, Claude stops when the code looks done and you become the test suite. (Claude Code best practices)

Test-driven development — write the test first, watch it fail, make it pass, clean up — is the most direct way to provide that check. It turns out to suit agents even better than it suits humans.

Why TDD works so well with agents

  • A precise target. A failing test is an unambiguous spec: input, expected output, edge cases.
  • A loop the agent can close alone. Claude runs the tests, reads the failures, edits, reruns — without waiting for you.
  • Protection against regressions from the agent's next change, or the one after.
  • Reviewable intent. Reviewing ten lines of tests is easier than reviewing two hundred lines of implementation. (Review an AI-written PR)

The workflow

1. Red: write the tests first, and only the tests

We're adding calculateInvoiceTotal(lines, { taxRate, discount }). Write tests only — no implementation. Cover: empty invoice, a single line, multiple lines, percentage discount, discount larger than the subtotal (total never below zero), rounding to 2 decimals, invalid tax rate throws. Use Vitest. Don't create the function yet.

Then:

Run the tests and confirm they fail for the right reason.

"Fails because the function doesn't exist" is fine. A test that passes before any implementation is testing nothing.

2. Review the tests yourself

This is your main job. Are the cases right? Is the rounding rule what the business wants? Is anything missing? Fix the tests now — they're the spec. (Unit vs integration vs E2E tests)

Commit them. (Claude Code and Git)

3. Green: implement until they pass

Implement calculateInvoiceTotal so all tests pass. Do not modify the test file. Run the full test suite when done and show me the output.

4. Refactor

Tests are green. Refactor for clarity without changing behaviour; rerun tests after each change.

Stop it gaming the tests

Agents under pressure to make tests pass sometimes take shortcuts: special-casing test inputs, weakening assertions, skipping tests, or mocking away the thing under test. Guard against it:

  • Say it explicitly: "Don't edit, skip or delete tests. If a test seems wrong, stop and tell me why."
  • Commit the tests first, so any change to them shows up in the diff.
  • Block edits with a hook — a PreToolUse hook can refuse writes to *.test.ts during implementation. (Claude Code hooks)
  • Prefer real dependencies (a test database) over mocks where practical. (How to test an API)
  • Get a second opinion: /code-review in a fresh context, or a subagent asked to check the implementation isn't special-casing.

Make the loop automatic

  • Put the test command in CLAUDE.md — "Run pnpm test before declaring any task done." (How to write a CLAUDE.md)
  • A Stop hook can run the suite and block Claude from finishing while it's red.
  • /goal sets a completion condition ("all tests in invoice.test.ts pass") that an evaluator checks after every turn, so Claude keeps going until it's met.

A skill for the team

Package the workflow as a skill so anyone can run it:

---
name: tdd
description: Implement a feature test-first
---
1. Write failing tests for the requested behaviour, covering edge cases. Do not implement.
2. Run them; confirm they fail for the right reason. Stop and show the tests for review.
3. After approval, implement until green without changing tests.
4. Refactor; rerun the full suite; show the output.

(Claude Code skills)

Where it falls short


EasySpawn gives Claude Code a persistent server with your test database and services already running, so the red-green loop works the same at 2 a.m. as when you're watching. See how it works or join the waitlist.

Related: Getting AI to Write Tests That Catch Bugs · Spec-Driven Development · Claude Code Best Practices · Vitest vs Jest

Keep reading