Blog
3 min read

How to Evaluate LLM Outputs: Evals for AI Features in Your App

You can't improve an AI feature you can't measure. How to build an eval set from real inputs, the kinds of checks (exact match, code-based assertions, LLM-as-judge, human review), avoiding judge bias, running evals in CI, and tracking quality, cost and latency when you change prompts or models.

Changing a prompt, switching to a cheaper model, adding a retrieval step — each can make your AI feature better or quietly worse. Evals (evaluations) are how you find out before your users do. They're to AI features what tests are to code.

Step 1: build an eval set from real inputs

Collect 30–100 real examples of what users send — from logs, support tickets, your own usage. Include:

  • typical cases,
  • hard cases (long, ambiguous, multilingual),
  • edge cases (empty input, off-topic, abusive),
  • cases that previously went wrong.

For each, write down what a good answer must do — not necessarily the exact answer.

{ "input": "Can I get a refund after 45 days?", "must": ["mentions 30-day policy", "offers to connect to support"], "must_not": ["promises a refund"] }

Store them in your repository, versioned like code.

Step 2: choose how to grade

Exact or structural checks (cheapest, most reliable)

For classification, extraction and structured output, compare directly:

expect(result.category).toBe('billing')
expect(Schema.safeParse(result).success).toBe(true)

(Structured output from LLMs)

Code-based assertions

Check properties: length limits, required keywords, no URLs outside your domain, valid JSON, response under N tokens, cites a source.

LLM-as-judge

For open-ended text, ask a model to grade against a rubric:

Grade the answer against the criteria. Reply with JSON:
{"mentions_policy": true|false, "promises_refund": true|false, "tone_ok": true|false, "reason": "..."}

Criteria: ...
Answer: ...

Make it work:

  • Specific, binary criteria beat "rate 1–10".
  • Ask for the reasoning before the verdict.
  • Check the judge: grade 20 examples by hand and compare. Judges drift and can favour longer answers or their own model family.
  • Use a different or stronger model as judge when you can.

Human review

Still essential for a sample, especially early, and for anything high-stakes. Lightweight: a spreadsheet of outputs with thumbs up/down and a note.

Step 3: measure more than quality

For every eval run, record:

  • pass rate per criterion,
  • cost (tokens × price), (Claude API pricing)
  • latency (time to first token and total),
  • failures — read them; they're where the insight is.

A change that improves quality 2% while doubling cost or latency may not be worth it.

Step 4: run them on every change

Treat prompts and model choices like code:

  1. Change the prompt, model, temperature or retrieval.
  2. Run the eval set.
  3. Compare against the previous run.
  4. Ship only if it's better (or equal and cheaper).

Put it in CI so a prompt change in a pull request shows its eval results. (GitHub Actions CI basics)

Tools exist for this — promptfoo, Braintrust, LangSmith, Langfuse and others — but a script with a JSON file and a results table is a fine start.

Step 5: learn from production

  • Log inputs and outputs (respecting privacy). (What is OpenTelemetry?)
  • Add a feedback button in your UI.
  • Turn every bad production output into a new eval case. Your eval set should grow from real failures.

Special cases

  • RAG: evaluate retrieval separately (did the right chunks come back?) from generation (did the answer use them faithfully?). (What is RAG?, RAG chunking)
  • Agents: evaluate final outcomes on realistic tasks, plus step counts and cost — not individual messages. (Evals for coding agents)
  • Safety: include adversarial inputs (prompt injection attempts, requests to break rules). (AI guardrails)

Common mistakes

  • Judging by "it looked good on the three examples I tried".
  • An eval set of only easy cases.
  • Trusting LLM-as-judge without checking it.
  • Not re-running evals when the provider updates the model.

EasySpawn gives your AI app a server where eval scripts, logs and the app itself live together — Claude Code can run the eval suite and compare results before you ship a prompt change. See how it works or join the waitlist.

Related: Evals for Coding Agents · What Is Prompt Engineering? · How to Build an AI App · AI Guardrails Explained

Keep reading