Blog
4 min read

Playwright Tutorial: End-to-End Tests That Aren't Flaky

Playwright drives real browsers to test your app the way users use it. Install it, write your first test, use role-based locators and auto-waiting assertions, log in once and reuse the session, run against a dev server, debug with UI mode and traces, and run it in CI.

Unit tests tell you functions work. End-to-end tests tell you the app works: a real browser loads your pages, clicks your buttons and checks what appears. (Unit vs integration vs E2E tests) Playwright is the most popular tool for this today — fast, reliable, and it drives Chromium, Firefox and WebKit (Safari's engine).

Install

In your project:

npm init playwright@latest

It asks a few questions, creates playwright.config.ts, an example test in tests/, and downloads browsers. Run the example:

npx playwright test
npx playwright show-report

Your first real test

// tests/signup.spec.ts
import { test, expect } from '@playwright/test'

test('visitor can sign up and reach the dashboard', async ({ page }) => {
  await page.goto('/signup')

  await page.getByLabel('Email').fill(`test+${Date.now()}@example.com`)
  await page.getByLabel('Password').fill('a-long-test-password')
  await page.getByRole('button', { name: 'Create account' }).click()

  await expect(page).toHaveURL(/\/dashboard/)
  await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible()
})

Read it like a user story. That's the goal.

Locators: find things the way users do

Prefer locators based on what a user sees and what assistive technology reads:

Locator Finds
getByRole('button', { name: 'Save' }) A button labelled Save — the best default
getByLabel('Email') The input with that label
getByText('Order confirmed') Visible text
getByPlaceholder('Search') Input by placeholder
getByTestId('cart-total') data-testid="cart-total" — when nothing else fits

Avoid CSS selectors like .btn-primary > span:nth-child(2) — they break whenever the markup changes. Role-based locators also nudge your app towards being accessible. (Web accessibility basics)

Auto-waiting: why Playwright isn't flaky (if you let it)

Actions like click() wait for the element to be visible, enabled and stable. expect(...) assertions retry until they pass or time out:

await expect(page.getByText('Saved')).toBeVisible()   // waits for it to appear

So:

  • Never use page.waitForTimeout(3000). Fixed sleeps are the main cause of flaky tests — too short on a slow CI machine, wasted time everywhere else.
  • Use web-first assertions (await expect(locator).toHaveText(...)) rather than reading a value and comparing it yourself.

Run against your dev server automatically

In playwright.config.ts:

export default defineConfig({
  use: { baseURL: 'http://localhost:3000', trace: 'on-first-retry' },
  webServer: {
    command: 'npm run dev',
    url: 'http://localhost:3000',
    reuseExistingServer: !process.env.CI,
  },
})

Playwright starts the app before tests and stops it after.

Log in once, reuse it

Logging in through the UI in every test is slow. Do it once in a setup project and save the session:

// tests/auth.setup.ts
import { test as setup } from '@playwright/test'

setup('authenticate', async ({ page }) => {
  await page.goto('/login')
  await page.getByLabel('Email').fill(process.env.E2E_USER!)
  await page.getByLabel('Password').fill(process.env.E2E_PASSWORD!)
  await page.getByRole('button', { name: 'Log in' }).click()
  await page.waitForURL('/dashboard')
  await page.context().storageState({ path: 'playwright/.auth/user.json' })
})
// playwright.config.ts
projects: [
  { name: 'setup', testMatch: /.*\.setup\.ts/ },
  {
    name: 'chromium',
    use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' },
    dependencies: ['setup'],
  },
]

Add playwright/.auth to .gitignore. Use a dedicated test account against a test database, never real users. (Database seeding)

Debugging failing tests

  • UI mode: npx playwright test --ui — watch tests run step by step, time-travel through each action, see the DOM at every point. The best way to understand a failure.
  • Trace viewer: with trace: 'on-first-retry', failed CI runs produce a trace with screenshots, network requests and console logs. Open it with npx playwright show-trace.
  • Codegen: npx playwright codegen localhost:3000 records your clicks as test code — a good starting point; clean up the locators afterwards.
  • await page.pause() stops a test and opens the inspector.

Keeping tests reliable

  • Independent tests. Each test sets up its own data; never rely on another test having run first.
  • Unique data. Use timestamps or random suffixes for emails and names so parallel runs don't collide.
  • Control the network when needed. page.route() can stub third-party APIs (payments, AI calls) so tests don't depend on them.
  • Test what matters. A handful of critical flows — sign up, log in, the core action, checkout — beats hundreds of fragile UI checks.

In CI

- run: npm ci
- run: npx playwright install --with-deps chromium
- run: npx playwright test
- uses: actions/upload-artifact@v7
  if: ${{ !cancelled() }}
  with:
    name: playwright-report
    path: playwright-report/

Upload the report so you can open traces from failed runs. (CI with GitHub Actions)

Playwright and AI agents

E2E tests are a superb "definition of done" for coding agents: "the feature is done when npx playwright test tests/checkout.spec.ts passes." Agents can also use Playwright directly to look at the app they're building. Ask them to write tests with role-based locators and no fixed waits. (Getting AI to write tests that catch bugs)

The summary

  • npm init playwright@latest, then write tests that read like user stories.
  • Use getByRole/getByLabel locators and web-first expect assertions — no sleeps.
  • Start the dev server via webServer; log in once with storageState.
  • Debug with UI mode and traces; run in CI and upload the report.

EasySpawn gives Claude Code a persistent server with your app, database and a real browser available, so it can run Playwright suites against the live app on every change. See how it works or join the waitlist.

Related: Unit vs Integration vs E2E Tests · How to Test Your App Before Launch · Getting AI to Write Tests · Set Up CI With GitHub Actions

Keep reading