Blog
4 min read

robots.txt and sitemap.xml Explained: A Beginner's Guide

Two small files that tell search engines (and AI crawlers) what to crawl and what exists on your site. How robots.txt and sitemap.xml work, examples, the robots.txt mistake that hides your whole site, why Disallow doesn't remove pages from Google, and how to generate both in Next.js.

Two small text files sit at the root of most websites and quietly shape how search engines see them: robots.txt and sitemap.xml. One says where crawlers may go; the other lists the pages you want found. Both take minutes to set up, and getting robots.txt wrong can hide your entire site from Google.

What a crawler is

Search engines discover pages with crawlers (also called bots or spiders) — programs that fetch pages and follow links. Google's is Googlebot. AI companies run their own crawlers too. Before crawling, well-behaved bots read your robots.txt.

robots.txt: the rules for crawlers

It lives at exactly https://yourdomain.com/robots.txt. A typical one:

User-agent: *
Disallow: /admin/
Disallow: /api/

Sitemap: https://yourdomain.com/sitemap.xml
  • User-agent: * — these rules apply to all crawlers.
  • Disallow: /admin/ — don't crawl anything under /admin/.
  • Sitemap: — where your sitemap is.

You can write rules for specific bots by name, for example to allow or block particular AI crawlers:

User-agent: GPTBot
Disallow: /

User-agent: *
Allow: /

The mistake that hides your whole site

User-agent: *
Disallow: /

This blocks everything. It's often added on a staging site to keep it out of Google — then accidentally copied to production. If your site has vanished from search, check this first.

robots.txt is not security or a "remove from Google" button

Two important misunderstandings:

  1. It's a polite request. Well-behaved crawlers obey it; malicious ones ignore it. Never rely on it to hide private pages — protect them with a login. (And listing /secret-admin-panel/ in robots.txt advertises it.)
  2. Disallow doesn't de-index. Blocking a page stops Google crawling it, but if other sites link to it, the URL can still appear in results. To keep a page out of search results, let it be crawled and add a noindex tag:
<meta name="robots" content="noindex">

sitemap.xml: the list of your pages

A sitemap is a list of the URLs you want search engines to know about:

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://yourdomain.com/</loc>
    <lastmod>2026-10-01</lastmod>
  </url>
  <url>
    <loc>https://yourdomain.com/pricing</loc>
    <lastmod>2026-09-15</lastmod>
  </url>
</urlset>

Things worth knowing:

  • Use full, canonical URLs — the exact version you want indexed (https, with or without www consistently; see www vs non-www).
  • lastmod helps if it's accurate. Google uses it when it reflects real changes. Don't set every page to today's date.
  • priority and changefreq are ignored by Google. You'll see them in old examples; you can leave them out.
  • Limits: 50,000 URLs or 50 MB per sitemap file. Larger sites use a sitemap index pointing at several sitemaps.
  • Only include pages you want indexed — no redirects, 404s, or noindex pages.

A sitemap doesn't guarantee indexing; it helps search engines find pages, especially on new sites with few links.

Generating them in Next.js

Next.js (App Router) can generate both from code — handy because the sitemap updates itself when you add pages:

// app/robots.ts
import type { MetadataRoute } from "next";

export default function robots(): MetadataRoute.Robots {
  return {
    rules: { userAgent: "*", allow: "/", disallow: ["/admin/", "/api/"] },
    sitemap: "https://yourdomain.com/sitemap.xml",
  };
}
// app/sitemap.ts
import type { MetadataRoute } from "next";

export default function sitemap(): MetadataRoute.Sitemap {
  return [
    { url: "https://yourdomain.com/", lastModified: new Date("2026-10-01") },
    { url: "https://yourdomain.com/pricing", lastModified: new Date("2026-09-15") },
  ];
}

For other frameworks, plugins or a static file in public/ work fine.

Submit your sitemap

Add your site to Google Search Console, then submit the sitemap URL under Sitemaps. Search Console will tell you how many URLs it found and any problems. See how to get your website on Google.

The summary

  • robots.txt tells crawlers where they may go; sitemap.xml lists the pages you want found.
  • Disallow: / blocks your entire site — check for it if you've vanished from Google.
  • robots.txt isn't security, and Disallow isn't removal; use noindex to keep pages out of results.
  • Keep sitemaps to canonical, indexable URLs with honest lastmod dates, and submit them in Search Console.

EasySpawn serves your app on your own domain with HTTPS from day one, so your robots.txt and sitemap have a real, public, canonical address from the start — and Claude Code can generate both and check them on the live site. See how it works or join the waitlist.

Related: SEO Basics for Your App · Anatomy of a URL · Subdomain vs Subdirectory · Dev, Staging, and Production Explained

Keep reading