Robots txt for SEO is mostly a story of a file being asked to do a job it was never designed for. People disallow pages to keep them out of Google, block folders to “save crawl budget” on sites with a few hundred URLs, and copy long rule sets from templates without knowing what each line does. The file is simpler than that, and the most important thing to know about it is what it cannot do.
What robots.txt actually controls
Google’s documentation defines it in one sentence: “A robots.txt file tells search engine crawlers which URLs the crawler can access on your site.” Access is the key word. The file governs crawling: whether a well-behaved bot requests a URL at all. It says nothing directly about whether that URL can appear in search.
A minimal file looks like this:
Each group names a user agent, then lists paths it should not request. The sitemap line is optional but useful: it tells crawlers where your list of indexable URLs lives.

What it doesn’t do: keep pages out of Google
This is the part that catches people out. Google states that robots.txt “is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with noindex or password-protect the page.” It goes further: “A page that’s disallowed in robots.txt can still be indexed if linked to from other sites.” The URL and the anchor text of those links can appear in results, even though Google never read the page.
| You want to | Use | Why not robots.txt |
|---|---|---|
| Stop crawlers requesting low-value URLs | robots.txt disallow | This is what it is for |
| Keep a page out of search results | noindex | A disallowed URL can still be indexed if linked from other sites |
| Keep private content private | Password protection | The file is public and only asks; it enforces nothing |
| Remove an indexed page using noindex | noindex, with crawling allowed | If the page is blocked, Google never sees the noindex |
| Opt out of Gemini training and grounding | robots.txt rule for Google-Extended | This one is a robots.txt job; it does not affect Search |
The fourth row is the trap. Add noindex to a page and block it in robots.txt at the same time, and Google cannot crawl the page to see the noindex. Our post on excluded by noindex tag covers how noindex works and how to check it is being read.

Is robots.txt necessary for SEO?
Not always. Google notes that not every site needs to edit one, and hosted platforms often manage it for you. If your site is a few hundred pages with clean URLs, a missing or near-empty file is fine. Nothing is blocked, and Google crawls what your links and sitemap point it to.
Robots.txt earns its keep when a site generates far more URLs than it has real pages. The usual culprits:
- Faceted navigation. Shop filters that combine colour, size, price and sort order into endless parameter URLs.
- Internal search results. Every query a visitor types becomes a crawlable page.
- Cart, checkout and account paths that have no search value and no reason to be fetched.
- Calendar and session-parameter URLs that go on forever.
Blocking those keeps crawlers spending their visits on pages you want found. If a site has a backlog of URLs Google has found but not crawled, this is one of the levers; see crawled vs discovered, currently not indexed.

Robots.txt for AI SEO: Google-Extended
The file has picked up a second job. Google uses a separate token, Google-Extended, to control whether content Google crawls may be used to train future Gemini models and for grounding in Gemini Apps and Vertex AI. To opt out:
Google says Google-Extended “does not impact a site’s inclusion in Google Search nor is it used as a ranking signal.” That makes the decision a business one rather than an SEO one. Blocking it will not cost you rankings, and allowing it will not earn you any. Whether it affects how often your content is cited in AI answers is a separate question that Google’s statement does not settle. Our post on AI Overviews and traffic covers the wider picture.
Other AI companies publish their own user-agent tokens. Each needs its own group in the file, and each only works if the crawler chooses to honour it.

The mistakes that cost traffic
A broken robots.txt is one of the few SEO errors that can take a whole site out of circulation in a day. The ones to check for:
disallow: /underuser-agent: *, usually left over from a staging site. It blocks everything.- Blocking CSS and JavaScript. Google renders pages; if it cannot fetch the files that build the layout, it sees a different page from your visitors.
- Disallowing pages you also noindexed, so the noindex is never read.
- Over-broad patterns. A rule meant for
/tag/written as/tblocks every path starting with that letter.
Yoast’s guide to robots.txt walks through the syntax rules in detail, including wildcards and how conflicting rules are resolved. Test any change before it goes live; the robots.txt report in Search Console shows which version Google last fetched.

What robots.txt cannot promise
The limitation is built into the format: robots.txt is a request, not a lock. Google and other major crawlers honour it. Scrapers and badly built bots can ignore it entirely, and the file itself is public, so listing a sensitive path in it advertises that the path exists. Anything that must stay private needs authentication.

The practical rule is short. Keep the file as small as your site allows. Block parameter spaces, internal search and utility paths if you have them in volume. Add a sitemap line. Use noindex, not disallow, for pages you want kept out of results, and make the AI-crawler decision on business grounds, knowing that Google-Extended does not touch Search.
Review the file whenever the site changes shape: a new platform, a new shop filter, a migration. Those are the moments when a leftover rule blocks something important, and the drop shows up in your reports weeks later. If indexing is the wider worry, our guide to getting a site indexed by Google covers the steps after crawling, and the strategy archive collects the related technical posts.
Frequently asked questions
Can robots.txt remove a page from Google?
No. Google says robots.txt is not a mechanism for keeping a page out of Google. A disallowed page can still be indexed if other sites link to it. Use noindex or password protection.
Does every site need a robots.txt file?
No. Not every site needs to edit one, and hosted platforms often manage it. It matters most on sites that generate many low-value URLs.
Does blocking Google-Extended hurt rankings?
No. Google says Google-Extended does not impact a site’s inclusion in Google Search and is not used as a ranking signal.
Should I block noindexed pages in robots.txt?
No. If the page is blocked, Google cannot crawl it to see the noindex, and the URL can stay in results.
The takeaway Robots.txt controls crawling, not indexing. Keep it short, use it for URL spaces with no search value, and use noindex when the goal is to stay out of results.

