Free to explore: filter winning sites by DR, traffic and niche  ·  Try the live explorer →

Meta robots tag and X-Robots-Tag explained

HTML meta tags in a code editor

A meta robots tag, in the definition used in Ahrefs’ guide by Michal Pecánek, is “an HTML snippet that tells search engine robots what they can and cannot do on a certain page.” It sits in the page’s <head> and looks like this:

That is the whole mechanism. The name attribute says which crawlers the rule applies to, with robots meaning all of them, and the content attribute lists the directives. The hard part is not the syntax. It is knowing which directive does what, and which other signals it collides with.

Meta robots tag directives, one by one

Each directive answers a narrow question. Most pages need none of them, because the default is that a page can be indexed and its links can be crawled. You add a directive only to take something away.

A noindex tag in page source
DirectiveWhat it tells crawlers
noindexDo not index this page
nofollowDo not crawl the links on this page
noneThe same as noindex plus nofollow
noarchiveDo not keep a cached copy of the page
nosnippetDo not show a text snippet for the page in results
max-snippetLimit the length of the text snippet
max-image-previewSet the image preview size: none, standard or large
unavailable_afterStop showing the page in results after a given date

Two of these are routinely misread. Page-level nofollow applies to every link on the page, which is rarely what anyone wants; if you mean to qualify individual links, the link attribute is the tool, as our guide to nofollow links explains. And none is not “no restrictions”. It is the most restrictive shorthand on the list.

The display directives are the ones worth using more often than people do. max-image-preview set to large allows a large image preview, while none suppresses it entirely, so a publisher that relies on visual results should check that a theme or plugin has not quietly set it lower. unavailable_after suits content with a natural end date, such as an event page or a time-limited offer, where you want the page out of results without having to remember to remove it.

Directives can be combined in one tag, separated by commas, as in the example above. Keep combinations deliberate. A long list copied from another site is a common way for a restriction nobody intended to end up on every page of a template.

What is the X-Robots-Tag?

The meta tag only works where there is HTML to put it in. For everything else there is a header. In Ahrefs’ words, “The X-Robots-Tag is an HTTP header sent from a web server”, and it accepts the same directives as the meta tag.

HTTP response headers in a browser panel

A response carrying it looks like this:

That makes it the answer for two jobs the meta tag cannot do. The first is non-HTML files: PDFs, images and other documents that have no <head>. The second is scale. A single server rule can apply a directive to a whole directory or file type, where the meta tag would have to be added to every template that renders those pages.

A folder of PDF documents

How to add an X-Robots-Tag header

The header is set in your server or CDN configuration rather than in page code. On Apache, a rule that keeps every PDF out of the index can be written in the site configuration or .htaccess:

Nginx, CDNs and most application frameworks have an equivalent way to add a response header. Whichever you use, check the result rather than the configuration, because what the server sends is what crawlers act on: request one of the affected files, for example https://example.com/files/brochure.pdf, and read the headers that come back. A rule that matches the wrong pattern fails silently, and so does one that matches too much.

A search snippet preview

Snippet directives deserve the same care. nosnippet and max-snippet change how a page appears in results, not whether it is indexed, so a mistake shows up as a worse-looking listing rather than a missing one. That makes it easy to miss for months.

Meta robots tag versus robots.txt

This is the collision that causes the most damage. Robots.txt controls crawling. The meta robots tag and X-Robots-Tag control indexing, and a crawler has to fetch the page to read them. Ahrefs’ warning is direct: do not put noindex on pages blocked in robots.txt, because crawlers cannot see the noindex, so the URL can stay indexed.

A robots file beside a meta tag

The received wisdom runs the other way. Blocking a page in robots.txt feels like the stronger instruction, so people add it on top of noindex for good measure, and the result is the opposite of what they wanted. If you want a URL out of the index, let it be crawled until the noindex has been seen. Our guide to robots.txt for SEO covers what blocking is actually for.

The same logic applies to mixed signals elsewhere. A page with noindex and a canonical pointing to another URL is asking Google to do two different things; our piece on the canonical tag explains when consolidation is the better choice than removal.

Auditing robots directives on a live site

Most robots problems are accidental: a staging noindex that shipped to production, a plugin setting that applied to a whole post type, a header rule written for one folder that matched another. Search Console reports pages it found with a noindex, and our guide to “Excluded by noindex tag” covers how to read that report and decide which entries are intended.

A developer editing server configuration

A crawl of your own site fills the gap Search Console leaves, because it shows you the directives on every URL you link to, including ones Google has not reported yet. Check both places a directive can live. A page can be clean in its HTML and still be noindexed by a header, and a check that only reads the source will miss it.

The limitation is the one every directive shares. These are instructions to crawlers, not access controls. Well-behaved search engine crawlers follow them; nothing in a meta tag or header stops a person or a badly behaved bot from fetching the page. If content must stay private, protect it with authentication. Use robots directives for what they are designed to do: decide what a search engine indexes and shows.

Frequently asked questions

What is a meta robots tag in SEO?

An HTML tag in a page’s head that tells search engine crawlers what they can and cannot do with that page, such as whether to index it or crawl its links.

What is a meta robots tag example?

<meta name="robots" content="noindex, nofollow"> tells all crawlers not to index the page and not to crawl its links.

What is the X-Robots-Tag?

An HTTP header sent by the web server that carries the same directives as the meta tag. Use it for non-HTML files such as PDFs and images, or to apply directives at scale.

Should I block a noindexed page in robots.txt?

No. If crawlers are blocked they cannot see the noindex, and the URL can stay indexed.

The takeaway Use the meta robots tag for HTML pages and the X-Robots-Tag for files and rules at scale. Never block a noindexed URL in robots.txt, and check the headers as well as the source when you audit.