Crawlability and indexability get treated as one property, usually summed up as “can Google see my page”. They are two separate gates with separate controls. A page can be perfectly crawlable and never indexed, and a page that Google is forbidden to crawl can still turn up in results. If you diagnose a missing page without knowing which gate it failed, you will reach for the wrong fix.
Crawling fetches, indexing decides
Conductor’s guide to controlling crawling and indexing draws the line cleanly. Crawlers find and fetch URLs. Indexers then analyse the content, process canonicals and decide which pages to include. The first step is about access; the second is about judgement.
That split matters because the two steps fail for different reasons. A crawl failure is mechanical: a disallow rule, a server error, a page no link points to. An indexing failure is often a decision. Google fetched the page, looked at it and chose not to keep it, or chose to keep a different URL in its place. The fixes do not overlap much.
- Crawlability asks whether a crawler can discover the URL and fetch it.
- Indexability asks whether a fetched page is allowed, and judged worth, a place in the index.

Which control affects crawlability and indexability
The most useful thing in Conductor’s guide is a plain table of which method touches which gate. Reproduced here in summary, it explains most of the confusion people bring to this topic.
| Method | Affects crawling | Affects indexing |
|---|---|---|
| robots.txt | Yes | Yes |
| Meta robots / X-Robots-Tag | Slightly | Yes |
| Canonical tag | No | Yes |
| hreflang | No | No (prevents duplicates) |
| Pagination attributes | No | No |
Two rows deserve a second look. robots.txt is the only method marked as affecting both, and its effect on indexing is not the one most people expect. And the canonical tag does nothing to crawling at all: Google still fetches the duplicate in order to read the tag. If you add canonicals hoping to save crawl activity, you have picked a tool for the wrong gate.

The robots.txt trap
The received wisdom is that robots.txt keeps pages out of Google. It does not. It keeps Google from fetching them. Conductor warns that robots.txt-blocked URLs “can still appear in search results”, typically when other pages link to them. Google knows the address exists; it simply cannot read what is there, so the listing appears without a proper description.
The worse mistake follows from that. Someone adds a noindex tag to a page and also disallows it in robots.txt, belt and braces. Conductor’s advice is to never combine the two, and the reason is mechanical: if Google cannot crawl the page, it never sees the noindex. The directive sits there unread while the URL lingers in results. Our guide to robots.txt for SEO covers what the file is actually good for, and blocked by robots.txt covers the Search Console status this produces.
The rule that follows is short. To keep a page out of the index, let Google crawl it and tell it noindex. To stop Google spending requests on a section, use robots.txt and accept that the URLs may still be listed.

Why indexability fails on crawlable pages
A page that passes the crawl gate can still fail the index gate for several reasons. Some are directives you set; others are judgements Google makes.
- A noindex directive, in the meta robots tag or an X-Robots-Tag header. This one is an instruction, and Google follows it. See our meta robots tag guide.
- A canonical pointing elsewhere. Conductor calls canonicals “a guideline, rather than a directive”. Google can ignore yours, in either direction, if the signals disagree.
- Duplication. If Google decides another URL says the same thing, it indexes that one and folds yours in.
- Quality. Google fetched the page and decided it was not worth keeping. That shows up as crawled, currently not indexed, and no technical fix will move it.
The last two are where most of the frustration lives, because nothing on the page is technically wrong. The page is crawlable and indexable in the technical sense, and still absent.

How to check website crawlability, then indexability
Work in order: the crawl gate first, because a page Google cannot fetch cannot be judged.
- Test robots.txt. Confirm the URL is not disallowed for Googlebot. A broad rule written for one folder often catches more than intended.
- Check the response. The page should return a 200. Server errors and redirect loops stop the crawl before content is read.
- Check discovery. Is the page linked from somewhere Google already crawls, and is it in the sitemap? An orphaned page is technically crawlable but rarely found.
- Read the directives. Look for noindex in the HTML and in the response headers. Header directives are easy to miss because they never appear in the page source.
- Compare canonicals. URL Inspection shows the canonical you declared and the one Google chose. When they differ, Google has overruled you.
- Read the Page indexing status. It tells you which gate the URL failed and, roughly, why.

On large sites the crawl gate also has a capacity dimension. Google will not fetch every URL as often as you might like, which is where crawl budget comes in. For many smaller sites it is not the binding constraint, and indexing quality is.
What the controls cannot do
The limitation is that only some of this is in your hands. Directives such as noindex are obeyed; signals such as canonicals are weighed against everything else Google sees. You can make a page crawlable and indexable in every technical sense and Google can still decline to index it. The table above tells you which lever reaches which gate. It cannot tell you whether Google will judge the page worth keeping.

That is also why fixing crawlability rarely rescues weak content. Opening robots.txt or adding a page to the sitemap gets it fetched. Whether it is kept depends on what is there once it is read. If a crawlable page has sat unindexed for months, the question worth asking is whether it says anything another page on your site does not already say. If it does not, merging it into the stronger page is usually a better use of time than another round of technical checks. And if a page should not be in search at all, the cleanest instruction is a noindex Google can actually read, with robots.txt left open for that URL.
Frequently asked questions
What is the difference between crawlability and indexability?
Crawlability is whether Google can find and fetch a URL. Indexability is whether a fetched page is allowed and judged worth including in the index. A page must pass the first to reach the second.
Does robots.txt stop a page being indexed?
Not reliably. It stops crawling, but blocked URLs can still appear in search results if other pages link to them. Use noindex on a crawlable page to keep it out.
Should I use robots.txt and noindex together?
No. If robots.txt blocks the page, Google cannot crawl it and never sees the noindex, so the directive has no effect.
Does a canonical tag guarantee which URL is indexed?
No. Conductor describes canonicals as a guideline rather than a directive. Google weighs them against other signals and can choose a different URL.
The takeaway Diagnose which gate a page failed before fixing anything. Use robots.txt to manage crawling, noindex to manage indexing, and never both on the same URL.

