Free to explore: filter winning sites by DR, traffic and niche  ·  Try the live explorer →

Crawl budget: who needs to care

A crawl stats report chart on a monitor

What is crawl budget in SEO? It is the number of URLs Google can and wants to crawl on your site in a given period. The term gets blamed for a great many indexing problems, usually on sites far too small for it to matter. The useful question is not how to increase your crawl budget. It is whether you have a crawl budget problem at all, and for most sites the answer is no.

Who Google says crawl budget is for

Google Search Central’s crawl budget guide is unusually specific about its audience. It names three kinds of site.

Site typeSizeHow often content changes
Large sites1 million+ unique pagesModerately often (once a week)
Medium or larger sites10,000+ unique pagesVery rapidly (daily)
Sites with an indexing backlogAnyMany URLs in “Discovered – currently not indexed”

If your site is a few hundred or a few thousand pages, updated now and then, you are not on that list. The received wisdom that every site should manage its crawl budget comes from advice written for large online shops and publishers and applied to everyone else. On a small site, Google can usually crawl everything it wants to.

A list of faceted navigation URLs

What is crawl budget in SEO, exactly?

Google describes two forces that together set how much it crawls.

  • Crawl capacity limit. The time Google’s crawler can spend on your server without overwhelming it. A fast, healthy server raises it; slow responses and errors lower it.
  • Crawl demand. How much Google wants to crawl, driven by “a site’s size, update frequency, page quality, and relevance, compared to other sites.”

Crawl budget is what you get where the two meet. Capacity is the ceiling your server sets; demand is how much of that ceiling Google chooses to use. The second point is the one most crawl budget advice misses. Google crawls more of a site it considers valuable, so the most reliable way to raise demand is to have pages worth crawling.

Duplicate parameter URLs in a crawl

How to optimise crawl budget

Google’s recommendations are almost all about waste: stopping the crawler spending time on URLs that should not exist or do not matter.

Manage your URL inventory

Large sites generate URLs faster than they generate content. Faceted navigation, sorting options, tracking parameters and session IDs can turn one category page into thousands of near-identical URLs. Google’s phrase for the goal is to focus on “unique content rather than unique URLs.”

Consolidate duplicates so one URL represents each piece of content; our guide to duplicate content covers the options. For URLs that should not be crawled at all, Google’s advice is to block them with robots.txt rather than noindex. That runs against a common habit. A noindexed page still has to be crawled for Google to see the tag, so noindex keeps the waste in place; a robots.txt rule stops the request.

The two tools do different jobs, and mixing them up causes problems of its own. Blocking a URL in robots.txt stops crawling, but it also stops Google seeing any canonical or noindex tag on that page, so do not block URLs whose signals you need Google to read. A canonical tag consolidates signals between duplicates Google can still crawl. Our guide to robots.txt for SEO covers how to write rules for parameters and faceted paths without blocking pages you want indexed.

Start with the biggest sources of waste. On most large sites that is faceted navigation: every combination of filters, sorts and page numbers creates a new address for content that already exists. Decide which filter combinations deserve their own indexable pages, because people search for them, and keep the crawler away from the rest.

A soft 404 pages report

Eliminate soft 404s

A soft 404 is a page that tells users the content is missing while returning a success code to crawlers. Google is blunt about the cost: soft 404 pages “will continue to be crawled, and waste your budget.” Return a real 404 or 410 for pages that are gone, and redirect only where a genuine replacement exists.

A server response time chart

Make crawling cheaper for your server

Since capacity depends on how your server copes, speed helps directly. Faster responses let Google fetch more in the same time. Support for HTTP caching helps too: when a page has not changed, your server can answer with a 304 Not Modified response instead of sending the whole page again, which saves work on both sides.

An XML sitemap with last-modified dates

Keep sitemaps current

A sitemap that lists only the URLs you want indexed, with accurate last-modified dates, tells Google where the new and changed content is. A sitemap full of redirected, blocked or deleted URLs does the opposite. Regenerate it automatically rather than by hand, so it never drifts out of date.

How to tell whether you have a problem

The clearest signal is the third item on Google’s list: a large and growing number of URLs in “Discovered – currently not indexed”. Discovered means Google knows the URL exists but has not crawled it yet. On a big site with a long backlog, that can be a crawl budget symptom.

A large online shop category tree

Be careful not to confuse it with “Crawled – currently not indexed”, which is a different problem. There, Google fetched the page and chose not to index it, usually for reasons of quality or duplication. More crawling will not change that decision. Our guide to crawled, currently not indexed covers what does.

For a direct view of what Googlebot actually requests, server logs are the best evidence. They show which URLs are crawled most and least, and how much of the crawl goes to parameter URLs and errors rather than pages you care about. Our guide to log file analysis for SEO covers how to read them.

The limitation is that Google does not publish a number for any site’s crawl budget, and its thresholds are guidance rather than hard cut-offs. A site of a few thousand pages with a badly built faceted navigation can behave like a much larger one, while a site with millions of clean, well-linked pages may never notice a constraint. Use the thresholds to decide whether to look, then let your logs and the indexing reports tell you whether there is anything to find.

If you run a small site and pages are not being indexed, start somewhere else: internal links, content quality, canonical tags and whether the pages are worth indexing at all. Those explain far more missing pages than crawl budget ever will.

Frequently asked questions

What is crawl budget in SEO?

The set of URLs Google can and wants to crawl on a site, determined by the crawl capacity limit and crawl demand.

Does my small site need to worry about crawl budget?

Usually not. Google’s guide is aimed at sites with 1 million+ pages, 10,000+ pages changing daily, or many URLs stuck in “Discovered – currently not indexed”.

Should I use noindex or robots.txt to save crawl budget?

Google recommends robots.txt for URLs you do not want crawled. A noindexed page still has to be crawled for the tag to be seen.

How do I increase Google’s crawl budget?

Make the server faster, support 304 responses, remove wasted URLs and soft 404s, and improve page quality, which raises crawl demand.

The takeaway Crawl budget matters for large or fast-changing sites. For everyone else, fix quality and internal linking first. If you do qualify, cut wasted URLs before anything else.