Free to explore: filter winning sites by DR, traffic and niche  ·  Try the live explorer →

Googlebot user agents and how to verify them

Server logs scrolling on a screen

The Googlebot crawler user agent is the first thing people search their logs for, and the least reliable thing to trust. It tells you which version of Googlebot claims to be visiting. It does not tell you whether the claim is true. This guide covers the two crawler types, what they fetch, and how to confirm a request really came from Google.

Two crawlers: smartphone and desktop

Google’s documentation on Googlebot describes two main crawler types. “Googlebot Smartphone: a mobile crawler that simulates a user on a mobile device” and “Googlebot Desktop: a desktop crawler that simulates a user on desktop.” Both identify themselves with the Googlebot token, and both obey robots.txt rules written for user-agent: Googlebot.

You cannot target one of them separately in robots.txt; a rule for Googlebot applies to both. You can tell them apart in logs, because the Googlebot mobile user agent string includes a mobile device description while the Googlebot desktop user agent string does not. Google publishes the full strings in its crawler documentation, and they change as the underlying browser version updates, so match on the Googlebot token rather than the whole string.

A smartphone next to desktop computer

Why the smartphone crawler matters more

Google states that “For most sites Google Search primarily indexes the mobile version of the content.” In practice, most Googlebot requests in your logs will come from the smartphone crawler, and what it sees is what gets indexed. If your mobile pages hide content, load it differently or block resources that desktop pages fetch, the mobile version is the one that wins.

That makes the split more than a curiosity. When you test rendering, test as the smartphone crawler. When a page ranks badly despite looking complete on a desktop, check what a mobile visitor receives. Our post on mobile-first indexing covers the parity problems that cause most of these gaps.

What Googlebot fetches, and how much

Googlebot does not read files of unlimited size. The documentation sets out the limits, and they are worth knowing if you serve very long pages or large documents.

ItemWhat Google’s documentation says
Googlebot SmartphoneA mobile crawler that simulates a user on a mobile device
Googlebot DesktopA desktop crawler that simulates a user on desktop
Which version is indexedFor most sites, primarily the mobile version
Supported file typesThe first 2MB is processed
PDFsThe first 64MB is processed
Crawl frequencyFor most sites, no more than once every few seconds on average

Google notes these limits apply to uncompressed data. For an ordinary HTML page, 2MB is a lot of markup. The limit tends to bite on pages that inline large amounts of data, script or base64 images into the HTML, pushing real content further down the file.

A network engineer at computer

The user agent is often spoofed

Google is blunt about this: “the HTTP user-agent request header used by Googlebot is often spoofed by other crawlers.” Scrapers and SEO tools borrow it because many sites treat Googlebot generously, letting it past rate limits, paywalls or bot protection. Anything that grants access based on the user-agent string alone can be walked through by anyone willing to type it.

That has two consequences. Log analysis that counts every request labelled Googlebot will overstate how often Google crawls you, sometimes badly. And any special treatment you give Googlebot should be tied to a verified identity, not a header. Our guide to log file analysis for SEO covers how to filter fake crawlers out before drawing conclusions.

A security professional at monitor

How to verify Googlebot

Google gives two ways to confirm a request is genuine: run “a reverse DNS lookup on the source IP”, or match the IP against Google’s published IP ranges. The DNS method works in three steps:

  1. Reverse lookup. Look up the hostname for the requesting IP address.
  2. Check the domain. A genuine Googlebot hostname ends in a Google-owned domain such as googlebot.com or google.com.
  3. Forward lookup. Resolve that hostname back to an IP address and confirm it matches the original. This stops anyone who controls their own reverse DNS from faking the result.

The output above is illustrative, using a documentation address. For checking traffic at scale, the published IP ranges are easier: download Google’s list and match against it in your firewall, CDN or log pipeline, refreshing it regularly as the ranges change.

A line chart on laptop screen

Crawl rate and blocking decisions

Google says that “For most sites, Googlebot shouldn’t access your site more than once every few seconds on average”. If you see far more than that from requests calling themselves Googlebot, verify them before assuming Google is hammering your server. The excess is often impostors.

A quick way to check is to take a day of logs, pull out every request whose user agent contains the Googlebot token, and group them by IP address. Verify the busiest addresses first. If the heaviest traffic comes from IPs that fail the reverse DNS check, you have found your load problem, and it is not Google. Rate-limit or block those addresses at the firewall or CDN; doing so has no effect on how Google crawls or ranks you, because Google never sent them. Then rerun the count on verified requests only. That figure is the one to compare against Google’s crawl stats report in Search Console, and it is usually much closer than the raw total.

A warning sign

If verified Googlebot traffic is genuinely heavy, the usual cause is a site generating more URLs than it needs to, through parameters, filters or infinite calendars. Our post on crawl budget covers how to cut that waste. Be careful with blocking: a firewall rule that drops verified Googlebot will remove pages from search, whereas blocking impostors costs nothing. And if your real concern is AI crawlers rather than Google Search, our guide to blocking AI crawlers covers the separate tokens they use.

An engineer analysing logs

One limitation applies to every method here. Verification tells you a request came from Google; it does not tell you which Google crawler made it or why, beyond what the user-agent string claims. Google runs other crawlers and fetchers alongside Googlebot, and some share its infrastructure. Treat a verified IP as proof of origin, and the string as a label to interpret.

The working rule is simple. Match on the Googlebot token to spot candidates, verify by DNS or published ranges before acting on them, and judge rendering from the smartphone crawler’s point of view. Our strategy archive collects the related posts on crawling and indexing.

Frequently asked questions

What is the difference between Googlebot Smartphone and Googlebot Desktop?

Googlebot Smartphone simulates a user on a mobile device; Googlebot Desktop simulates a user on desktop. For most sites, Google primarily indexes the mobile version of the content.

Can I trust the Googlebot user agent in my logs?

No. Google says the user-agent header is often spoofed by other crawlers. Verify with a reverse DNS lookup on the source IP or Google’s published IP ranges.

How much of a page does Googlebot read?

Googlebot processes the first 2MB of a supported file type and the first 64MB of a PDF, measured uncompressed.

How often should Googlebot crawl my site?

Google says that for most sites Googlebot shouldn’t access the site more than once every few seconds on average.

The takeaway Googlebot has smartphone and desktop crawlers, and the smartphone one decides what is indexed for most sites. Never trust the user-agent string alone; verify by IP.