Log file analysis for SEO means reading your server’s access logs to see which URLs search engine bots requested, when, and with what response. A site crawler tells you what a bot could reach by following links. Search Console tells you a sampled, delayed version of what Google chose to report. The access log is neither: it is a line written by your own server every time something asked for a page.
That makes it the closest thing to ground truth in technical SEO. It also makes it tedious, which is why most people skip it, and for most small sites skipping it is the right call.
What log file analysis for SEO can tell you
Screaming Frog’s SEO Log File Analyser is a useful map of the questions a log can answer, because the tool is built around them. Its own description lists the main ones, and they hold whichever tool you use.
| Question | What the log shows | Why it matters |
|---|---|---|
| Which URLs were crawled? | Every request from search bots such as Googlebot and Bingbot, and from AI bots | Confirms whether important pages are being visited at all |
| Where does crawl activity go? | The most and least frequently crawled URLs | Shows whether bots spend time on the pages you care about |
| What broke? | 4XX and 5XX responses returned to bots | Errors bots hit that a one-off crawl can miss |
| What are you forgetting? | Orphan URLs: requested in the logs but unknown to you | Old or unlinked pages still consuming crawl activity |
| Is that really Googlebot? | Requests that can be verified against Google | Exposes spoofed requests pretending to be a search bot |
None of these needs an opinion. A URL was requested or it was not; it returned a 200 or a 500. That is what makes logs different from most SEO data, which is estimated, modelled or sampled before you see it.

Crawl frequency is the finding, not the volume
The first thing most people do with a log file is count total bot hits and watch the number go up or down. That is the least useful view. A rising total can mean Google is paying more attention to the site, or that it has found a faceted navigation producing thousands of near-identical URLs.
Group the requests by URL, or by template, and sort them. The pattern you are looking for is a mismatch: money pages crawled rarely while parameter URLs, old tag archives or internal search results are crawled constantly. That mismatch is what our post on crawl budget is about, and the log is the only place you can see it directly rather than infer it from indexing reports.
The least frequently crawled URLs are as interesting as the most. A page that has not been requested for months is a page Google has little reason to revisit, often because it is weakly linked internally or adds nothing the site does not say elsewhere.

Verify Googlebot before you trust the numbers
Any client can send a request with a Googlebot user agent. Scrapers, SEO tools and less polite bots do it routinely, so filtering a log by user agent string alone mixes real search engine activity with impersonators. Screaming Frog’s tool lists verifying Googlebot to expose spoofed requests as one of its core functions, and it is the step that most often changes the conclusion.
Verification works by checking the requesting IP address against Google rather than believing the label. Do it before you draw any conclusion about crawl frequency, or you may spend a week optimising for a scraper.

Errors and orphan URLs: what a crawler cannot see
A site crawler starts from your homepage and follows links, so it only finds what is linked. Bots do not work that way. Google remembers URLs it has seen before, from old sitemaps, removed navigation, external links and redirects that no longer exist, and it keeps requesting them.
That is why logs surface orphan URLs: addresses bots still request that do not appear in your own crawl. Some are harmless. Others are old pages returning errors, or live pages with no internal links pointing to them. Our guide to orphan pages covers what to do with each kind once you have the list.

Server errors deserve the same attention. A 5XX that appears for a few minutes during a deployment will not show up in a crawl you run the next morning, but it is in the log with a timestamp. If bots repeatedly hit errors on the same section, that is a reliability problem first and an SEO problem second.
AI crawlers now belong in the same report
Screaming Frog’s tool explicitly covers AI bots alongside Googlebot and Bingbot, and that is the most practical new reason to read logs. Whether an AI crawler is visiting your site at all, which pages it requests and whether your robots.txt rules are being respected are questions only the log can answer.

Treat this as a separate segment. AI crawlers often behave differently from search bots, and lumping them together hides both patterns. If visibility in AI answers matters to you, our post on ranking in ChatGPT explains why being crawled is the precondition for being cited.
Getting started without a big budget
You need three things: access to the raw access logs, enough days of them to see a pattern, and something to parse them. Your host or CDN may store logs for only a short time, so check retention first and start keeping copies before you need them.
Format is rarely the obstacle. Screaming Frog’s tool reads Apache, W3C Extended and Amazon ELB formats, which between them cover logs from Apache, IIS and NGINX servers. Its free version handles 1,000 log events in a single project, which is enough to learn the method on a small slice of data, though not to analyse a busy site.

The most useful exercise is a comparison. Run a normal crawl of the site, import the logs, and line the two lists up. URLs in the crawl but absent from the logs are pages bots are ignoring. URLs in the logs but absent from the crawl are orphans or leftovers. URLs in both, crawled often and returning 200, are the healthy core.
Here is the limitation. Logs tell you what bots requested, not what they did with it. A URL crawled every day can still be excluded from the index, and a log will never show you a ranking or a reason. Pair it with the indexing reports in Search Console, and treat the log as evidence of attention, not of approval.
The received wisdom is that log analysis is for enterprise sites only. That is half right. If your site has a few hundred pages and they are all indexed, you have better uses for the afternoon. But the moment indexing stalls, a migration goes wrong or you want to know whether AI crawlers can reach you, the log is the first place to look, not the last. More measurement guides sit in our data archive.
Frequently asked questions
What is log file analysis in SEO?
Reading your server’s access logs to see which URLs search and AI bots requested, how often, and what status code each request returned.
Do small sites need log file analysis?
Usually not. It earns its time when indexing stalls, after a migration, or when a large site suspects crawl activity is being wasted.
Why verify Googlebot in log files?
Because any client can claim to be Googlebot in its user agent. Verification separates real Google requests from spoofed ones.
Can log files show AI crawler activity?
Yes. AI bots appear in access logs like any other client, so you can see whether they visit and which pages they request.
The takeaway Logs are the only record of what bots actually did on your site. Verify the bots, sort by URL rather than counting hits, and compare against a crawl to find what is ignored and what is forgotten.

