If you want to know how to find a sitemap for a website, start where the site is most likely to tell you: /robots.txt. Guessing URLs works too, but it is slower and easier to get wrong. A sitemap is the list of pages the owner wants search engines to find, so for competitor research it is a quick view of what a site considers its important content, how it groups that content, and sometimes when it last changed.
How to find a sitemap for a website: start with robots.txt
The sitemap protocol defines a standard way to announce a sitemap’s location. In the words of the sitemaps.org protocol: “You can specify the location of the Sitemap using a robots.txt file. To do this, simply add the following line including the full URL to the sitemap:”
So load https://example.com/robots.txt in your browser and search the page for “Sitemap”. A site can list more than one, and the URL can live anywhere on the domain, not only at the root. Copy the URL exactly as written, including the protocol and any subdomain.
The line is optional. Plenty of sites submit their sitemap through Search Console and never add it to robots.txt, so a missing line tells you nothing yet. Our guide to robots.txt for SEO covers the rest of the file, which is worth a glance while you are there: disallowed folders often reveal sections a site does not want crawled.

Step two: try the common sitemap paths
When robots.txt is silent, try the paths that platforms and plugins create by default. Type each one after the domain in your address bar. A working sitemap returns XML, often styled by a stylesheet into a readable table; a missing one returns a 404 or redirects to the homepage.
Watch for two variations. Some sites serve sitemaps compressed, with a .xml.gz extension, which your browser will download rather than display. Others keep a separate sitemap on each subdomain or language folder, so a sitemap on the main domain may cover only part of the site. If the site runs a blog on a subdomain, check that subdomain’s robots.txt as well.
| Where to look | What you will find | Notes |
|---|---|---|
/robots.txt | A Sitemap: line with the full URL | The method the sitemaps.org protocol describes; optional for owners |
/sitemap.xml | A single sitemap or an index file | The most common default location |
/sitemap_index.xml | A sitemap index pointing to child sitemaps | Common where a plugin splits sitemaps by content type |
/wp-sitemap.xml | The sitemap WordPress core generates | Usually replaced or redirected when an SEO plugin takes over |
| CMS or plugin settings | The configured sitemap URL | Only for sites you manage |
| Search Console Sitemaps report | Submitted sitemaps and their status | Only for properties you own |
If none of the paths respond, a site: search combined with inurl:sitemap occasionally surfaces a sitemap that sits somewhere unexpected. Our guide to Google search operators explains how to combine them. Treat a miss there as inconclusive, since search results are not a complete record.

Reading a sitemap index file
On larger sites the first file you open will often not list pages at all. It lists other sitemaps. That is a sitemap index, and it exists because of the protocol’s limits: a single sitemap file can hold up to 50,000 URLs and be no larger than 50MB, so bigger sites split their URLs across several files and point to them from one index.
The way a site splits its index is useful in itself. Child sitemaps are usually named after content types or sections: posts, pages, products, categories, authors. That naming is a free outline of how the site is organised, and of which sections the owner wants indexed. Open each child file to see the page URLs, and look for a lastmod date beside each one if the site provides it.

A sitemap index can also point to sitemaps for images, videos or news. Those are worth noting for competitor research: a site that maintains a news sitemap is publishing for Google News, and one with a video sitemap is investing in video search.
For research, the counts matter more than the individual URLs. Paste each child sitemap’s URLs into a spreadsheet and you can see how many products, posts or category pages a competitor publishes, which sections it is growing, and which it has let go stale. Repeat the check a few months later and the difference shows where the site is investing. It takes minutes and needs no paid tool.
Your own site: CMS settings and Search Console
For a site you manage, you do not need to guess. Your CMS or SEO plugin settings show the sitemap URL it generates, and usually let you choose which content types it includes. WordPress core produces /wp-sitemap.xml on its own; most SEO plugins replace it with their own index.

Search Console’s Sitemaps report lists every sitemap submitted for the property, when Google last read it and whether it hit errors. If the report shows a failure, our guide to fixing “Sitemap could not be read” walks through the usual causes. For what belongs in the file and what to leave out, see XML sitemaps for SEO.

The limitation: a sitemap is the owner’s list, not the whole site
A sitemap shows what the owner wants search engines to index. It is not a complete inventory, and it is not proof that Google indexed anything in it. Sites forget to add sections, leave out old content on purpose, or keep stale URLs that now redirect. Some sites, often small ones, have no sitemap at all and rely on internal links for discovery.
Treat lastmod with the same caution. Some systems update it on every rebuild, whether or not the page changed. Use it as a hint, then check the page itself. The same goes for comparing sitemap sizes between competitors: a larger sitemap means more URLs the owner submitted, not more pages that rank or earn. Thin tag archives, filtered product listings and paginated pages can inflate a count without adding anything of value, so look at what the URLs are before drawing conclusions.

If you need every page rather than the owner’s chosen list, combine the sitemap with a crawl and the other sources in our guide to finding all pages on a website. The gap between the two lists is often the most interesting part: pages the crawler finds but the sitemap omits, and sitemap URLs nothing links to. For more research methods like this, browse the strategy archive.
Frequently asked questions
How do I find a website’s sitemap?
Open the site’s /robots.txt and look for a line starting “Sitemap:”. If there is none, try /sitemap.xml, /sitemap_index.xml and, on WordPress, /wp-sitemap.xml.
What is a sitemap index file?
A file that lists other sitemaps instead of pages. Sites use one because a single sitemap can hold up to 50,000 URLs and 50MB.
Does every website have a sitemap?
No. A sitemap is optional, and some sites, especially small ones, have none and rely on internal links for discovery.
Does a sitemap list every page on a site?
Not necessarily. It lists the pages the owner wants indexed, which can leave out sections or include outdated URLs.
The takeaway Check robots.txt first, then the standard paths. Read the sitemap index as an outline of the site, and remember it is the owner’s list, not a full inventory.

