A sitemap is one of the few files on a site that exists purely to talk to search engines, which makes it easy to build once and never look at again. That is the usual route to a file that quietly lists pages the site no longer wants found. This article sets out what Google says a sitemap is for, what belongs in one, the limits and fields that matter, and the ways a sitemap stops being honest. Where we go beyond what Google states, we say so.
What a sitemap is for, and what it is not
Google’s page Learn about sitemaps describes a sitemap as a file where you provide information about the pages, videos and other files on your site, and the relationships between them. Search engines read it to crawl a site more efficiently. It tells them which pages and files you think are important, and can carry details such as when a page was last updated and its alternate language versions.
The same page is explicit about the limit. A sitemap helps search engines discover URLs, but it does not guarantee that everything in it will be crawled and indexed. Google’s guidance on submitting adds that submitting a sitemap is merely a hint: it does not guarantee that Google will download the sitemap or use it for crawling the site’s URLs.
Google also says a sitemap is not always needed. It may help if a site is large, new with few external links, or heavy in video, image or news content. It may be unnecessary if the site has about 500 pages or fewer that matter for search, is comprehensively linked internally, and has little rich media. In our reading, that makes the sitemap a safety net and a statement of intent, not a repair for weak internal linking.
What belongs in the file, and what the limits are
When you create a sitemap, Google says, you are telling search engines which URLs you prefer to show in results, and those are the canonical URLs. If the same content is reachable at several addresses, choose the preferred one and list that instead of every variant. Google generally shows canonical URLs in results, and says sitemaps are one way to influence that choice. Our article on how Google picks a canonical URL covers the other signals, where Google’s documentation treats sitemap inclusion as a weak one.
The build page lists practical requirements:
- URLs must be fully qualified and absolute, and Google will try to crawl them exactly as listed.
- The file must be UTF-8 encoded.
- A sitemap posted at the site root can affect all files on the site, which is where Google recommends posting it. Elsewhere, it affects only descendants of its parent directory, unless it is submitted through Search Console.
Google supports XML, RSS or Atom feeds and plain text files, and says it has no preference between them. A text sitemap can list only HTML and other indexable pages, one URL per line.
On size, every format limits a single sitemap to 50MB uncompressed or 50,000 URLs. Beyond that you must split the file, and Google’s page on sitemap index files explains the tidy way to do it. An index file lists other sitemaps and is submitted as a single file. It can hold up to 50,000 sitemap locations, its referenced sitemaps must sit on the same site (unless cross-site submission is set up) and in the same directory or lower, and up to 500 index files can be submitted per site in Search Console. The Search Console help adds that an uncompressed size over 50MB is a reportable error and that compression errors are reported too, which implies compressed files are accepted. The sitemaps protocol says files may be compressed with gzip, but the limit applies to the uncompressed size.
lastmod, priority and changefreq
Google’s wording on the optional XML fields is precise. It ignores the priority and changefreq values. It uses lastmod if the value is consistently and verifiably accurate, for example by comparing it with the page’s last modification. The value should reflect the last significant update to the page: Google gives changes to the main content, structured data or links as examples, and a changed copyright date as one that is not significant.
The practical consequence is that lastmod is the only field where honesty has a direct payoff, and also where dishonesty is easiest. A generator that writes the build date into every entry breaks the condition Google sets. We would treat that as a defect to find in any audit, though Google does not say what happens to a site whose lastmod values are unreliable beyond not using them.
Telling Google where the sitemap is
Google lists three ways to make a sitemap available: submit it in Search Console using the Sitemaps report, use the Search Console API, or add a line to robots.txt in the form Sitemap: https://example.com/my_sitemap.xml. The robots.txt line can appear anywhere in the file, can be repeated, and is picked up the next time Google crawls robots.txt.
The Sitemaps report adds useful behaviour. Submitting means telling Google where the file is, not uploading it. The report shows when Google last read the file and what status it returned, such as Success, Has errors or Couldn’t fetch. And the report lists only sitemaps submitted through it or the API, not those discovered through robots.txt, though you can submit a discovered one to track it. Owner permission on the property is needed to submit, otherwise robots.txt is the alternative.
How a sitemap stops telling the truth
Google’s documentation names several failures directly. A sitemap blocked by robots.txt cannot be fetched, because Google respects robots.txt when fetching sitemaps. URLs that cannot be crawled are reported as not accessible. Redirects are called out specifically: Google suggests replacing redirect URLs in sitemaps with the URLs that should actually be crawled. URLs on a different domain or at a higher directory level than the sitemap are reported as not allowed.
Other failures follow from the stated purpose. Google says to list the URLs you want to see in search results. In our reading, that means a sitemap should not contain URLs that are noindexed, return errors, or are non-canonical duplicates; Google’s pages do not enumerate those cases, but each contradicts the file’s own instruction.
Our reading of the routine is simple. Compare the sitemap with what the site actually serves, and compare the count of submitted URLs with the count that are indexed. Google documents that the Page indexing report can be filtered by sitemap, and that the Sitemaps report’s discovered-pages figure counts URLs parsed, with no guarantee that any were crawled or indexed. A large gap between the two is a question to investigate, not a verdict.
Where this fits at Cultured Digital: the sitemap as evidence
Sitemaps and robots.txt sit within the crawling and indexing part of our technical SEO work, alongside canonicals and duplication and log-file analysis. When the subject is a whole site’s outbound links, the starting point matters too: Link Signals, our own tool, scans a site starting from its sitemap, including sitemap indexes and compressed files.
For sites where crawling capacity is the real question, our article on when crawl budget matters sets out where Google says it applies.