Learn · Technical SEO · Beginner

How to check if crawl budget is a problem for your site

Learn what crawl budget means, how to test whether your site is big enough for it to matter, and where to look in Search Console for evidence before changing anything.

By Dean Cruddace · 45 minutes to do · Updated · Last reviewed

What crawl budget is

Google cannot explore every address on the web, so there is a limit to the time and resources it spends crawling any one site. Google calls the allocation a site's crawl budget. It is set by two things: how much crawling your server can handle (the crawl capacity limit), and how much Google wants to crawl (crawl demand). Put together, it is the set of addresses Google can and wants to crawl.

This guide is a short procedure for finding out whether this applies to you at all, and what to look at if it does.

Why crawl budget matters for only some sites

Crawl budget is often blamed when a site underperforms, but Google's guide says it is advanced guidance meant for particular sites. If your site does not have a large number of pages that change rapidly, or its pages seem to be crawled the same day they are published, Google says you do not need it. Checking first stops you spending time on the wrong problem. See Google's crawl budget guide.

What you need to check if crawl budget is a problem for your site

  • Access to Google Search Console, with a Domain property or a root-level URL-prefix property (the Crawl Stats report needs one)
  • A rough count of how many unique pages your site has
  • An idea of how often those pages change
  • A web browser

Check if crawl budget is a problem for your site, step by step

  1. 1

    Compare your site with Google's description

    Google says its guide is intended mainly for three kinds of site: large sites (1 million or more unique pages) whose content changes about weekly; medium or larger sites (10,000 or more unique pages) with very rapidly changing content, meaning daily; and sites with a large portion of their addresses shown in Search Console as Discovered - currently not indexed. Google says these numbers are rough estimates, not exact thresholds.

    You will know it worked when You have written down your page count, how often the content changes, and which of the three descriptions, if any, fits.

  2. 2

    Check how your pages are being indexed

    In Search Console, open the Page indexing report. Google says that for Search, keeping your sitemap up to date and checking this report regularly is adequate for a site outside the groups above. Look at the reasons listed for pages that are not indexed. Google describes Discovered - currently not indexed as pages it found but has not yet crawled, typically because crawling was expected to overload the site and was rescheduled. Read the Page indexing report help.

    You will know it worked when You know roughly how many addresses fall under Discovered - currently not indexed compared with your total.

  3. 3

    Open the Crawl Stats report

    In Search Console choose Settings, then Crawl stats. Google says the report is aimed at advanced users and that a site with fewer than a thousand pages should not need it. Note that the counts show actual requested addresses, every step of a redirect chain counts as a separate request, and some requests may not be counted, so figures can differ slightly from server logs. See the Crawl Stats report help.

    You will know it worked when You can see total crawl requests, average response time, host status and crawl responses for your site.

  4. 4

    Look for signs of a capacity problem

    Google says the crawl capacity limit goes down if your site slows, returns server errors (5xx) or sends rate-limiting signals such as HTTP 429. Check the host status (ideally green) and the crawl responses table for server errors.

    You will know it worked when Host status is green and server errors are rare, or you have a list of the error types to take to whoever runs your server.

  5. 5

    Look for wasted crawling

    Google names the factor you can control most as perceived inventory: if many known addresses are duplicates or unwanted, crawling time is wasted. Click into the example addresses in the Crawl Stats report and look for duplicates, filter or sort addresses created by faceted navigation, removed pages, soft 404 pages and redirect chains. Google's guidance is to consolidate duplicates, block unimportant ones with robots.txt (not noindex, which still gets requested), return 404 or 410 for permanently removed pages, fix soft 404s, avoid long redirect chains and keep sitemaps up to date.

    You will know it worked when You have a short list of the address types that Google is spending requests on but that you do not want crawled.

  6. 6

    Fix what you found, then recheck

    Make the changes the guide describes for each item on your list. For filter addresses, Google's faceted navigation guidance says that if you do not need them indexed, prevent crawling with robots.txt rules. Come back to both reports later. Google does not give a timescale or promise a result.

    You will know it worked when The unwanted address types no longer appear among your example crawled addresses, or you can see exactly what still does.

Common mistakes when you check if crawl budget is a problem for your site

  • Diagnosing crawl budget on a small or stable site. Google says its guide is not needed for sites that are crawled the day they publish.
  • Using noindex to stop crawling of unimportant pages. Google says it still requests the page and wastes crawling time.
  • Using robots.txt to temporarily shift budget to other pages. Google says it will not move the freed budget unless the site is already at its capacity limit.
  • Leaving removed pages blocked instead of returning 404 or 410. Google says blocked addresses stay in the crawl queue much longer.
  • Treating Crawl Stats numbers as exact server logs. Google says some requests might not be counted.

Terms used when you check if crawl budget is a problem for your site

Crawl budget
The set of addresses on your site that Google can and wants to crawl.
Crawl capacity limit
How much crawling your server can handle without being overwhelmed, adjusted by Google up or down.
Crawl demand
How much Google wants to crawl your site, based on factors such as size, popularity and how stale its copy is.
Faceted navigation
Filters and sort options on a listing page that often create many address variations.
Soft 404
A page that tells visitors it does not exist but does not return a proper 404 status.
A large site where pages seem slow to be crawled?

Describe how many addresses the site has, how filters create new ones and what Crawl Stats shows. Large-site diagnostics are part of the investigation work.

Talk to an SEO specialistTechnical SEO: architecture →