Technical SEO · Analysis

Technical causes of a traffic drop: crawling, indexing and serving errors

Google defines a technical cause of a traffic drop as an error that stops it crawling, indexing or serving your pages. This article goes through the errors Google documents, why some bite fast and others slowly, and what the Crawl Stats report, the Page indexing report and URL Inspection can and cannot confirm.

By Dean Cruddace · Published · Updated · Read 6 min

Key findings

  • Google describes technical issues as errors that can prevent crawling, indexing or serving, and separates site-wide faults such as an outage from page-wide ones such as a misplaced noindex tag, which depends on recrawling and so produces a slower drop.
  • A robots.txt file that cannot be fetched is handled differently from one that returns a 4xx: for a 5xx, Google stops crawling the site for the first 12 hours, while a 4xx (other than 429) is treated as if there were no robots.txt file.
  • The three Search Console tools answer different questions: Crawl Stats shows requests and availability, the Page indexing report shows status by URL group, and URL Inspection shows one URL, from the index or from a live test, and Google says it does not test everything needed to appear.

Of the causes Google lists for a fall in search traffic, the technical ones are the easiest to describe and the easiest to misdiagnose. Our overview of diagnosing a traffic drop covers every cause briefly; this article takes one in depth. It assumes you have already ruled out a reporting problem and the calendar, as set out in checking whether the drop is real.

What Google counts as a technical cause

In its guide to debugging drops in Search traffic, Google defines technical issues as errors that can prevent it from crawling, indexing or serving your pages, giving server availability, robots.txt fetching and "page not found" as examples. A site-wide issue, such as the website being down, affects everything. A page-wide issue, such as a misplaced noindex tag, affects the pages carrying it, and because Google has to crawl a page to see the tag, the drop in traffic would be slower.

Our reading of that timing: a cliff edge on one day points toward something that changed for the whole site, or for Google’s ability to reach it, while a slope that steepens over weeks is consistent with a fault that registers only as pages are recrawled. It is a pointer, not proof. Google’s noindex documentation says only that revisiting a page may take months, depending on its importance.

Availability and status codes: what each family does to indexed URLs

Google’s page on how HTTP status codes affect its crawlers says, in summary:

  • 5xx and 429. Crawlers temporarily slow down. For Google Search, indexed URLs are preserved but eventually dropped if the errors persist. Once the server returns 2xx again, Google gradually increases the crawl rate.
  • 4xx (other than 429). Google does not index these URLs, and indexed URLs that start returning one are removed from the index. Google says not to use 401 or 403 to limit crawl rate.
  • 2xx. The content is considered for indexing, though 200 does not guarantee it. If the content is an empty page or an error message, Search Console shows a soft 404.
  • 3xx. Googlebot generally follows up to 10 hops. A 301 is a strong signal that the target should be processed, a 302 a weak one.

The soft 404 is the quiet one: the server reports success while the page says otherwise. Our reading is that a template fault or failed data call can produce it across many URLs at once. Our article on redirects covers the 3xx side, and the crawl budget article covers how server health affects how much Google asks for.

Robots.txt, DNS and network failures

Some errors happen before any page is requested. Google’s robots.txt specification says that a 4xx other than 429 is treated as if no valid robots.txt existed, so no crawl restrictions apply. For a 5xx, Google stops crawling the site for the first 12 hours while it keeps trying; for the next 30 days it uses the last good version if it has one. A file that cannot be fetched because of DNS or networking problems, such as timeouts or reset connections, is treated as a server error.

The Crawl Stats help page lists the same families as fetch outcomes: DNS unresponsive, DNS error, fetch error, page timeout, redirect error and "page could not be reached". Google notes that for the last of these the request never reached your server, so it will not appear in your logs. Our reading: clean logs are not proof that Google had no trouble.

Noindex, canonicals and blocked resources: errors that look like decisions

Here the server is healthy, and something told Google to leave pages out. Google’s noindex documentation says that once Googlebot crawls a page and extracts the tag or header, the page is dropped from results, and that a page blocked by robots.txt never shows the rule at all.

Canonical tags are a signal about which URL is the main one; how Google picks a canonical URL covers how the choice is made. Our reading: a wrong canonical in a shared template is the same kind of fault as a stray noindex, with the same slow, page-by-page effect.

Blocked resources are the third. Google’s robots.txt guide says unimportant resource files can be blocked, but not ones whose absence makes the page harder for Google to understand. A sitemap that lists blocked or noindexed URLs is a symptom worth noticing, as our article on keeping sitemaps honest explains.

Site-wide or page-specific: reading the pattern

Google’s debugging guide sends you to the Pages table in the Performance report to see whether the loss is across the site, a group of pages or one important page. Our reading for technical causes: site-wide and abrupt suggests availability, DNS, a robots.txt change or a redirect fault; a group of pages sharing a template or folder suggests a template-level noindex, canonical or status-code fault; a single page suggests something about that page alone. These are hypotheses to test, not conclusions.

What Crawl Stats, Page indexing and URL Inspection can and cannot show

Google’s guide names Crawl Stats and Page indexing as the places to look for a matching spike in issues. Each has limits its own documentation states.

  • Crawl Stats shows requests, response codes, host status and fetch problems. Google says it is aimed at advanced users, counts each hop of a redirect chain separately, does not count client-side redirects, may differ slightly from server logs, and is available only for root-level properties.
  • Page indexing shows which URLs are indexed or not and why, including soft 404, server error (5xx) and exclusion by noindex. Example lists are capped at 1,000 URLs and may not show every URL in a status, and the report is not for investigating specific pages. Not indexed is not necessarily bad.
  • URL Inspection shows what Google knows about one URL. The indexed result is from the last crawl, not the live page, and the live test cannot predict the canonical. Google says it does not check manual actions or security issues.

Our reading: Crawl Stats says whether Google could reach the site, Page indexing says what became of groups of URLs, and URL Inspection checks a sample. None says why traffic fell; they confirm or rule out a technical explanation.

Where this fits at Cultured Digital: crawling and indexing faults as tickets

Our SEO audit is a focused investigation that finds the cause, then says what to do first, and crawling and indexing is one of the areas it can cover, depending on the site and the problem. The crawling and indexing page, under technical SEO, sets out the crawl and index steps: whether a page is allowed and affordable to crawl, and which version counts. A technical fault is written up as a ticket with the fault, the evidence, the fix and the acceptance test. We do not promise rankings or traffic numbers. If the evidence points elsewhere, see what to check after a Google update, security and spam causes and drops after a site move or release.

Further reading on crawling and indexing errors

Written by Dean Cruddace

Founder of Cultured Digital. Working in SEO since 2001, across independent consultancy, in-house and agency roles, with a focus on technical SEO, strategy and development.

About Dean →