Broken-link reports are one of the most routine outputs in SEO, and the advice that follows them is routine too: fix it, redirect it or remove it. That advice is only sound if the report is right. In testing link checks we saw two well-known websites reported as unreachable although both answer normally for people. One sits behind a bot challenge. The other sends a first visit through an address that sets a cookie and then sends you back. Neither was down, and neither was broken. Both were behaving as designed towards a visitor that could not pass a challenge or keep a cookie.
What the status codes actually say
The HTTP specification is precise about the codes that cause most false alarms. A 403 means the server understood the request but refuses to fulfill it
. That is a refusal, not an absence. A 429 is defined in RFC 6585 as the user has sent too many requests in a given amount of time
: it is rate limiting. Neither tells you that a page does not exist, and removing a link or a page on the strength of one is a mistake.
Why this is an SEO problem, three times over
1. Your own site and search engine crawlers
This is the most important one. Google’s documentation says that all 4xx errors except 429 are treated the same by its crawlers: the content is treated as not existing, and a URL that was indexed is removed from the index. A 429 is treated as a signal that the server is overloaded, which makes the crawlers slow down. It also states that 401 and 403 should not be used to limit the crawl rate.
So a bot-protection rule, firewall or rate limiter that answers Googlebot with a 403 or a challenge can cause pages to drop out of the index, and one that answers with 429 or 5xx can slow crawling of the whole site. Protection is good practice; protection that treats search engines as attackers is an indexing problem. If you run a firewall or CDN with bot rules, check what a search engine crawler receives, not just what a browser receives. Google publishes how to verify that a request really comes from Googlebot, so that you can allow the real crawler without allowing impostors.
2. Link audits that delete good citations
If a broken-link report flags a working source as dead, the obvious action is to remove the citation or the partner link. Do that at scale and you have thinned out the references that make your pages credible, because of a classification error. Before acting on a “broken” label, open the link in a browser. If it works there, the report is wrong or incomplete.
3. Outreach that pitches the wrong thing
Broken-link building depends on finding genuinely dead pages. If your prospecting list includes a site that is only refusing automated visitors, your pitch to replace the dead resource will land on a live one.
Cookies and redirects
A pattern that confuses automated checkers is a first visit that is redirected through a URL that sets a cookie and returns to the original address. A browser keeps the cookie and carries on. A client that does not keep cookies loops until it gives up and reports a failure. It is a reasonable thing for a site to do, and a reminder to test your own redirect and consent behaviour with a client that stores no cookies, to see what one of those clients sees.
What a trustworthy report distinguishes
A report that is useful for SEO decisions separates at least four situations, because each needs a different response:
- Not found or gone (404, 410): the destination is missing. Replace or remove the link.
- Moved (3xx): update the link to the final address.
- Couldn’t check (401, 403, 429, a bot challenge, a redirect loop it cannot pass): the destination answered but would not let the checker read it. Verify by hand.
- Unreachable (DNS failure, refused or timed-out connection, certificate problem): the destination really could not be reached, for a stated reason.
Link Signals separates “Couldn’t check” from “Unreachable” in this way. The wording matters because a site owner who reads “unreachable” beside a working link will reasonably conclude the link is dead.
When the challenge page returns a 200
Not every protection layer answers with an error code. Some serve a challenge page with a normal 200 status, because from the server’s point of view the request succeeded and a page was delivered. A status-code report cannot see the difference. A checker or a crawler that receives that page has fetched a page, but not yours. If your own site works this way for search engine crawlers, what they index or process may be the challenge, not your content. The remedy is to look at the content, not only the status: compare the HTML a crawler receives with what a browser shows, which is what the “View crawled page” function in the URL Inspection tool is for.
Allowing the crawlers you want
The aim of bot protection is to stop abusive traffic, not search engines. Sensible practice is to allow verified search engine crawlers, to apply rate limits rather than blanket blocks to unknown traffic, and to review the rules whenever the CDN or firewall configuration changes, because a rule that was safe last quarter can behave differently after an update. After a change, retest with URL Inspection and watch the Crawl Stats report in Search Console. A sudden rise in 403, 429 or 5xx responses there is the early warning.
Rate limits and large sites
Because 429 and 5xx responses make Google’s crawlers slow down, an over-eager rate limit does more than reject a few requests. On a site with a large number of URLs, slower crawling means new and changed pages are discovered later. If your pages are time-sensitive, or you have just migrated, that delay has a direct cost.
A triage guide
- 404 or 410 from the destination: the page is gone. Replace or remove the link.
- 301 or 308: the page has moved. Update the link to the new address.
- 403, 429 or a challenge when checked automatically, but fine in a browser: not broken. Leave the link and verify manually.
- Timeouts, connection refused, certificate errors: test again later and from another network before concluding anything. If it persists, the destination really is down.
- A page that loads but is not what you cited: that is a different problem, covered in the article on expired domains.
Checking by hand
- Open the URL in a normal browser, logged out. If it loads, the destination is alive.
- For your own pages, use the URL Inspection tool in Search Console to see what Google receives.
- Look in your server or CDN logs for 403, 429 and 5xx responses to verified Googlebot requests.
Further reading on false broken links
- IETF RFC 9110: HTTP Semantics, 403 Forbidden
- IETF RFC 6585: Additional HTTP Status Codes (429 Too Many Requests)
- Google Search Central: How HTTP status codes, and network and DNS errors affect Google Search
- Google Search Central: Verifying Googlebot and other Google crawlers
- Cloudflare: Challenges documentation
- Search Console Help: URL Inspection tool
