Technical SEO · Analysis

Why links built by script can be invisible to crawlers

A menu that works perfectly in a browser can leave a crawler with nothing to follow. Google documents which link formats it can use, and how to check what it actually saw.

By Dean Cruddace · Published · Updated · Read 6 min

Key findings

  • Google says it can generally crawl a link only if it is an anchor element with an href attribute; most other link formats will not be parsed.
  • Links inserted by JavaScript are fine when they use that same markup, but they depend on rendering having worked.
  • The URL Inspection tool shows the raw HTML, the rendered page and the JavaScript console output, so the question can be answered with evidence.

A site can look perfect in a browser and still give a crawler very little to follow. If the destinations are held in scripts rather than in links, the pages behind them may never be discovered through the site itself. It is an expensive fault because nothing looks broken to the people who built it.

This article sets out what Google documents about crawlable links and JavaScript, what that implies, and how to check a real page. Where we are giving our own reading rather than Google’s wording, we say so.

The pattern is familiar. A section is not being crawled or indexed as expected, yet the internal links appear to be in place. The cause is often that the links exist only after a script has run, or are not links at all as far as a crawler is concerned.

Here is a hypothetical example. A category page lists forty products. Each card is a div with a click handler that sends the browser to the product URL. A person sees forty links. A crawler reading the HTML sees forty boxes. Nothing in the markup names a destination in a way Google documents as reliable.

The rule: an anchor element with an href

Google’s documentation on link best practices is direct. It says that, generally, Google can only crawl a link if it is an <a> element with an href attribute, and that most links in other formats will not be parsed and extracted by its crawlers. It adds that Google cannot reliably extract URLs from anchor elements without an href, or from other tags that act as links because of script events.

The same page gives a short list of formats Google can parse, including a plain href, a relative path, and an anchor that has both an href and an onclick handler. In that last case the href carries the destination and the script is an addition.

The documentation also lists formats it does not recommend, though it says Google may still attempt to parse them. They include an anchor with a framework attribute such as routerLink in place of an href, a span with an href, an anchor that only has an onclick handler, and an href that contains a javascript: call rather than an address.

The URL in the href should also resolve into an actual web address a crawler can request. And the JavaScript SEO documentation says single-page applications should use the History API, not URL fragments, to load different content, because Googlebot cannot reliably resolve those URLs.

Our reading: if the destination is not written as an address in an href, you are relying on Google to guess.

Where rendering sits between fetch and index

None of this means JavaScript is a problem in itself. Google’s JavaScript SEO basics says Google runs JavaScript with an evergreen version of Chromium, and that it is fine to inject links into the page this way as long as they follow the crawlable link practices above. Links are also crawlable when inserted dynamically, provided they use the anchor and href markup.

What the page describes is a sequence. Googlebot fetches a URL it is allowed to crawl and parses the response for URLs in the href of HTML links. Pages with a 200 status are then queued for rendering, unless a robots meta tag or header says not to index them. The page may stay in that queue for a few seconds, but it can take longer. When resources allow, a headless Chromium renders the page and runs the JavaScript. Googlebot then parses the rendered HTML for links again and queues what it finds, and it uses the rendered HTML to index the page.

The page does not use the phrase “two waves”; its description is crawling, then a rendering queue, then indexing. For links, a link that is in the server’s HTML can be found on the first parse. A link that only exists after scripts have run has to wait for rendering, and it only exists at all if rendering succeeds.

The page adds that Google Search will not render JavaScript from blocked files or on blocked pages, and that for a non-200 status code rendering might be skipped. It also says server-side or pre-rendering is still a good idea, because not all bots can run JavaScript.

Checking what Google actually saw

Google’s guide to fixing search-related JavaScript problems points to the Rich Results Test and the URL Inspection tool to test how Google crawls and renders a URL, and says they show loaded resources, JavaScript console output and exceptions, the rendered DOM and more.

The URL Inspection tool documentation describes a live test, and says that choosing to view the crawled page shows the raw HTML returned, the HTTP headers, JavaScript console output and the page resources loaded. It also offers a screenshot of how the page is seen. The extra response data is available only for URLs that are on Google, and live inspections have a daily limit per property. The link best practices page adds a specific instruction: if JavaScript inserts your anchor text, use URL Inspection to make sure it is present in the rendered HTML.

A sensible routine, and this is our reading rather than a Google procedure, is to compare two things for the same URL: the links in the raw HTML and the links in the rendered HTML. If a navigation block, a pagination control or a set of related-page links appears only in one of them, you know which question to ask next.

Google’s troubleshooting guide lists behaviour of the rendering service that catches out script-built pages. It says to expect Googlebot to decline permission requests, such as a camera, and not to rely on data persistence: local storage, session storage and cookies are cleared across page loads. Our reading: links that appear only after a stored preference or a login state may never appear for Google.

The guide also says client-side analytics may not give a full or accurate picture of Googlebot activity, and points to the Crawl Stats report in Search Console instead. Our reading is that a script which fails in Googlebot may therefore leave no trace in your analytics.

Rendering is one of the areas we cover within technical SEO, under the question of whether the content is really there when Google looks. Our rendering guide covers JavaScript SEO and rendered versus raw HTML, and our architecture guide covers internal linking, which is where this fault does its damage. Our web development page also states the principle we build to: content in the HTML, with JavaScript kept minimal.

We write findings as tickets that include the fault, the evidence, the fix and the acceptance test. We can work with your developers or make the change ourselves where that is faster.

One of our tools touches the same ground from the outbound side. Link Signals reads outbound links across a site and detects hidden links, cloaked links and script-written links. It states its own limit plainly: links built only while scripts run may be missed.

Written by Dean Cruddace

Founder of Cultured Digital. Working in SEO since 2001, across independent consultancy, in-house and agency roles, with a focus on technical SEO, strategy and development.

About Dean →