Strategy · Analysis

Site structure is a strategy decision: what Google documents about URLs and links

Google documents the mechanics of URLs, links and hostnames in some detail, and says openly that several structural choices are the business's to make. This article separates what the documentation covers from what it leaves to strategy.

By Dean Cruddace · Published · Updated · Read 6 min

Key findings

  • Google's URL guidance sets requirements for a crawlable structure (no fragments to change content, common parameter encoding) and warns that URLs which fail them will likely be crawled inefficiently.
  • Google says keywords in the domain or URL path alone have hardly any effect on ranking beyond appearing in breadcrumbs, and that the subdomain-or-subdirectory choice should be made on business grounds.
  • Google has no page on information architecture, topic strategy or entity strategy as disciplines; the documentation covers URLs, links and directories, and the rest is judgement.

A site’s structure is a set of decisions that tend to be made once and lived with for years: what each URL looks like, which pages link to which, and which hostname or folder a section sits on. Google documents more of the mechanics than people often expect, and says less about the strategy than they often assume. This article sets out what its documentation covers on URLs, links and hostnames, what it explicitly leaves to the business, and where it has no page at all. Everything attributed to Google comes from pages we read for this article; our own readings are labelled as such.

Where structure appears in Google’s own guidance

The SEO Starter Guide says that when you set up or redo a site, organising it logically can help search engines and users understand how pages relate to the rest of the site. It is candid about the limits. It tells readers not to drop everything and start reorganising, because search engines will likely understand pages as they are now, and says the suggestions help most over the long term, especially on larger sites.

The guide then names three practical areas. Descriptive URLs, because parts of a URL can appear in results as breadcrumbs. Grouping topically similar pages in directories, which it says can affect how Google crawls and indexes a site once it has more than a few thousand URLs, because Google can learn how often content in each directory changes. And links, which it describes as how the vast majority of new pages are found each day.

Descriptive URLs and the requirements behind them

Google’s page on URL structure best practices opens with requirements, not tips. If URLs do not meet them, it says, Google will likely crawl the site inefficiently, in the extreme with very high crawl rates or not at all. The requirements are to follow the URL standard (IETF STD 66, with reserved characters percent encoded), not to use URL fragments to change page content (use the History API instead), and to use the common parameter encoding of an equals sign between key and value and an ampersand between pairs.

The best practices that follow are mostly about being intelligible:

  • Use readable words rather than long ID numbers.
  • Use words in your audience’s language.
  • Separate words with hyphens, not underscores, and do not join them together.
  • Use as few parameters as you can, trimming those that do not change the content.
  • Remember that URLs are case sensitive: Google treats /APPLE and /apple as different URLs, so if your server treats them the same, convert everything to one case.

It helps to hold that alongside another line in the Starter Guide, which says keywords in the domain name or URL path alone have hardly any effect on ranking beyond appearing in breadcrumbs. Our reading, not Google’s wording: descriptive URLs are a clarity decision for people and crawlers, not a ranking lever, and the cost of changing them later is a reason to settle the pattern early.

The same page lists what inflates URL counts: additive filtering, irrelevant parameters such as session IDs and referral tags, infinite calendars and broken relative links. We do not repeat that ground here; our article on when crawl budget matters covers how Google frames the cost of excess URLs.

Google’s link best practices begin with a statement of purpose: Google uses links as a signal when determining the relevancy of pages, and to find new pages to crawl. Generally, it can crawl a link only if it is an <a> element with an href attribute. Links inserted by JavaScript are fine if they use that markup; elements that act as links only through script events are not reliably extracted.

On anchor text, the page says good text is descriptive, reasonably concise and relevant to both the page it sits on and the page it points to. It lists generic examples such as “click here” and “read more” as too generic, suggests reading the anchor text out of context to see whether it makes sense alone, and cautions against chaining links together and against forcing keywords in. If an anchor is empty Google can fall back on the title attribute, and for an image link it uses the alt text.

For internal links specifically, the page says paying attention to anchor text can help people and Google make sense of a site, that every page you care about should have a link from at least one other page, and that there is no magical ideal number of links per page, although if you think it is too many, it probably is. Our reading: navigation and templates repeat across many pages, so a choice made once in a template is effectively a site-wide linking decision.

Hostnames, folders and what Google leaves to the business

On hostnames, Google’s position is mostly that it is your call. The Starter Guide says that for subdomains versus subdirectories you should do whatever makes sense for the business, noting that segmenting by subdirectory may be easier to manage, while partitioning topics into subdomains can suit some sites. It says the top-level domain only matters if you are targeting a specific country’s users, and even then is usually a low-impact signal.

For multi-regional sites, the URL structure page recommends a structure that makes geotargeting easy and shows both a country-specific domain and a country-specific subdirectory on a generic domain as recommended examples. It does not rank one above the other. Our reading: with no verdict from Google, the deciding factors are operational. Domain and folder strategy is part of the international side of our Technical SEO scope for that reason.

Where Google has no page: architecture and topics

Google does not publish a page on information architecture as a discipline, or on topic strategy or entity strategy. What exists is the set of fragments above, plus adjacent guidance. Its helpful content guidance asks whether a site has a primary purpose or focus, and lists among its warning signs producing lots of content on many different topics in the hope some of it performs. That is guidance about content, not a method for structuring a site.

That gap is where strategy sits. Which sections the business needs, which searches each should serve and how the site will grow without becoming tangled are decisions the documentation does not make for you. Anything we say about them is judgement, informed by what Google does document.

Where this fits at Cultured Digital: structure as a decision

We treat structure as both a strategy question and a technical one. On the strategy side, information architecture, topic and entity strategy and internal linking strategy make up the “Shape the site” group in our SEO Strategy scope; our Shape the site guide sets that out. On the technical side, information architecture, internal linking, faceted navigation and parameters and large website analysis sit in the architecture part of Technical SEO, where the work is diagnosis and implementation.

For builds, our Web Development page lists URLs as one of the things decided at the start: logical, stable and human-readable. And on our technical SEO page, the first step of the path from crawl to understanding is whether a page can be found, through internal links, sitemaps and orphan pages. We do not promise rankings or traffic; we set out the opportunity and the evidence, and measure against it.

Further reading on site structure

Written by Dean Cruddace

Founder of Cultured Digital. Working in SEO since 2001, across independent consultancy, in-house and agency roles, with a focus on technical SEO, strategy and development.

About Dean →