How to Audit a Topic Cluster and Decide What Each URL Needs

A topic cluster can look tidy in a diagram and still fail on the live site. Two spokes may answer the same question. A valuable page may sit three clicks away behind a vague “Learn more” link. The hub may list every article yet give readers no reason to choose one. Meanwhile, an old page can keep attracting impressions for advice that no longer deserves to be followed.

topic-cluster audit: a closed site folder, a tablet showing an abstract node chart, and pruning shears arranged left to right, magnifying glass, chain links, desk tray, pencil cup

A useful topic-cluster audit resolves those problems at URL level. It compares what each page is meant to do with the queries it appears for, the answer it actually gives, the internal paths that lead to and from it, its canonical state, and the parts most likely to have gone stale. The result is not a score. It is one bounded decision for every URL: retain, consolidate, relink, refresh, create, or retire.

That distinction matters because the same symptom can point to different repairs. Query overlap may reveal duplicate intent, or it may show that a broad hub and a specialist spoke both deserve visibility. Low traffic may reflect weak discovery, or it may be entirely reasonable for a narrow but commercially important page. Audit the cause before prescribing the change.

Build a ledger that makes vague overlap visible

Begin by fixing the boundary of the cluster. Name the hub, the audience, the market or language, the topic the cluster owns, the neighboring topics it does not own, and the period you will use for search-performance comparisons. Without that boundary, almost any related URL can be pulled into the exercise, and the audit turns into a site-wide content inventory.

Build the candidate list from several views of the site: the hub and its linked pages, an internal crawl, the XML sitemap, CMS categories or tags, and Google Search Console query-page data. Each view has a blind spot. A sitemap can include a page that no reader can reach through an internal link; a CMS tag can group pages that share a label but not a reader task; a crawl can miss a page blocked by the path used; and Search Console does not show every query.

Google’s documentation says the Performance report can be grouped and filtered by query and page, but it also omits anonymized queries and truncates rows because of internal limits. Most performance data is credited to the canonical URL rather than a duplicate. Treat the report as evidence of how URLs surfaced, not as a complete map of demand or proof that two pages are interchangeable.

One row per canonical candidate is enough, provided the row forces a clear comparison:

FieldWhat to captureWhat it changes
URL stateLive URL, status code, declared canonical, Google-selected canonical when available, redirect target, sitemap presenceReveals whether the page being evaluated is the page search systems are consolidating around
Page taskOne reader, one unresolved question, and the action or understanding the page should enableMakes duplicate intent testable without relying on shared keywords
ScopeIncluded and excluded questions, market, product version, and audience levelSeparates legitimate specialization from accidental repetition
Query-page evidenceRelevant queries, impressions, clicks, position context, filters, country, device, and date rangeShows observed overlap while preserving the conditions under which it appeared
Inbound pathsLinking URLs, placement, anchor text, and whether the link is crawlableDistinguishes a discoverable spoke from an orphan or a buried destination
Outbound handoffsThe next reader task, destination URL, anchor, and surrounding explanationShows whether the page helps a reader continue rather than merely accumulating links
Content comparisonDirect answer, distinctive evidence, examples, and material shared with other URLsExposes duplicated answers and unique value that should survive a merger
Freshness and obligationsChange-sensitive claims, source dates, product or policy dependencies, conversions, backlinks, legal needs, and review responsibilityPrevents a content decision from ignoring business or maintenance consequences
Proposed remedyRetain, consolidate, relink, refresh, create, or retire, with rationale and dependencyConverts diagnosis into an implementable page-level choice

The page-task field does most of the intellectual work. Write it after reading the page, not by copying its target keyword. A practical form is: “For this reader, resolve this uncertainty so they can take this action, while excluding this neighboring task.” Two pages about “enterprise data governance” may be distinct if one helps a buyer compare platforms and the other helps an administrator configure access. Two pages with different titles may still duplicate each other if both stop at the same answer for the same reader.

Work through the ledger in a fixed order so that later decisions inherit the right evidence:

  1. Freeze the cluster boundary. State the hub, audience, topic, exclusions, market, and comparison period. If stakeholders cannot agree on that sentence, resolve the boundary before evaluating individual pages.

  2. Collect canonical candidates. Reconcile the hub, crawl, sitemap, CMS, and Search Console lists. Keep alternate URLs attached to the candidate they may duplicate rather than counting each variation as an independent editorial page.

  3. Declare each page’s job. Read the title, lead, headings, answer, and calls to action. Write the intended reader task, then note where the live content does something different. That mismatch may call for a refresh even when no other page overlaps.

  4. Compare query-page evidence under stable filters. Use the same property, search type, country, device treatment, and date range for pages being compared. Flag sustained overlap for inspection, but preserve the limits of anonymized, truncated, and canonically aggregated data.

  5. Trace internal paths and URL signals. Check incoming and outgoing links in rendered pages and HTML, their anchors and context, status codes, canonicals, redirects, and sitemap entries. A link count alone cannot show whether a reader understands where the link goes.

  6. Inspect answers, evidence, and freshness. Put overlapping pages side by side. Mark genuinely distinct reasoning or proof, repeated passages, change-sensitive instructions, broken destinations, and claims whose source no longer supports them.

  7. Choose one remedy per URL. Add the rationale, the surviving destination where applicable, implementation dependencies, and the observation that would trigger review or rollback. “Improve” is not a remedy because it does not say what changes.

  8. Implement, recrawl, and remeasure. Prepare destination content before removing access to a source URL. After the site and reporting systems have had time to reflect the work, compare canonical state, internal paths, query-page distribution, reader behavior, and the business outcome that justified the page.

Diagnose the live cluster before choosing the fix

The ledger becomes useful when it eliminates plausible but wrong explanations. A page with few entrances may be orphaned, but it may also be an intentionally narrow destination linked from the only place that matters. Two pages may rank for the same phrase, but one may answer “what is it?” while the other answers “which option fits?” The diagnosis rests on the combination of page job, delivered answer, query evidence, link path, URL state, freshness, and business need.

Duplicate intent or a legitimate division of labor?

Start with the direct answers, not the keyword lists. Duplicate intent is likely when two URLs address the same audience at the same decision stage, resolve the same uncertainty, cover substantially the same boundaries, and leave the reader with the same next action. Repeated headings, interchangeable introductions, and links whose anchors cannot explain the difference strengthen the case. So does persistent appearance of both pages for the same relevant queries—but only as supporting evidence.

Do not merge pages merely because they share vocabulary or appear for the same query. A hub may orient the reader across a field while a spoke handles one consequential choice in depth. Search Console’s omissions and canonical aggregation also mean the visible query-page table cannot prove that the underlying intent is identical.

If the jobs are distinct but the pages look interchangeable, sharpen their boundaries. Give each a direct answer that belongs to its stage, remove material that drifts into the neighboring job, and rewrite internal-link context so the choice is legible. If the jobs and answers truly duplicate one another, select the survivor by its fitness for the task: completeness, accuracy, distinctive evidence, external and internal link value, conversion role, URL obligations, and maintainability all matter. Traffic alone can preserve the weaker answer.

Google’s SEO Starter Guide says duplicate content is not inherently a spam violation, but it can confuse users and consume crawling resources. That is a useful boundary. Consolidation is a repair for redundant destinations, not a penalty response and not a general way to make a cluster look smaller.

Orphaned spoke or weak handoff?

An orphaned spoke lacks a crawlable internal path from relevant site content. Confirm the condition in the live implementation. Google generally expects a crawlable link to be an HTML <a> element with an href; script-driven elements that merely behave like links may not provide the same reliable path. Its link guidance also recommends descriptive, concise anchor text and links placed in context.

An internal crawl is the starting clue, not the entire test. Inspect the hub, adjacent spokes, navigation components, and rendered HTML. Check whether the destination resolves, whether the canonical points elsewhere, and whether a template or JavaScript behavior changes what the crawler sees. A URL printed as text, included only in a report, or present in a sitemap does not create a reader’s route through the site.

Weak handoffs are subtler because a link exists. A hub can contain every spoke and still make selection difficult through generic card labels, an undifferentiated alphabetical list, or links placed after the point where the reader needed them. Read the sentence around each link. It should explain the question the destination answers or the action it enables. “Compare deployment models for regulated workloads” gives a reason to continue; “Read more” makes the reader reconstruct the architecture.

Relink from the smallest set of pages where the destination is a genuine next step. Adding the spoke everywhere dilutes context and can turn an information problem into a navigation problem. The useful test is whether a qualified reader can enter at the hub or a related spoke, recognize the destination’s role, and reach it through a normal crawlable link.

Scope, freshness, and the remedy decision

A scope gap exists when the cluster’s intended reader has a relevant next task and no existing page handles it adequately. An unused keyword is not enough. Look for the task in customer or sales questions, product and support needs, query evidence, and the declared boundary of the cluster. Then ask whether the hub can answer it briefly. A new spoke earns its URL only when the task needs a distinct answer, evidence set, or decision path.

Define that job before commissioning the page: reader, question, direct answer, inclusions, exclusions, evidence needed, inbound path, and intended handoff. Otherwise the “gap” page often becomes a second, weaker version of something already present.

Freshness is equally specific. A stable definition and a product configuration procedure do not age at the same rate. Inspect the claims most exposed to change: feature names, screenshots, policies, prices, version-dependent steps, benchmarks, recommendations, and external links. A recent modified date proves only that the file changed. Refresh when the page’s job remains valuable but its facts, execution, or proof no longer support the answer. Retire when the job itself is no longer useful or cannot be maintained—and only after checking inbound links, user expectations, conversions, contractual or legal duties, and whether a relevant replacement exists.

Use the diagnosis to constrain the remedy:

RemedyChoose it whenDo not proceed until
RetainThe page has a distinct useful job, a working path, and supportable contentIts boundary and next review trigger are explicit
ConsolidateMultiple pages materially duplicate one reader task and one destination can carry the useful materialUnique value is merged, a relevant survivor is chosen, and affected links and URL signals are mapped
RelinkThe page is useful but its inbound path, anchor context, or next-step handoff is weakThe linking placement answers a real reader need and the target URL is canonical and live
RefreshThe job remains valid while facts, examples, proof, or instructions have degradedChange-sensitive claims are rechecked against current authoritative material
CreateA verified in-scope task has no adequate destinationThe new page has a distinct contract, evidence requirement, inbound path, and handoff
RetireThe page has no current useful job, no defensible answer, or no viable maintenance pathObligations, backlinks, conversions, user paths, and relevant replacement destinations have been reviewed

Implementation order protects the reader from half-finished fixes. Complete the surviving or replacement content first. Then update internal links, navigation, and canonical declarations so they agree on the preferred destination. Apply redirects only after the target can genuinely replace the old page. Recrawl the affected paths, inspect important URLs, and annotate the change date in the reporting view. Measure the intended result rather than looking only for a traffic increase: a consolidation may reduce the number of ranking URLs while clarifying which page carries the task.

Make the first pass reversible

Start with one bounded cluster and preserve the pre-change ledger. It gives editors and developers the same object to inspect, and it makes a disputed decision reversible: the old task, evidence, paths, and URL dependencies are still visible. Prepare and verify the surviving destination before any redirect removes the old page as an independent route. Google describes permanent server-side 301 and 308 redirects as signals that the target should replace the old URL in Search, so that implementation belongs after the editorial decision, not before it.

Do not judge the work the morning after release. Google notes that search changes may take from hours to months to appear and suggests waiting at least a few weeks when assessing effects. Recrawl sooner to catch broken paths and conflicting canonicals; evaluate search distribution and business outcomes on a window long enough to reduce day-to-day noise. The first question is still the simplest one: can every retained URL now explain why it exists?

Frequently asked questions

How often should you run a topic-cluster audit?

Use change triggers rather than a universal calendar. Reopen the cluster when a product, policy, market, source, or search demand changes materially; when a migration alters URLs or navigation; when several pages begin surfacing for the same important queries; or when an assigned review trigger fires. Stable pages can remain untouched while a version-dependent spoke is checked immediately after the underlying product changes.

Should overlapping pages use a redirect or a canonical?

A permanent redirect fits a page that should stop being an independent destination and has a relevant replacement. A rel="canonical" signal fits duplicate or very similar URLs that remain accessible while one URL is preferred for Search. Google treats redirects and rel="canonical" as strong canonicalization signals, while sitemap inclusion is weaker, and advises linking internally to the preferred canonical URL. Neither mechanism determines whether two editorial jobs are actually the same.

Is there an ideal number of spokes or internal links?

There is no universal ideal. The cluster needs enough pages to serve distinct, evidenced reader tasks and no extra pages whose only purpose is to fill a diagram. Google likewise states that there is no magical ideal number of links on a page. Judge each link by whether it helps a reader or crawler discover and understand a relevant destination.

Can a sitemap fix an orphaned page?

A sitemap can help search engines discover a URL, but it does not create a contextual internal path for readers and is only a weak canonicalization signal in Google’s guidance. Repair an important orphan with at least one relevant, crawlable internal link whose anchor and surrounding sentence explain the destination. Keep the sitemap entry if the URL belongs there; do not treat it as a substitute for site architecture.

Run your growth team from one screen.

Invite only