Here Designs

Keyword Clustering for Better Content Structure: A Practical Guide

Keyword clustering isn't about sorting keywords into buckets—it's about deciding who owns what before Google decides for you. Here's how to handle the messy edge cases nobody talks about.

Keyword Clustering for Better Content Structure: A Practical Guide

You type "keyword clustering" into Google and you get four thousand results telling you to group keywords by intent. Great. Nobody tells you what happens when a keyword belongs to two clusters at once, or what to do the day your beautiful pillar page starts ranking for its own children.

I ran into that problem myself on a B2B site I look after. One page about invoice software kept outranking every article we wrote about "invoice templates," "invoice format," and "free invoice generator." Six months of content work, and one old page was eating all of it. That's what keyword clustering is really about — not sorting keywords into buckets, but deciding who owns what before Google decides for you.

Key Takeaways

  • Keyword clustering groups queries by search intent, not by string similarity. Two keywords can look identical and belong in different clusters.
  • A cluster is only useful if it maps to one URL. If two pages target the same cluster, you have cannibalization, not structure.
  • SERP overlap is the most reliable signal: if the same five pages rank for two queries, those queries belong together.
  • The hard part isn't building clusters. It's the ambiguous keywords, and you need a rule for them before you start.
  • Cluster structure should show up in your URL architecture, your internal links, and your breadcrumbs — not just in a spreadsheet.
  • Review clusters 90 days after publishing. If two pages keep swapping positions for the same query, merge them.

What keyword clustering actually means for content structure

A keyword cluster is a set of queries that share the same intent and should be answered by a single page. That's it. No pyramid diagrams needed.

The confusion starts when people treat clustering as a research exercise. It isn't. It's an architecture decision. You are deciding which URL gets to be the answer for which group of questions, and everything else downstream — internal links, headings, breadcrumbs, sitemap — flows from that decision.

Intent first, not string first

Take these three queries:

  • "how to invoice a client"
  • "invoice template"
  • "what should an invoice include"

All three contain "invoice." None of them belong together. The first is a process question, the second is a download, the third is a definition. If you build one page for all three, you'll write a Frankenstein article that ranks for nothing because it satisfies no one completely.

Now these two:

  • "best invoicing software for freelancers"
  • "invoicing tools for small business"

Different words, nearly identical intent. Same cluster. Same page.

That's the whole game. Semantic similarity is a trap — it groups words, and you need to group needs.

Why this matters more now than it did a few years ago

Search engines stopped parsing strings as strings a while ago. They match a query to a body of content that satisfies the underlying need, which means a page written for one narrow phrase can absolutely rank for a dozen related ones. That's good news and bad news.

Good: one well-built page can capture a whole cluster's worth of traffic.

Bad: one poorly-built page can steal that traffic from the page you actually wanted to rank.

Which brings us to the part nobody advertises.

Keyword cannibalization: the real reason you need clusters

I've watched a site lose roughly 30% of its organic impressions on a topic after a content team "improved coverage" by publishing four new articles on the same subject. The pages competed. Google picked one, then changed its mind the following week, then changed back.

Keyword cannibalization: the real reason you need clusters

Rankings oscillated for two months. The client thought they'd been hit by an update. They hadn't. They'd cannibalized themselves.

How to spot it before it costs you

Open your Search Console and pull queries where two or more of your URLs appear in the top 20. If you do this once a quarter you'll find it fast. The pattern looks like this:

  • URL A ranks position 4 one week, position 11 the next
  • URL B does the opposite, in the same week
  • Neither page ever stabilizes
  • Click-through rate on both is weaker than the position suggests

That last point is the tell people miss. When two of your pages alternate, users see a result they didn't expect, click, bounce, and Google reads the bounce as a relevance problem.

The ambiguous keyword problem

Here's where most clustering tutorials wave their hands. What do you do with a keyword that genuinely fits two clusters?

My rule, and I'll defend it: assign it to the cluster where it's most commercially valuable, and mention it once in the other cluster. You get a supporting mention, not a second competing page. The keyword gets coverage, but only one URL is built to win it.

I made the opposite mistake early on. I built a dedicated page for every ambiguous term, thinking I was being thorough. I ended up with 14 pages that were all roughly about the same thing, and none of them ranked. Consolidating them into three pages recovered more traffic in six weeks than the previous year of publishing had generated.

How to build a cluster map, step by step

This is the process I use now. It's not elegant, but it holds up.

How to build a cluster map, step by step

Step 1: pull your queries, not your keyword list

Start with queries your site already appears for, not aspirational keywords from a research tool. You want the raw list — hundreds or thousands of rows, including the ugly ones with two impressions. Those reveal intent patterns that head terms hide.

Step 2: group by SERP overlap, not by keyword tool grouping

For each query, look at which pages rank on page one. If two queries share three or more of the same ranking URLs, they belong in the same cluster. This is more reliable than any semantic similarity score because it reflects what the search engine itself considers equivalent.

Tools that do this for you: Keyword Insights is the one I've used most, and clustering tools built on top of DataForSEO or similar APIs let you set the overlap threshold yourself. Free keyword clustering tools exist too, and they're fine for a few hundred keywords — they usually fall apart past that, and they almost never let you tune the threshold.

Step 3: name the cluster by intent, not by keyword

"Invoice software comparison" is a cluster name. "Invoice" is not. The name should describe what the searcher wants, because that name will become your page brief.

Step 4: map one cluster to one URL

This is where the structure actually gets designed. Build a table like this:

Cluster Primary intent Target URL Page type
Invoice software comparison Commercial /invoicing/software/ Pillar
Invoice template downloads Transactional /invoicing/templates/ Pillar
How to write an invoice Informational /invoicing/how-to-write/ Spoke article
Late payment rules Informational /invoicing/late-payments/ Spoke article
Invoice numbering best practice Informational /invoicing/numbering/ Spoke article

Note that the pillar isn't always the broadest term. Sometimes a commercial page deserves to be the hub because that's where the money is, and the informational pages link up to it.

Bringing clusters into your actual site structure

A cluster map in a spreadsheet changes nothing. It has to show up in your URLs, your navigation, and your internal links, or it's just a research artifact.

Bringing clusters into your actual site structure

URL and breadcrumb mapping

Keep cluster URLs siblings under a shared parent. If "invoice templates" and "invoice software" live at /invoicing/templates/ and /invoicing/software/, the breadcrumb tells both users and crawlers that these belong to the same topic. The parent page becomes a natural hub.

One caution: don't nest so deep that a cluster sits five levels down. Two levels below the root is usually enough. I've seen sites bury a pillar under /resources/guides/finance/invoicing/ and then wonder why it never ranks.

Internal linking rules that hold up

Spokes link up to the pillar. The pillar links down to every spoke. Spokes cross-link to each other only when there's a genuine topical bridge.

What I don't do anymore: link every page to every other page in a cluster. It dilutes the signal and makes the structure meaningless. Pick the links that help a reader move to the next logical question.

How this changes the reading experience

Structure isn't just a crawler concern. If your cluster is mapped properly, a reader landing on "how to write an invoice" can find the template they actually need in one click, from a link that makes sense. That's the point. The SEO benefit follows from the reader benefit, not the other way around.

Measuring whether your clusters work

Most people build clusters and then never check whether they did anything. Don't be most people.

The three metrics that matter

  1. Cluster visibility — the total impressions for every query in the cluster, tracked as a group. If the cluster grows, the structure is pulling.
  2. Position stability — how often your top URL for the cluster changes week to week. Instability means two pages are still fighting.
  3. Internal link equity — whether your pillar page picks up links from its spokes over time. If it doesn't, your linking rules aren't being followed.

When to merge or split a cluster

Roughly 90 days after publishing, review. If two pages in the same cluster keep trading positions for the same query, merge them. If one page ranks for two clearly distinct intents and neither gets served well, split it.

I've done both and merging is almost always the right first move. Splitting feels productive, but it usually creates a new cannibalization problem you'll discover six months later.

Where clustering goes wrong

Three failure modes, all of which I've lived through.

Cluster sprawl. You build 60 clusters for a site that has the resources to maintain 12 well. Half the pages go stale, the internal linking rots, and the structure collapses in on itself. Start with your five most valuable clusters and expand from there.

Tool worship. Any clustering tool gives you a starting point, not an answer. The threshold settings change the output dramatically, and no tool knows your business model. A cluster that looks clean in the interface can be commercially useless, and a messy cluster can be exactly where your buyers are.

Ignoring the pillar. Teams get excited about spoke content because it's easier to write. Eighteen months later there are 40 spoke articles and no pillar, and the whole cluster has no anchor. Build the pillar first, even if it's harder.

Questions I get asked about clustering

How many keywords should go in a cluster?

Enough to cover the intent completely, and no more. Some clusters legitimately need 40 keywords because the topic is broad. Others work with three. The number isn't the point — whether one page can honestly serve all of them is.

Do I need a paid tool to cluster keywords?

For under 500 keywords, no. A spreadsheet and manual SERP checks get you most of the way. Past a few thousand, the manual approach stops being feasible and you'll want a tool that automates SERP overlap analysis. Free keyword clustering tools handle the small stuff but rarely let you control the similarity threshold, which is the setting that determines whether your output is usable.

What's the difference between keyword clustering and topic clusters?

Keyword clustering is the research step. A topic cluster is the content structure you build from it — a pillar page plus supporting pages, connected by internal links. You cluster keywords to design the topic cluster. They're two stages of the same job.

And here's the thing I keep coming back to: the best cluster map I ever built wasn't the most detailed one. It was the one with five clusters, each mapped to a page that already existed and just needed restructuring. Total time invested: about nine hours of actual work. The team had spent three weeks on a 60-cluster plan the year before, and roughly 40 of those pages never got written.

Fewer clusters, mapped properly, beats a beautiful spreadsheet nobody builds from. Every time.

Simone Prescott

Simone Prescott is a technical SEO consultant who helps organizations diagnose and resolve complex search visibility issues. Her expertise spans technical SEO audits, site architecture optimization, Core Web Vitals, and crawl budget management, allowing her to bridge the gap between engineering teams and marketing goals. Known for a clear, collaborative approach, she turns dense technical findings into actionable roadmaps that support sustainable organic growth.

See all articles →

Related articles