Classic keyword research produced a spreadsheet: 5,000 strings, each with a volume and a difficulty score, each destined, in theory, for its own optimized page. That workflow made sense when Google matched strings. It stopped making sense when Google started resolving queries to meanings: today a single ranking page routinely covers hundreds or thousands of phrasings of the same underlying need, and building a page per string produces a site full of near-duplicates competing with each other.
Semantic keyword research inverts the process. Instead of asking "which strings have volume?", it asks "which entities and intents exist in this topic, and how do queries group around them?" The output isn't a keyword list, it's a set of clusters, each representing one intent, each mapped to exactly one planned page in a topical architecture.
This guide walks the full workflow: seeding from entities, expanding the query universe, clustering by intent, mapping clusters to a topical map, and prioritizing what to build first.
Why string-based keyword research broke
Three shifts in how search works made the old workflow obsolete:
- Query understanding got semantic. Hummingbird, RankBrain, BERT, and MUM progressively taught Google to resolve queries into meaning, synonyms, rephrasings, and even never-before-seen queries collapse to known intents. "cheap flights to tokyo" and "low cost airfare tokyo japan" are one problem to Google, so they must be one page for you. This is the same machinery covered in how search parses entities.
- Pages rank for query families, not strings. Study any top-ranking page in your tool of choice: the typical page ranking #1 for a head term also ranks for hundreds of related queries it never 'targeted.' Coverage of a meaning earns the whole family.
- One-page-per-keyword now actively hurts. Ten thin pages targeting ten phrasings of one intent split internal links, dilute authority, and cannibalize each other in the SERP. The tactic didn't just stop working, it became a liability.
“Volume tells you how often a string is typed. It tells you nothing about how many pages the underlying need deserves, and that's the question keyword research exists to answer.”
None of this means keywords are dead. Queries are still the observable data, the fossil record of demand. What changed is the unit of analysis: you collect strings in order to discover intents, and you build pages for intents. That reframing is the heart of semantic SEO.
Step 1: Start from seed entities, not seed keywords
A seed keyword is a string you hope has volume. A seed entity is a thing your topic is made of, a concept, tool, method, problem, or audience segment. Entities generate keywords systematically; keywords generate only their own variations. For a site about email marketing, seed entities include: deliverability, list segmentation, automation workflows, subject lines, ESPs (and each major ESP as its own entity), sender reputation, re-engagement campaigns.
Build the seed list from four sources:
- The topic's own structure. How do textbooks, documentation, and Wikipedia decompose it? Category trees and tables of contents are pre-built entity inventories.
- Audience problems. Sales calls, support tickets, community threads, the problems people actually phrase, which rarely match tool-suggested keywords.
- Competitor coverage. Crawl the top 2–3 topical competitors' URL structures. Every hub and category they maintain is an entity they've judged worth owning.
- SERP artifacts. People Also Ask chains, related searches, and the bolded terms in snippets reveal which sub-entities Google itself associates with the topic.
Step 2: Expand each entity into its query universe
Now the tools earn their keep. For each seed entity, collect every query phrasing you can find, this is deliberately a wide net, because the clustering step will compress it. Pull from: keyword tool matching- and related-terms reports, autocomplete and PAA scrapes, your own Search Console queries (the only source showing demand you already partially capture), and competitor ranking keywords for their page covering the same entity.
At this stage, resist two temptations. Don't filter by volume yet, zero-volume queries cluster into intents with real aggregate demand, and tools chronically under-report long-tail and B2B volumes. And don't assign anything to pages yet; that's what clustering is for. A healthy expansion for a mid-sized topic yields 2,000–10,000 raw queries across 30–80 entities.
Step 3: Cluster queries by shared intent
Clustering answers the central question: which of these queries can one page satisfy? The most reliable method is SERP overlap clustering: if two keywords share a meaningful number of URLs in their top 10 results (3–4 is the common threshold), Google is already answering them with the same pages, proof they're one intent. Most modern tools automate this; the judgment calls are in the settings and the review.
| Method | How it groups | Strength | Weakness |
|---|---|---|---|
| Lemma / string matching | Shared words and stems | Fast, free, no API costs | Groups by spelling, not meaning, splits synonyms, merges homonyms |
| SERP overlap | Shared top-10 URLs between queries | Reflects Google's actual intent judgment; the industry default | Costs SERP data; unstable on volatile SERPs |
| Embedding similarity | Semantic distance between query vectors | Catches synonyms with zero string overlap; cheap at scale | Similar meaning doesn't always mean same SERP intent, needs SERP validation |
In practice, use embeddings or lemma grouping for a cheap first pass on huge lists, then confirm page boundaries with SERP overlap. After the automated pass, review clusters manually for the two failure modes: over-merging (informational and commercial queries fused because big publishers rank for both, split them by modifier: "best," "vs," "pricing" signal commercial) and over-splitting (one intent fragmented across volatile SERPs, merge clusters whose top results keep interleaving). Every boundary decision is really an intent decision, which is why clustering and search intent analysis are the same discipline at different zoom levels.
Label each finished cluster with three fields: a primary query (the phrasing you'll use in the title), the intent type, and total cluster volume, the aggregate across all member queries, which matters far more than any single keyword's number.
Step 4: Map clusters onto a topical map, not one page per keyword
A pile of clusters is still just a smarter spreadsheet. The final transformation is architectural: arrange clusters into a hierarchy of hubs and spokes, where each cluster becomes exactly one planned page with a defined slot, and every page's relationships to its neighbors are explicit.
- Group clusters by parent entity. All clusters about deliverability, definition, spam filters, warmup, testing tools, form one topical neighborhood.
- Assign the hub. The broadest, usually most competitive cluster in the neighborhood becomes the hub page; the rest become spokes.
- Define one query family per page. Each cluster's member keywords become that page's title, H2, and FAQ fodder, and no other page may target them. This single rule eliminates cannibalization by design.
- Plan the links before the pages. Spokes link up to hubs, hubs link down to spokes, hubs link to related hubs, the linking plan falls directly out of the map, per our internal linking strategy.
- Mark coverage gaps. Entities with no cluster (because tools showed no volume) still get planned pages if the topic's completeness requires them, completeness itself is a ranking asset under topical authority.
Step 5: Prioritize clusters and put the research to work
You can't build the whole map at once, so sequence it. Score each cluster on four factors:
- Business value, how directly the cluster's intent connects to revenue. A 200-volume transactional cluster beats a 20,000-volume trivia cluster.
- Cluster opportunity, aggregate volume weighted by realistic difficulty. Judge difficulty at the cluster level by who actually ranks: weak or mismatched incumbents matter more than any difficulty score.
- Cluster completeness leverage, clusters that complete a neighborhood get a bonus, because finishing a hub-and-spoke set lifts the whole set's rankings, not just the new page's.
- Existing equity, clusters where you already half-rank (check Search Console) convert fastest into wins.
Then execute cluster-by-cluster, not page-by-page: complete one neighborhood, hub, spokes, and internal links, before starting the next. Each cluster's data flows straight into its pages' content briefs: the query family defines the title and headings, the member questions become FAQ sections, and the map defines the links. Revisit the research twice a year, new queries appear, intents drift, and the map should absorb both without changing its skeleton.
Key takeaways
- The unit of keyword research has changed: you're no longer collecting strings to target one-by-one, you're discovering the entities and intents in a topic and grouping queries that one page can satisfy.
- Start from seed entities, your core topic's sub-concepts, attributes, and audience problems, not from whatever autocomplete happens to suggest.
- Cluster by SERP overlap: if two keywords share several top-10 URLs, Google considers them the same intent and they belong on the same page.
- The deliverable isn't a spreadsheet of keywords, it's clusters assigned to slots in a topical map, each slot becoming exactly one page with one query family.
- Prioritize clusters by business value and cluster-level opportunity, not by any single keyword's volume; long-tail queries within a cluster usually deliver most of the traffic.
Frequently asked questions
Is keyword research still relevant for semantic SEO?
Yes, queries remain the observable record of what people want, and no entity model replaces that data. What changed is the unit of analysis: you collect keywords to discover intents and entities, cluster the strings that share an intent, and build one page per cluster instead of one page per keyword.
What is keyword clustering and how does it work?
Clustering groups keywords that a single page can rank for. The most reliable method is SERP overlap: if two queries share three or more URLs in their top 10 results, Google is already treating them as the same intent, so they belong together. Embedding-based and lemma-based grouping are useful cheap first passes, but SERP overlap should confirm final page boundaries.
How many keywords should one page target?
One cluster, which might be five queries or five hundred. The bound is intent, not count: every query one page targets should be satisfiable by the same content in the same format. If satisfying two queries would require different page types or different depth, they're different clusters and different pages.
Should I create pages for zero-volume keywords?
Often, yes. Tools under-report long-tail, B2B, and new-category queries badly, and zero-volume phrasings frequently cluster into intents with real aggregate demand. Beyond traffic, covering low-volume subtopics completes your topical map, and completeness itself strengthens rankings across the whole cluster, plus long-tail questions are disproportionately what AI search engines cite.
What's the difference between a keyword list and a topical map?
A keyword list is an inventory of demand: strings, volumes, difficulties. A topical map is an architecture: clusters assigned to planned URLs, organized into hubs and spokes, with explicit link relationships and one query family per page. The list answers 'what do people search?'; the map answers 'what do we build, and how does it connect?'
How often should keyword research be redone?
Do the full entity-and-clustering exercise once, then refresh incrementally every six months or so: pull new Search Console queries, re-check volatile SERPs for intent drift, and slot newly discovered clusters into the existing map. The map's skeleton, its entities and hubs, should stay stable for years; it's the leaves that grow.
Founder & Semantic SEO Lead · Permanent SEO
Writes about entity SEO, topical authority, and how modern and AI-powered search actually rank content.
View full profile