What is keyword clustering?
Clustering groups keywords by whether Google returns the same pages for them. If two queries share results in the SERP, one page can serve both - and that is the only reliable criterion, because it comes from the search engine rather than from word similarity.
In practice a cluster answers the question of how many texts you need at all. A dozen queries can fit inside one topic, while two seemingly close phrases may demand separate pages.
What do you get?
- Cluster cards with a label, intent and combined volume, each holding a table of phrases - a ready basis for a content plan.
- A pillar page suggestion per cluster: the topic that would serve the whole group of queries.
- Separate buckets for phrases that formed no cluster, have no search results, or whose results failed to load. Nothing disappears quietly - you know what stayed out of the analysis and why.
- Cluster expansion on demand: 10-15 related phrases with a type (long-tail, question, variant, subtopic) and a ready article title. That is a separate action costing 1 credit.
- CSV export with columns: keyword, cluster, volume, intent, pillar page.
Why cluster keywords?
Without clustering you get the most expensive mistake in content: several separate texts targeting phrases that mean the same thing to a search engine. The pages start competing with each other, none accumulates full strength, and the editorial budget goes into duplicating the same material.
Word similarity does not settle this - “car insurance” and “auto policy” share no word, yet their search results largely coincide. Conversely, two phrases differing by a single word can have disjoint SERPs, because the intent behind them differs.
So we group by overlapping URLs in the results: a cluster is a set of phrases one page can genuinely serve.
How does clustering work?
- 1
Fetching search results
For every phrase on the list we pull the search results - those are the input data, not the phrases themselves.
- 2
Phrase × URL matrix
We build a matrix of URL presence in the results and compute the similarity of every phrase pair.
- 3
Grouping into clusters
Phrases linked by similarity above the threshold land in shared groups.
- 4
Describing the cluster
The model labels the cluster, identifies its intent and proposes a pillar page topic.
- 5
Search volumes
We attach monthly search volumes and sort clusters by their combined volume.
What you set before the run
| Parameter | Range and default |
|---|---|
| Similarity threshold | 0.5-1.0 (default 0.95) - the higher, the tighter the clusters |
| Minimum cluster size | 1-20 (default 1) - groups smaller than this never become a cluster, and their phrases go to the “Unclustered” group |
| SERP results considered | 1-10 (default 5) - how many positions we compare |
What does it cost?
- 1 credit per clustering run - the same as a page audit.
- Expanding a cluster costs a separate credit and can be done once per cluster.
- If a job ends in an error, the credit is returned automatically.
Frequently asked questions
How is this different from clustering by word similarity?
The criterion differs: what counts is overlapping search results, not textual similarity. That is why “car insurance” and “auto policy” can land in one cluster despite sharing no word - Google returns the same pages for both.
Which similarity threshold should I pick?
The default 0.95 produces tight, confident clusters - good when you plan separate pages. Lowering it merges broader topical groups, useful when building a site section. You can run the same list twice with different thresholds and compare.
What happens to phrases that joined no cluster?
They go into a separate bucket and stay in the result. We distinguish three cases: the phrase had results but did not join others; the phrase has no results at all; fetching its results failed and is worth retrying.