📚 General & Other

Detecting Target Overlap Across a Page Inventory

Overlap detection is a pairwise comparison problem, and the naive implementation gets expensive fast: comparing every page against every other is O(n²), which is fine at 200 pages and unpleasant at…

Overlap detection is a pairwise comparison problem, and the naive implementation gets expensive fast: comparing every page against every other is O(n²), which is fine at 200 pages and unpleasant at 50,000.

Exact matching is not enough

Comparing target keyword strings for equality catches only the obvious cases. running shoes and shoes for running collide in the index and not in a string comparison. Normalising first — case-folding, sorting tokens, stripping stop words — catches far more of the real overlaps.

Where the signal actually is

The strongest evidence of cannibalization is not in your keyword mapping at all; it is in Search Console. Filter to a single query and look at which pages received impressions for it. If the URL receiving impressions changes over time, the index is switching between candidates — which is cannibalization observed directly rather than inferred.

Structuring the check

Group by normalised target, then report any group with more than one URL. That is a single pass with a hash map rather than a pairwise sweep, and it scales to any inventory you are likely to have.

Try it: Keyword Cannibalization Checker on SeoWolf's Notepad.