๐Ÿ“š General & Other

Uniqueness Depends on How You Normalise Words

Extracting distinct words is a set operation over tokens, and every decision about what counts as the same token changes the result.

Extracting distinct words is a set operation over tokens, and every decision about what counts as the same token changes the result.

Case folding is the obvious one

Without it, SEO, Seo and seo are three entries. Lowercasing merges them โ€” and also merges US the country with us the pronoun, which for a term-frequency profile is usually an acceptable loss.

Stemming versus lemmatisation

run, runs, running are three tokens and one concept. Stemming chops suffixes mechanically and is fast, but produces non-words: running becomes run, while business becomes busi. Lemmatisation maps to real dictionary forms and needs a language model behind it. For quick content profiling, stemming is usually enough; for anything user-facing, its output is too ugly to display.

Unicode normalisation before comparison

The same accented character has multiple valid byte representations. Two visually identical words can compare as different unless both are normalised to the same form first โ€” normalise to NFC before building the set, or the count is quietly wrong on any non-English text.

Try it: Extract Unique Words on SeoWolf's Notepad.