Extracting distinct words is a set operation over tokens, and every decision about what counts as the same token changes the result.
Case folding is the obvious one
Without it, SEO, Seo and seo are three entries. Lowercasing merges them โ and also merges US the country with us the pronoun, which for a term-frequency profile is usually an acceptable loss.
Stemming versus lemmatisation
run, runs, running are three tokens and one concept. Stemming chops suffixes mechanically and is fast, but produces non-words: running becomes run, while business becomes busi. Lemmatisation maps to real dictionary forms and needs a language model behind it. For quick content profiling, stemming is usually enough; for anything user-facing, its output is too ugly to display.
Unicode normalisation before comparison
The same accented character has multiple valid byte representations. Two visually identical words can compare as different unless both are normalised to the same form first โ normalise to NFC before building the set, or the count is quietly wrong on any non-English text.
Try it: Extract Unique Words on SeoWolf's Notepad.