📚 General & Other

Word Counting Is a Tokenisation Decision

gives a word count per line by splitting on whitespace. That definition is a choice, and different tools make it differently — which is why two word counts of the same text disagree.

awk '{print $0 " " NF}' gives a word count per line by splitting on whitespace. That definition is a choice, and different tools make it differently — which is why two word counts of the same text disagree.

Hyphens and apostrophes

Is state-of-the-art one word or four? Is don't one or two? Whitespace splitting says one and one. A tokeniser splitting on non-alphanumerics says four and two. Neither is wrong; they answer different questions. Just do not compare counts produced by both.

Runs of whitespace

Splitting on a literal single space turns a double space into an empty token, inflating the count. awk's default field splitting handles runs correctly, but a naive split(' ') in most languages does not — that is the usual source of counts that are mysteriously a few too high.

Numbers and punctuation as words

Top 10 tips! is three words plus a number, or four tokens, depending on whether digits count. For keyword length banding it rarely matters, as long as one rule is applied to the entire list. Mixing rules within a dataset is what produces bands that do not mean anything.

Try it: Append Word Count per Line on SeoWolf's Notepad.