๐Ÿ“š General & Other

Word Frequency Is a Hash Map and Three Hard Decisions

The classic implementation is a hash map from token to count, or in shell . The algorithm is settled. The preprocessing is where results are won or lost.

The classic implementation is a hash map from token to count, or in shell tr -s '[:space:]' '\n' | sort | uniq -c | sort -rn. The algorithm is settled. The preprocessing is where results are won or lost.

Raw counts favour long documents

A word appearing 40 times in a 10,000-word page is rarer than one appearing 5 times in a 200-word page. Comparing documents of different lengths on raw counts is meaningless โ€” normalise to a rate, or use TF-IDF, which weights each term by how unusual it is across the whole corpus rather than how often it appears in one document.

Case and punctuation change the answer

Shoes, shoes and shoes, are three keys unless you fold case and strip punctuation first. That alone typically merges 10-20% of a naive count's distinct entries.

Bigrams carry meaning unigrams lose

Single-word counts destroy phrases. running and shoes appearing frequently does not establish that running shoes appears at all. Counting adjacent pairs alongside single words recovers most of what tokenisation threw away.

Try it: Word Frequency Counter on SeoWolf's Notepad.