๐Ÿ“š General & Other

Punctuation Is a Unicode Category, Not a Character List

removes ASCII punctuation. Real text contains a great deal that is not ASCII.

tr -d '[:punct:]' removes ASCII punctuation. Real text contains a great deal that is not ASCII.

Smart quotes are not quotes

Word processors and CMS editors convert ' to ' and " to ". Those are different code points, outside ASCII, and [:punct:] in a C locale does not match them. Text that appears cleaned still carries typographic quotes, and they break the comparisons the cleaning was meant to fix.

Match by category

A Unicode-aware engine can match \p{P} for all punctuation and \p{S} for symbols. That covers em dashes, ellipsis characters, guillemets and every other mark a real document contains, rather than the sixty-odd ASCII marks.

Decide about the middle cases first

Apostrophes inside words and hyphens inside compounds are the ones where removal changes tokenisation. A common approach is to remove punctuation at word boundaries and keep it word-internally, which preserves don't and state-of-the-art while still stripping trailing commas. Whichever rule you choose, apply it to the whole corpus โ€” mixing rules produces counts that cannot be compared.

Try it: Remove Punctuation on SeoWolf's Notepad.