๐Ÿ“š General & Other

Whitespace Is More Than Space and Tab

squeezes repeated spaces and misses most of what causes matching failures in real data.

tr -s ' ' squeezes repeated spaces and misses most of what causes matching failures in real data.

The characters that break comparisons

U+00A0 non-breaking space is the usual culprit, arriving from HTML and PDFs. There are also en and em spaces, the zero-width space, and the byte-order mark, which frequently appears at the start of the first line of a UTF-8 file and makes the first value mysteriously unmatchable.

Regex whitespace classes vary

\s in a Unicode-aware engine matches the full whitespace category; in a byte-oriented one it matches only ASCII. [[:space:]] in POSIX tools depends on locale. Two implementations of "strip whitespace" can therefore disagree on the same input.

The reliable sequence

Normalise to NFC, replace every Unicode whitespace character with a plain space, collapse runs, then trim the ends. Doing this on ingest โ€” rather than at the point of comparison โ€” means everything downstream can assume clean input, which is a much easier property to reason about than remembering to normalise before each operation.

Try it: Normalize Whitespace on SeoWolf's Notepad.