๐Ÿ“š General & Other

Trim Removes More Than the Space Character

trims, then drops empties. What counts as space is the part that varies.

sed 's/^[[:space:]]//; s/[[:space:]]$//' trims, then grep -v '^$' drops empties. What counts as space is the part that varies.

The characters that survive a naive trim

Non-breaking space arrives from HTML and PDFs and is not matched by an ASCII space class. The zero-width space is invisible and defeats both trimming and visual inspection. A UTF-8 byte-order mark at the start of a file attaches to the first value and makes exactly one row unmatchable, which is a memorable afternoon to debug.

Order matters

Trim before removing empty lines. A line containing only spaces is not empty until it has been trimmed, so reversing the order leaves whitespace-only rows in the output.

Language trim functions differ

Java's trim() removes characters below U+0020 only; strip() is Unicode-aware. Python's strip() handles Unicode whitespace. JavaScript's trim() does too, including the BOM. Knowing which behaviour you have decides whether the invisible characters actually go away.

Try it: Trim & Remove Empty Lines on SeoWolf's Notepad.