๐Ÿ“š General & Other

Sorting by Length Requires a Computed Key

Sorting by length is not a comparison of the strings themselves, so it needs the length computed first and sorted on.

Sorting by length is not a comparison of the strings themselves, so it needs the length computed first and sorted on.

The decorate-sort-undecorate pattern

awk '{print length($0)"\t"$0}' | sort -rn | cut -f2- prefixes each line with its length, sorts numerically in reverse, then strips the prefix. This is the general shape for sorting by any derived value, and it is worth recognising because it generalises to word counts, dates and anything else you can compute per line.

Numeric sort is not lexical sort

Without -n, sort compares the prefixes as text and places 100 before 99. This is the most common bug in hand-rolled length sorts, and it is easy to miss because the output still looks broadly ordered.

Length means characters, not bytes

awk's length() counts characters in a UTF-8 locale and bytes otherwise. For accented or non-Latin text the two differ, and the sort order changes accordingly. Set the locale deliberately rather than inheriting whatever the environment provides.

Try it: Sort by Length (Longest First) on SeoWolf's Notepad.