๐Ÿ“š General & Other

Sorting on a Computed Field Without Losing the Line

Sorting by word count needs the count computed, sorted on, and then discarded โ€” the decorate-sort-undecorate pattern.

Sorting by word count needs the count computed, sorted on, and then discarded โ€” the decorate-sort-undecorate pattern.

The pipeline

awk '{print NF"\t"$0}' | sort -n | cut -f2- prefixes each line with its field count, sorts numerically, then strips the prefix. The same shape works for any derived key.

Numeric sort is mandatory

Without -n the prefixes sort as text and 10 precedes 2. The output still looks ordered at a glance, which is what makes this bug survive review.

Ties need a secondary key

Thousands of lines share a word count, and their relative order among equals is arbitrary unless specified. sort -k1,1n -k2 sorts by count then alphabetically within each count, producing a deterministic result. Without it, two runs over the same data can differ โ€” which is confusing when someone tries to reproduce your output.

Delimiter collision

If lines can contain tabs, a tab-delimited prefix breaks cut. Choose a delimiter absent from the data, or use awk to strip the prefix by field position instead.

Try it: Sort by Word Count (Fewest First) on SeoWolf's Notepad.