πŸ“š General & Other

Why uniq Only Removes Adjacent Duplicates

is the classic source of confusion in shell pipelines: it removes only adjacent duplicate lines, which is why is the idiom and alone rarely does what people expect.

uniq is the classic source of confusion in shell pipelines: it removes only adjacent duplicate lines, which is why sort | uniq is the idiom and uniq alone rarely does what people expect.

The design is deliberate

Operating on a stream one line at a time means constant memory regardless of input size. Full deduplication requires remembering everything seen so far; adjacent deduplication requires remembering one line. That is the trade, and it is why uniq handles files larger than memory.

Useful flags

uniq -c prefixes each run with its length, which turns a sorted list into a frequency count. uniq -d shows only values that repeated; uniq -u only those that did not.

When order is data, do not sort

Reaching for sort -u on sequential data destroys the sequence that gave the data meaning. If you need full deduplication with order preserved, awk '!seen[$0]++' is the tool β€” at the cost of holding the distinct set in memory.

Try it: Remove Consecutive Duplicate Lines on SeoWolf's Notepad.