๐Ÿ“š General & Other

Set Difference and the Cost of Sorted Input

returns lines in A but not B. It requires both inputs sorted and will produce wrong output without warning if they are not.

comm -23 a.txt b.txt returns lines in A but not B. It requires both inputs sorted and will produce wrong output without warning if they are not.

Sorted versus hashed

comm streams two sorted files in one pass with almost no memory. grep -vxFf b.txt a.txt needs no sorting but holds all of B in memory. For large inputs the first is dramatically cheaper; for unsorted inputs the second avoids a sort you would otherwise pay for.

Locale changes sort order

sort respects the locale, so the collation used to sort A must match the one used for B. Sorting one file under a different LCALL than the other produces interleaved output that comm misreads entirely. Set LCALL=C for both to get deterministic byte ordering.

Duplicates distort the result

If A contains a value twice and B once, set semantics say the value is excluded. Line-based tools may emit the extra copy. Deduplicate both inputs first so the operation matches the set logic you actually intended.

Try it: List Difference (Keyword Gap Finder) on SeoWolf's Notepad.