wc -l, wc -w and wc -c give lines, words and bytes. Each definition has an edge that surprises people.
The final line without a newline
wc -l counts newline characters, not lines. A file whose last line has no trailing newline reports one fewer than it contains. Concatenating two such files also joins their last and first lines into one โ a genuine data corruption that no error reports.
Bytes versus characters
wc -c counts bytes; wc -m counts characters. On any non-ASCII text they differ, and the one you want depends on whether the downstream limit is a storage constraint or a display one.
Word definition
wc -w splits on whitespace, so state-of-the-art is one word and hello,world is one word. Anything expecting a tokeniser's definition will disagree.
Counting distinct values
wc -l after sort -u gives distinct lines, and the difference against the raw count is the duplicate volume. That pair of numbers is usually more informative than either alone.
Try it: Line & Word Statistics Counter on SeoWolf's Notepad.