๐Ÿ“š General & Other

Character Counts Depend on What You Mean by Character

appends a length to every line. Whether that number is the one you need depends on an encoding question most tooling leaves implicit.

awk '{print $0 " " length($0)}' appends a length to every line. Whether that number is the one you need depends on an encoding question most tooling leaves implicit.

Bytes, code points, and what users see

In UTF-8, cafรฉ is four characters and five bytes. An emoji is one visible character and four bytes. A flag emoji is one visible character built from two code points. wc -c counts bytes, wc -m counts characters, and neither necessarily matches what a platform's own validator counts.

Where the target system enforces a byte limit โ€” a database column declared in bytes, for instance โ€” a character count will pass text that the insert then rejects.

Combining marks split the difference

รฉ can be one code point or two: e plus a combining accent. Both render identically. A naive count reports different lengths for two visually identical strings, and normalising to NFC first is what makes them agree.

Practical guidance

For SERP and ad limits, count characters rather than bytes, normalise first, and leave headroom. Exact-fit strings are the ones that break when the data changes underneath you.

Try it: Append Character Count per Line on SeoWolf's Notepad.