๐Ÿ“š General & Other

Sentence Boundaries Are Harder Than a Period

Splitting on and capitalising what follows handles the easy cases and mangles a surprising number of real sentences.

Splitting on . and capitalising what follows handles the easy cases and mangles a surprising number of real sentences.

Abbreviations are not sentence ends

Dr. Smith, e.g., Inc., 3.5 million โ€” every one contains a period that does not end a sentence. Naive splitting capitalises mid-sentence and produces Dr. Smith Was Here from a single clause. Real sentence segmentation uses an abbreviation list plus heuristics, and still gets edge cases wrong.

Case conversion is locale-dependent

Turkish has dotted and dotless i, so lowercasing I correctly produces ฤฑ in Turkish and i everywhere else. Using an invariant lowercase where a locale-aware one is needed, or the reverse, is a genuine and well-known source of bugs.

German รŸ has no single-character uppercase

Uppercasing รŸ traditionally yields SS, which changes the string length โ€” so uppercase-then-lowercase is not an identity operation. Any code assuming case conversion preserves length will be wrong on German text.

Practical approach

For headline lists, lowercase everything, capitalise the first character, then restore a maintained list of proper nouns and acronyms. Simple, predictable, and easy to inspect.

Try it: Convert Text to Sentence Case on SeoWolf's Notepad.