๐Ÿ“š General & Other

Why You Cannot Strip HTML With a Regex

removes anything between angle brackets. It works on well-formed fragments and fails on a great deal of real-world HTML.

s/<[^>]*>//g removes anything between angle brackets. It works on well-formed fragments and fails on a great deal of real-world HTML.

Script and style content is not markup

The tags around a <script> block are removed; the JavaScript inside is not. Strip a full page that way and the output is a wall of code interspersed with prose. Those elements must be removed with their contents, which a single tag-matching expression cannot express.

Angle brackets appear in attributes and comments

<img alt="a > b"> terminates the naive pattern at the wrong >, leaving b"> in the output. Conditional comments and CDATA sections break it in similar ways.

Entities survive the strip

Removing tags leaves &amp;, &nbsp; and &#8217; intact. Text that still contains entities is not plain text, and a word count over it is still wrong โ€” decoding has to happen as a separate pass after stripping.

The right tool

Parse the document and read its text content. Every language has a parser that handles malformed markup the way browsers do. For a quick one-off paste, a stripping tool is fine; for anything in a pipeline, use the parser.

Try it: Remove HTML / BBCode Tags on SeoWolf's Notepad.