๐Ÿ“š General & Other

There Is No Correct Regex for Matching URLs

Every URL-matching regex is a trade-off between catching valid addresses and rejecting invalid ones, and the specification is permissive enough that no single pattern gets both right.

Every URL-matching regex is a trade-off between catching valid addresses and rejecting invalid ones, and the specification is permissive enough that no single pattern gets both right.

Where the boundary is genuinely ambiguous

Read more at https://example.com/page. โ€” is the final period part of the URL or the end of the sentence? Both are legal. Most implementations strip trailing punctuation, which is right almost always and wrong for the rare URL that legitimately ends in one.

Parentheses are worse. Wikipedia URLs contain balanced parentheses, so a link inside a parenthetical aside requires counting bracket depth to terminate correctly.

Practical extraction

Match a scheme, then consume until whitespace, then trim trailing .,;:!? and unbalanced closing brackets. That handles the overwhelming majority of real text.

Extracting from HTML is a different job

Do not regex HTML for links. Parse it and read href attributes. A regex cannot reliably tell an attribute from text that looks like one, and it will miss relative URLs entirely โ€” which are usually the ones you actually wanted from a page's source.

Try it: Extract URLs from Text on SeoWolf's Notepad.