๐Ÿ“š General & Other

Fragments, Encoded Question Marks, and Other Query Traps

truncates at the first . That is correct for most URLs and wrong for a few in ways worth knowing.

sed 's/?.*//' truncates at the first ?. That is correct for most URLs and wrong for a few in ways worth knowing.

The fragment comes after the query

/page?a=1#section truncates to /page, taking the fragment with it. Usually fine, since fragments are client-side only โ€” but if the site uses fragment routing, the fragment was the page identity and you have just discarded it.

Encoded question marks inside parameters

A parameter value containing an encoded ? is safe, because the encoded form is %3F. A raw ? in a value is technically permitted after the first one, so truncating at the first ? is correct while truncating at the last is not.

Parameter order makes duplicates

?a=1&b=2 and ?b=2&a=1 are the same page and different strings. If you are keeping parameters rather than stripping them, sort them before deduplicating or the same page counts twice.

Prefer a URL parser

Every language has one. It handles the component boundaries correctly, and reassembling from parsed components is safer than pattern-matching the string.

Try it: Strip Query Strings from URLs on SeoWolf's Notepad.