Sorting tokens within a record to produce a comparable key is a standard normalisation move — the same idea behind sorting a set before hashing it. The implementation details are where results diverge.
Your sort is locale-dependent
Plain byte sorting puts every uppercase letter before every lowercase one, so Paris flights and flights Paris do not converge. Case-fold before sorting, or the whole exercise produces false negatives on any list with inconsistent capitalisation.
Beyond ASCII it gets worse: é sorts after z bytewise but next to e under a locale-aware collation. For keyword data in any European language, that distinction decides whether the canonical forms match.
Whitespace defines your tokens
Splitting on a single space breaks on double spaces and tabs, producing empty tokens that sort to the front and leave a leading space in the output. Split on runs of whitespace, and normalise the separator on the way out.
Punctuation attaches to words
shoes, and shoes are different tokens. Strip punctuation before sorting if the goal is comparison, or accept that a stray comma will keep two otherwise-identical phrases apart.
Try it: Alphabetize Words Within Each Line on SeoWolf's Notepad.