A blog with two hundred posts produces a crawl export of several thousand URLs, most of them paginated archive pages: /blog/page/2/ through /blog/page/40/, repeated for every category and tag.
Reducing to the real pages
Removing lines containing digits strips the pagination and leaves the canonical content URLs. The site's actual shape becomes visible for the first time in the export.
Sanity-check what left
Posts with a year or a number in the slug get removed too. If the site uses dated URLs, this filter is the wrong tool and a more specific pattern is needed β check a sample of what disappeared before trusting the result.
What the ratio tells you
If pagination outnumbers real content by ten to one, crawlers are spending most of their budget on archive pages that rank for nothing. That is worth addressing at the source, with noindex on deep archive pages, rather than only in the export.
Try it: Remove Lines With Digits on SeoWolf's Notepad.