A crawl of a mid-size site returns forty thousand URLs. Perhaps twelve thousand are pages. The rest are images, stylesheets, scripts, PDFs and feeds — and if they stay in the export, every percentage you report is wrong.
Filtering by extension
Remove lines ending in .jpg, .png, .css, .js, .svg, .woff2 and .pdf, one pass each. The list shrinks toward the set of things that are actually pages.
Why this changes the numbers
"18% of URLs are missing a meta description" reads very differently once you know the denominator included every image on the site. Cleaning the denominator is not cosmetic — it is the difference between a finding and a misleading statistic.
Keep the PDFs somewhere
Do not discard them entirely. PDFs are indexable and sometimes rank, so they deserve their own small review — just not to be counted among the HTML pages.
Try it: Remove Lines Ending With on SeoWolf's Notepad.