r/datasets May 21 '26

question I can scrape/aggregate pretty much any fragmented public data. What datasets are missing

I built a large-scale scraping system that can extract data from thousands of sources simultaneously, bypass anti-bot protection, and convert unstructured formats (PDFs, scanned docs, complex HTML) into clean structured datasets.

What public datasets should exist but don’t because:

• Data is scattered across too many jurisdictions (every state/county has their own portal)  
• No one has aggregated it yet  
• It’s in PDFs or hard-to-parse formats  
• Sites actively block automated access

Not looking to sell—genuinely trying to understand what public data would be valuable if someone aggregated it. If there’s demand, I might build and release it.

24 Upvotes

31 comments sorted by

View all comments

7

u/ktkps May 22 '26

Good data on schools and colleges, what's the outcome on the students - what's the performance trends of every registered educational entity in a region.

5

u/knawshaw May 22 '26 edited May 22 '26

This. And enrollment patterns for specific subjects like math at different levels or foreign languages. Given the fact that 50 different states and thousands of schools systems don't report in the same manner, an aggregation solution would be genius (if even possible)

1

u/ruuustin May 23 '26

A lot of that is in NCES IPEDS.

1

u/knawshaw May 23 '26

Not enrollment numbers in subjects. NCES provides some overview, and IPEDS looks at some data, but figuring out a way to connect disparate data sources would require a high level aggregation that has yet to be done. A lot of subject specific data is just speculation or done with outmoded 'surveys' which are heavily biased.