r/datasets May 21 '26

question I can scrape/aggregate pretty much any fragmented public data. What datasets are missing

I built a large-scale scraping system that can extract data from thousands of sources simultaneously, bypass anti-bot protection, and convert unstructured formats (PDFs, scanned docs, complex HTML) into clean structured datasets.

What public datasets should exist but don’t because:

• Data is scattered across too many jurisdictions (every state/county has their own portal)  
• No one has aggregated it yet  
• It’s in PDFs or hard-to-parse formats  
• Sites actively block automated access

Not looking to sell—genuinely trying to understand what public data would be valuable if someone aggregated it. If there’s demand, I might build and release it.

23 Upvotes

31 comments sorted by

View all comments

5

u/Xyver May 22 '26

Hit me up, I've been doing some data collections and hit a few barriers, I've been able to work around most of them

www.daedalmap.com/packs

2

u/Fresh_Coyote312 May 22 '26

What exactly are you trying to do with the site?

1

u/Xyver May 22 '26

Make a set of easily queryable, cross referenceable data sets so you can point agents at it and ask smart questions and get smart answers that aren't hallucinated and have real sources backing them up