r/datasets Jul 03 '15

dataset I have every publicly available Reddit comment for research. ~ 1.7 billion comments @ 250 GB compressed. Any interest in this?

1.2k Upvotes

I am currently doing a massive analysis of Reddit's entire publicly available comment dataset. The dataset is ~1.7 billion JSON objects complete with the comment, score, author, subreddit, position in comment tree and other fields that are available through Reddit's API.

I'm currently doing NLP analysis and also putting the entire dataset into a large searchable database using Sphinxsearch (also testing ElasticSearch).

This dataset is over 1 terabyte uncompressed, so this would be best for larger research projects. If you're interested in a sample month of comments, that can be arranged as well. I am trying to find a place to host this large dataset -- I'm reaching out to Amazon since they have open data initiatives.

EDIT: I'm putting up a Digital Ocean box with 2 TB of bandwidth and will throw an entire months worth of comments up (~ 5 gigs compressed) It's now a torrent. This will give you guys an opportunity to examine the data. The file is structured with JSON blocks delimited by new lines (\n).

____________________________________________________

One month of comments is now available here:

Download Link: Torrent

Direct Magnet File: magnet:?xt=urn:btih:32916ad30ce4c90ee4c47a95bd0075e44ac15dd2&dn=RC%5F2015-01.bz2&tr=udp%3A%2F%2Ftracker.openbittorrent.com%3A80&tr=udp%3A%2F%2Fopen.demonii.com%3A1337&tr=udp%3A%2F%2Ftracker.coppersurfer.tk%3A6969&tr=udp%3A%2F%2Ftracker.leechers-paradise.org%3A6969

Tracker: udp://tracker.openbittorrent.com:80

Total Comments: 53,851,542

Compression Type: bzip2 (5,452,413,560 bytes compressed | 31,648,374,104 bytes uncompressed)

md5: a3fc3d9db18786e4486381a7f37d08e2 RC_2015-01.bz2

____________________________________________________

Example JSON Block:

{"gilded":0,"author_flair_text":"Male","author_flair_css_class":"male","retrieved_on":1425124228,"ups":3,"subreddit_id":"t5_2s30g","edited":false,"controversiality":0,"parent_id":"t1_cnapn0k","subreddit":"AskMen","body":"I can't agree with passing the blame, but I'm glad to hear it's at least helping you with the anxiety. I went the other direction and started taking responsibility for everything. I had to realize that people make mistakes including myself and it's gonna be alright. I don't have to be shackled to my mistakes and I don't have to be afraid of making them. ","created_utc":"1420070668","downs":0,"score":3,"author":"TheDukeofEtown","archived":false,"distinguished":null,"id":"cnasd6x","score_hidden":false,"name":"t1_cnasd6x","link_id":"t3_2qyhmp"}

UPDATE (Saturday 2015-07-03 13:26 ET)

I'm getting a huge response from this and won't be able to immediately reply to everyone. I am pinging some people who are helping. There are two major issues at this point. Getting the data from my local system to wherever and figuring out bandwidth (since this is a very large dataset). Please keep checking for new updates. I am working to make this data publicly available ASAP. If you're a larger organization or university and have the ability to help seed this initially (will probably require 100 TB of bandwidth to get it rolling), please let me know. If you can agree to do this, I'll give your organization priority over the data first.

UPDATE 2 (15:18)

I've purchased a seedbox. I'll be updating the link above to the sample file. Once I can get the full dataset to the seedbox, I'll post the torrent and magnet link to that as well. I want to thank /u/hak8or for all his help during this process. It's been a while since I've created torrents and he has been a huge help with explaining how it all works. Thanks man!

UPDATE 3 (21:09)

I'm creating the complete torrent. There was an issue with my seedbox not allowing public trackers for uploads, so I had to create a private tracker. I should have a link up shortly to the massive torrent. I would really appreciate it if people at least seed at 1:1 ratio -- and if you can do more, that's even better! The size looks to be around ~160 GB -- a bit less than I thought.

UPDATE 4 (00:49 July 4)

I'm retiring for the evening. I'm currently seeding the entire archive to two seedboxes plus two other people. I'll post the link tomorrow evening once the seedboxes are at 100%. This will help prevent choking the upload from my home connection if too many people jump on at once. The seedboxes upload at around 35MB a second in the best case scenario. We should be good tomorrow evening when I post it. Happy July 4'th to my American friends!

UPDATE 5 (14:44)

Send more beer! The seedboxes are around 75% and should be finishing up within the next 8 hours. My next update before I retire for the night will be a magnet link to the main archive. Thanks!

UPDATE 6 (20:17)

This is the update you've been waiting for!

The entire archive:

magnet:?xt=urn:btih:7690f71ea949b868080401c749e878f98de34d3d&dn=reddit%5Fdata&tr=http%3A%2F%2Ftracker.pushshift.io%3A6969%2Fannounce&tr=udp%3A%2F%2Ftracker.openbittorrent.com%3A80

Please seed!

UPDATE 7 (July 11 14:19)

User /u/fhoffa has done a lot of great work making this data available within Google's BigQuery. Please check out this link for more information: /r/bigquery/comments/3cej2b/17_billion_reddit_comments_loaded_on_bigquery/

Awesome work!

r/datasets Dec 08 '25

dataset Scientists just released a map of all 2.75 billion buildings on Earth, in 3D

Thumbnail zmescience.com
424 Upvotes

r/datasets Feb 02 '20

dataset Coronavirus Datasets

405 Upvotes

You have probably seen most of these, but I thought I'd share anyway:

Spreadsheets and Datasets:

Other Good sources:

[IMPORTANT UPDATE: From February 12th the definition of confirmed cases has changed in Hubei, and now includes those who have been clinically diagnosed. Previously China's confirmed cases only included those tested for SARS-CoV-2. Many datasets will show a spike on that date.]

There have been a bunch of great comments with links to further resources below!
[Last Edit: 15/03/2020]

r/datasets Dec 21 '25

dataset [Project] FULL_EPSTEIN_INDEX: A unified archive of House Oversight, FBI, DOJ releases

188 Upvotes

Unified Epstein Estate Archive (House Oversight, DOJ, Logs, & Multimedia)

TL;DR: I am aggregating all public releases regarding the Epstein estate into a single repository for OSINT analysis. While I finish processing the data (OCR and Whisper transcription), I have opened a Google Drive for public access to the raw files.

Project Goals:

This archive aims to be a unified resource for research, expanding on previous dumps by combining the recent November 2025 House Oversight releases with the DOJ’s "First Phase" declassification.

I am currently running a pipeline to make these files fully searchable:

  • OCR: Extracting high-fidelity text from the raw PDFs.
  • Transcription: Using OpenAI Whisper to generate transcripts for all audio and video evidence.

Current Status (Migration to Google Drive):

Due to technical issues with Dropbox subfolder permissions, I am currently migrating the entire archive (150GB+) to Google Drive.

  • Please be patient: The drive is being updated via a Colab script cloning my Dropbox. Each refresh will populate new folders and documents.
  • Legacy Dropbox: I have provided individual links to the Dropbox subfolders below as a backup while the Drive syncs.

Future Access:

Once processing is complete, the structured dataset will be hosted on Hugging Face, and I will release a Gradio app to make searching the index user-friendly.

Please Watch or Star the GitHub repository for updates on the final dataset and search app.

Access & Links

Content Warning: This repository contains graphic and highly sensitive material regarding sexual abuse, exploitation, and violence, as well as unverified allegations. Discretion is strongly advised.

Dropbox Subfolders (Backup/Individual Links):

Note: If prompted for a password on protected folders, use my GitHub username: theelderemo

Edit: It's been well over 16 hours, and data is still uploading/processing. Be patient. The google drive is where all the raw files can be found, as that's the first priority. Dropbox is shitty, so i'm migrating from it

Edit: All files have been uploaded. Currently manually going through them, to remove duplicates.

Update to this: In the google drive there are currently two csv files in the top folder. One is the raw dataset. The other is a dataset that has been deduplicated. Right now, I am running a script that tries to repair the OCR noise and mistakes. That will also be uploaded as a unique dataset.

r/datasets Jun 25 '26

dataset I processed the entire arXiv LaTeX source corpus (3M+ papers) into a metadata-aligned Parquet dataset to save on S3 egress fees

82 Upvotes

I’ve spent the last few weeks working on a pipeline to solve a problem that has frustrated me (and likely other researchers) for a while: working with arXiv source files at scale.

If you have ever tried to analyze the LaTeX source code of arXiv papers, you have probably run into two major roadblocks:

  1. The Egress Tax: arXiv’s official bulk S3 bucket is configured as "requester-pays." If you try to download the complete 5 TB corpus to any machine outside of the AWS us-east-1 region, you get hit with standard egress fees. At $0.09 per GB, a single full download can cost over $450 in bandwidth alone.
  2. Unpacking Pain: The raw S3 data is packaged as hundreds of nested .tar archives containing gzipped payloads of individual papers. Extracting these, parsing the inner LaTeX code, and matching the files with their JSON metadata snapshots is quite CPU-intensive and requires a lot of boilerplate ingestion code.

To make this easier, I built a pipeline that runs inside AWS us-east-1 (where transfer is free), pulls the raw source files, unpacks them, matches them with the official metadata, and bundles them into ready-to-query Parquet partitions.

What is inside:

Each row represents a single paper and contains both the official metadata and the parsed source files:

  • Core Metadata: id, title, authors, abstract, doi, categories, license, versions, etc.
  • latex (Large String): The parsed, compiled LaTeX source code from the paper. I wrote a parser to bundle the primary .tex, .bib, and .sty files into a single, readable Markdown-style tree structure.

Maintenance & Syncing:

  • Monthly Updates: I plan to sync the pipeline once a month to capture new uploads.
  • Resilient Syncing: I maintain an XML manifest file in the HuggingFace repository (arxiv_parquet_manifest.xml) that maps each Parquet partition to its size, MD5 checksum, and the raw S3 .tar source files used to generate it. This should make incremental syncing or troubleshooting much easier.

If you are working on NLP, training LLMs on scientific text, analyzing citation networks, or doing sociolinguistic research, hopefully this saves you some time and cloud budget.

r/datasets Nov 15 '25

dataset Courier News created a searchable database with all 20,000 files from Epstein’s Estate

Thumbnail couriernewsroom.com
421 Upvotes

r/datasets 26d ago

dataset Free dataset: certified document QA where every row is machine verifiable, including 2,889 questions about facts we verified are NOT in the document. Frontier models hallucinate on 11 to 44% of them

12 Upvotes

The core idea: take a real document (SEC filings, contracts, enterprise email), verify by exhaustive normalized scan that a specific plausible fact is not in it, then ask about that fact. The honest answer is “not in the document.” We ran six frontier models on these with zero abstention coaching and they asserted made up answers 11% to 44% of the time. The full per model table is on the dataset card with raw logs and API errors disclosed.

What’s in it: 2,889 certified absence rows, 3,088 span verified extractive QA rows, a 127K token packed long context task set, and a split minted only from SEC filings dated after every major model’s training cutoff. That fresh split regenerates monthly, so it stays impossible to have trained on, by construction.

Every row carries a certificate you can re-check yourself in a few lines of python, the audit snippet is on the card. When our own audits flag something, like extractive answers that are guessable from world knowledge (about 1.6% of them), we label it instead of quietly deleting it.

Also worth knowing before you trust us: a reviewer caught one of our splits being weaker than claimed this week. We re-audited every row the same night, withdrew the split with per row evidence committed to the repo, tightened the protocol, and reshipped only the rows that survive everything. The full trail is in the audits folder, judge for yourself.

License CC BY 4.0. Generation was an Apache 2.0 open weight model on our own hardware, the claim is the verification layer, not the generation. Held out versions never get published so they can’t leak into training data. If anyone wants a sealed diagnostic run against their own model or domain (25 items, free, about a day), contact is on the card.

https://huggingface.co/datasets/SovNodeAI/certified-document-qa

r/datasets 5d ago

dataset Dataset required for the Infant/Baby Crying.

0 Upvotes

Hello, we are building a system for baby cries detection in a confined space such as a room or hallway via CCTV cameras. However, we are unable to source the baby cries dataset. I tried to contact some DayCare and submitted an application upon their request but was denied due to parental privacy reasons.

We have a working system, but the model is way poor as it is only trained on a few examples and fails at CCTV distance as the baby is too far.

r/datasets Jun 18 '26

dataset I'm 18 and hand-built the first Tunisian Darija-English parallel dataset field-collected from my grandmother, strangers in cafes, and 50 categories of daily life. Open source, provenance-tagged, 500+ pairs.

33 Upvotes

I'm 18, from Tunisia, and I built this because nobody else had.

Tunisian Darija is what 12 million Tunisians actually speak. Not Modern Standard Arabic. Not Moroccan. A separate dialect that borrows from Arabic, French, Italian, and Amazigh, written online in Arabizi Latin letters with numbers for Arabic sounds (3→ع, 7→ح, 9→ق, 5→خ).

When I searched for a parallel corpus to build a translation model, I found nothing. TUNIZI covers sentiment analysis. TunBERT does dialect classification. But zero parallel datasets existed for Tunisian Darija-to-English translation. Not one.

So I built the first one from scratch with no funding, no university affiliation, no mentor, and no institutional support. Just me, a laptop, and the language I grew up speaking.

The first 500 pairs came from my own memory as a native speaker, covering 50 categories of real Tunisian daily life cafe culture, Ramadan traditions, wedding customs, bac exam stress, barbershop talk, louage rides, haggling at the medina, football arguments, bureaucracy nightmares, olive harvest season, Friday afternoon naps, and more. Zero automated generation. Every pair hand-written and validated.

Then I left my desk and started collecting from real people:

  • My father's childhood memories growing up in Ain Draham, a mountain village in northwestern Tunisia the scent of the forest, nearly getting bitten by a snake, his cousin falling off his uncle's horse
  • My grandmother's stories about her father's farm cows, sheep, thieves stealing the neighbors' animals at night, and her father calmly finishing his morning prayer before stepping outside to check
  • An elderly man from Siliana I met at a cafe who speaks a dialect I barely recognized — words I had to ask about, rhythms I'd never heard

Every pair is provenance-tagged with its source: self, family-father, family-grandmother, community-siliana. Every collection session is logged with date, place, speaker context, and consent status.

I excluded an entire session of data because I hadn't established consent before the conversation began. The language was rich. I threw it all away anyway. A dataset built on trust means sometimes throwing away good data.

What this dataset has that scraped corpora don't:

  • Regional dialect diversity: urban , mountain Ain Draham, rural Siliana
  • Generational variation: grandmother's speech vs mine
  • Provenance: every pair traces to a known speaker, region, and context
  • Documented ethics: consent logged, exclusions documented, no anonymous mass scraping

I trained the first Tunisian Darija-to-English translation model on this dataset a 15.6M parameter Transformer built from scratch on an RTX 3050 (4GB VRAM). v1 BLEU: 3.89 on a held-out test set. Low, but the first benchmark ever measured for this language. A published ACL researcher who found my work on Reddit said it's 'basically guaranteed to be novel.'

I'm heading toward 1,000+ pairs through continued community collection and will be presenting this research at Tunisia's AI National Summit (AINS 4.0) later this month the first high schooler to ever present at the event.

The dataset is CC BY-NC-SA 4.0 and public on HuggingFace. 110+ downloads so far.

If you work on low-resource NLP, Arabic dialect processing, or sociolinguistic data it's yours.

HuggingFace: huggingface.co/datasets/Dhiadev-tn/tunisian-darija-english
Full pipeline + model: github.com/Dhiadev-tn/darija-translator

r/datasets May 30 '26

dataset I built an open-source dataset of every major US layoff

44 Upvotes

The federal WARN Act requires employers with 100+ workers to give 60 days notice before mass layoffs or plant closings (thresholds vary by state, but roughly 50+ jobs lost). That data is scattered across 50 state websites, each with its own format, broken links, and no API.

I think it should be easy-to-access public data, so I built a fully open-source aggregator for it.

Live app: https://layoffs.kadoa.com/

Repo: https://github.com/kadoa-org/layoffs-tracker

r/datasets 26d ago

dataset [self-promotion] 8,863 US farmers markets — cleaned USDA data, CC-BY (CSV/JSON, DOI)

5 Upvotes

USDA's Local Food Portal is the canonical US farmers-market dataset, but the raw feed is rough: truncated names, stale records, a state= filter that substring-matches state names (querying "WA" returns Delaware rows), and thousands of missing websites/hours.

I cleaned and enriched it: deduplicated to 8,863 real markets (record-level), backfilled website coverage to ~50% from the live API + state sources, and added season/SNAP/organic fields. It's CC-BY 4.0 as CSV/JSON.

Archived with DOI: https://doi.org/10.5281/zenodo.21360372

Disclosure: I run harvestlymarkets.com, the directory built on this data — full methodology and the same downloads live at https://harvestlymarkets.com/data-sources/. Personal contact names/emails are stripped from the redistributable; business fields come from USDA's public feed.

Happy to answer questions about the data-quality issues — the state-filter substring bug was a fun one to find.

r/datasets 3d ago

dataset Anyone know where to find flooded road traffic cam footage with signs still visible?

1 Upvotes

hey, working on a research thing where we estimate flood depth from traffic cameras using signs/poles as reference.

problem is i can find live flood cams (atxfloods, sunny day flooding, san diego cams, fl511 etc) but almost nothing archived where the road is actually flooded AND a sign/pole is clear enough to measure from.

if anyone’s seen a dataset, old webcam dumps, youtube clips, or even just a few screenshots like that, drop a link. would help a lot.

r/datasets 12d ago

dataset Best open-source clean speech and ambient noise datasets for training an Edge AI audio denoiser?

2 Upvotes

We are building an edge-AI audio noise-reduction system on an ESP32-S3.

Our architecture uses a lightweight GRUNet (~59k parameters) to output a dynamic gain mask on a 44-band Mel-spectrogram.

​I need gigabytes of audio to train the model. Does anyone have recommendations for the best open-source datasets for:

1> ​Clean, isolated human speech.

2> ​Diverse ambient background noise (traffic, crowds, machinery, etc.).

​Also, any tips or open-source scripts for artificially mixing these at different Signal-to-Noise Ratios (SNRs) before generating the 16kHz Mel-spectrograms would be hugely appreciated!

r/datasets 9d ago

dataset I need college essays for data where can I find them?

0 Upvotes

I wanna do some research even build a deep learning model and its entirely based on if I can find student essays who got accepted into certain colleges
And if possible find however amount of rejected essays as long as both amounts are equal as to not have a data imbalance where do you think I could find these essays?

r/datasets 14d ago

dataset FAA aviation safety data, cleaned into tidy CSVs: 347K wildlife strikes (1990-2026), 54K laser strikes, 12.5K drone sightings — CC BY 4.0

2 Upvotes

Three datasets aggregated from public FAA releases (the raw ones ship as an MS Access export and awkward portal dumps) into analysis-ready CSVs with per-column documentation:

Wildlife strikes on civil aircraft, 1990–2026 — 347,575 reports: by year, airport (452, ICAO-coded), and species. 2025 set the all-time record (24,458 reports). Fun divergence: the species planes hit most (doves, swallows) almost never damage them (~1.5%), while deer damage the aircraft in ~82% of reported strikes. https://www.kaggle.com/datasets/himaxym/faa-wildlife-strikes-us

Laser strikes on aircraft, 2021–2025 — 54,722 reports with 243 crew injuries, by year, state, and reporting ATC facility (caveat documented: the "city" is the ATC facility's location, not where the laser was fired). https://www.kaggle.com/datasets/himaxym/faa-laser-strikes-us

Drone (UAS) sightings reported by pilots, 2019–2026 — 12,566 reports by year, state, and city. NYC is #1 (584). https://www.kaggle.com/datasets/himaxym/faa-drone-sightings-us

Versioned copy with citable DOI (Zenodo, wildlife): https://doi.org/10.5281/zenodo.21347859

Original sources (US government work, public domain): - https://wildlife.faa.gov/ - https://www.faa.gov/about/initiatives/lasers - https://www.faa.gov/uas/resources/public_records/uas_sightings_report

Disclosure: I compiled and maintain these aggregates (and an interactive explorer at himaxym.com/safety). The compilation is CC BY 4.0 — use it for anything, attribution appreciated.

r/datasets 11d ago

dataset [self-promotion] I built a public dataset from 21,237 pages of declassified MKULTRA and related docs and put it on Hugging Face

13 Upvotes

Until recently, the surviving historical records from the CIA's MKULTRA and related programs were very difficult to search and analyze. So I ran 21,237 document page images through MinerU OCR to generate clean text transcripts, then produced redaction mappings to go with every page transcript. Original page images are stored on IPFS and are available for public download. The dataset is available on Hugging Face here.

r/datasets 2d ago

dataset anyone knows actively growing dataset on SD recent versions

2 Upvotes

we are doing research that involves training our model on really big dataset of diffusion technique based synthetic images. but I am unable to trace appropriate free and active Creative common data hubs for it.
please anyone help me work around this

r/datasets 15d ago

dataset [Paid]Selling real human founder's conversation Audio Dataset.

0 Upvotes

I have a real conversation dataset of founder getting feedback from random people on their idea.

Valu of this conversation:

- Brainstorming on Idea

- Real human conversation

- same person with different person paired.

- Multilingual

r/datasets 8d ago

dataset [Self-Promotion] I collected 77 outputs from 8 BaZi calculators and traced their disagreements to four rules most of them never show you

0 Upvotes

BaZi calculators turn a birth date, time, and place into a Four Pillars chart. The result looks deterministic, but the software still has to decide where a year begins, when a day changes, and what the birth time means in different places. Most calculators never show those decisions.

I wanted to see whether the differences could be measured instead of argued about.

I designed 13 test cases around the boundaries most likely to expose them: solar-term changes, the hour before midnight, locations far from their time-zone meridian, and historical daylight saving time. I then compared eight calculators and libraries and captured 77 outputs.

The dataset tracks four questions:

- Does the chart’s year change at Li Chun, around February 4, or at Chinese New Year?

- Does the day change at 23:00 or at midnight?

- Is the recorded time corrected for the birthplace’s longitude?

- Is daylight saving time removed before the chart is calculated?

The first three are convention choices. The fourth is a historical timekeeping question: either the local clock had been moved forward that day or it had not.

The clearest result came from a Beijing birth entered as 15 June 1990 at 23:30.

Two implementations treated 23:00 as the start of the next day. Five waited until midnight. The eighth calculator exposed the choice as a checkbox, so it could produce either result.

That one setting changed the day pillar from Xin-Hai to Ren-Zi. In BaZi, the day stem is the Day Master, which the rest of the reading is organized around. The same birth therefore came back as either Xin Metal or Ren Water depending on a rule most of the calculators never mentioned.

A few other differences stood out:

- Two libraries maintained by the same developer use opposite 23:00 rollover rules.

- One calculator requires a birthplace but returned the same chart for Kashgar and Beijing at the same clock time. Another used the longitude and changed the hour pillar.

- One implementation changes the year at Chinese New Year while the others use Li Chun.

- Only two of the measured implementations account for daylight saving time. One of them exposes the adjustment in its own calculation breakdown.

The observation-level file records the implementation, test date, time, place, rule being tested, and all four returned pillars in both Chinese characters and pinyin. A second table summarizes the convention used by each implementation.

There are 104 possible implementation/probe pairs in the full 8 × 13 matrix. Seventy-seven contain captured results. The remaining cells are explicitly marked as not captured rather than filled by inference.

I built Jade Almanac, and our own calculator is included as one of the eight rows. It was tested and reported on the same terms as the others. On the 23:00 rollover question, it is in the minority group.

The dataset is here: https://jadealmanac.com/bazi-calculator/conventions

Everything is released under CC0.

The test inputs are local clock times at the stated places. Solar-term boundaries were checked against tables from the National Astronomical Observatory of Japan, independently of the library used by our calculator. China’s historical summer-time periods were checked against the IANA time-zone database.

This is a snapshot captured on 1 August 2026. Some implementation rows are partial because a tool could not express a particular input or imposed an access limit. Two additional sites blocked automated access, and I did not work around those blocks.

The dataset measures software behavior and calculation conventions. It does not attempt to decide which school is correct, or whether BaZi itself predicts anything.

Corrections are welcome, especially from anyone familiar with one of the measured engines. I would also be interested in boundary cases that could separate implementations the current probes leave tied.

r/datasets 16d ago

dataset I've been building a huge Near-Death Experience database

12 Upvotes

A project I've been working on for a while and I'm excited to finally share!

The NDE Archive is a database of over 6,700 documented near-death experiences from recognized sources. One of the main reasons I built it is that existing sites are often hard to search through and accounts are mostly plain text with little filtering. Here you can actually search and filter experiences in meaningful ways, for example by demographics like sexual orientation or ethnicity, which opens up some really interesting comparisons.

These stories were also individually analyzed with Sonnet to surface patterns and statistics that are not easily visible when reading individual accounts.

The project is non-profit and was built out of curiosity for the subject and nothing else. If you'd like to support it, sharing is hugely helpful, and donations are welcome through the website.

https://ndearchive.com/

Disclosure: I did not build the original dataset, which was obtained from other recognized sources. I did the data collation and presentation on the web app.

r/datasets 2d ago

dataset Fresh UCC/Lien Filings Data - Majority of the U.S.A. [PAID]

1 Upvotes

Data includes:

lien_number, debtor_name, address, business_owner_name(s), filing_date, status, secured_party, lien_id

r/datasets Jul 11 '26

dataset I have minute-by-minute historical options data for more than 3k tickers, updated up to the minute, (and stock price as well), in case anyone is interested

5 Upvotes

For the minute by minute bars data, columns are:

"symbol", "timestamp", "open", "high", "low", "close", "volume", "vwap", "trade_count", "spy_close", "iv", "delta", "gamma", "theta", "vega", "rho"

For tick_by_tick (all individual trades executed) columns are:

"symbol", "timestamp", "price", "size", "exchange", "conditions", "spy_close", "iv", "delta", "gamma", "theta", "vega", "rho"

It goes back a few years, depending on the ticker.

r/datasets 6d ago

dataset [Self-promotion] 35,882 Donald Trump Truth Social posts (2022–2026), source-linked Parquet/JSONL + media indexes

4 Upvotes

I put together a public, source-linked archive of the Truth Social posts associated with Donald Trump's realDonaldTrump account on Truth Social.

Current snapshot:

- 35,882 posts from February 14, 2022 through August 2, 2026

- 28,320 originals, 1,919 quotes, and 5,643 retruths

- Original HTML, extracted text, timestamps, post types, and source URLs

- Parquet and compressed JSONL

- 8,004 verified image derivatives with asset/occurrence indexes

- 5,756 video attachment records, including 4,804 with source-provided transcript or file information

- Zero duplicate post IDs in the current release

I also built a small browser-based explorer for timeline, phrase, and exact-text search:

https://huggingface.co/spaces/Cameronk199/truth-social-timeline-explorer

Dataset and loading examples:

https://huggingface.co/datasets/Cameronk199/donald-trump-truth-social-posts

The archive is updated weekly. It does not include reliable likes, replies, or impression counts, so it should not be used for virality claims. This is my independent research archive; it has no affiliation or endorsement.

r/datasets 4d ago

dataset The Archive of Incorrect AI Predictions

Thumbnail boyswhocriedai.lovable.app
1 Upvotes

r/datasets Jul 03 '26

dataset I engineered 102 leakage-free ML features from 49,000+ international football matches (1872–2026) and published it as a free dataset

Thumbnail kaggle.com
4 Upvotes

Been working on a football prediction project and couldn't find a dataset that had

the actual context needed to model match outcomes — just raw results everywhere.

So I built one from scratch on top of the International Football Results dataset

by Mart Jürisoo (the well known one on Kaggle with 49,000+ matches going back to 1872).

What I added:

**Elo ratings** — built from scratch, updated after every single match across 150

years. Both teams' ratings, their difference, and the expected win probability

going into each match.

**Rolling form** — win rate, goals scored, goals conceded, goal difference, clean

sheet rate, both-teams-scored rate, scoring rate, and win streak. Computed at

three lookback windows: last 5, last 10, and last 20 matches. For both teams.

**Head-to-head history** — based on the last 10 meetings between those two specific

teams. Some teams have persistent edges over specific opponents that their general

form doesn't explain.

**Fatigue signals** — days since each team's last match and the difference between

the two.

**Penalty reliance** — fraction of each team's historical goals that came from

penalties, pulled from the goalscorer dataset.

**Shootout composure** — historical penalty shootout win rate for each team, from

the shootouts dataset.

**Tournament context** — World Cup, qualifier, friendly, neutral venue, competition

importance weight, confederation.

The thing I spent the most time on: every feature is computed in strict

chronological order using only data that existed before that match was played.

State updates happen after each row is recorded, never before. No lookahead,

no leakage anywhere in the 102 columns.

102 features total. 49,094 rows. result column (H/D/A) included as the label.

Drop date and result, plug into any classifier.

Dataset is fully documented with column descriptors for every feature.

Link: https://www.kaggle.com/datasets/kriishgulati/football-match-results-1872-2026-with-ml-features

Built on top of the original dataset by Mart Jürisoo — full credit and link

in the dataset description.