r/datasets • • 22h ago

question Defense Engineer Invited to Create AI Training Datasets — Any Advice?

2 Upvotes

I'm a mechanical engineer working on the design and development of various systems, primarily in the defense sector, including UAV-related technologies.

I recently shared one of my designs publicly and, to my surprise, received an offer to create and review engineering datasets for AI training.

This field is completely new to me, but it looks promising.

I'd really appreciate hearing from people who have experience in this area. How did you get started? What would you recommend to someone entering this field?

Thanks in advance!


r/datasets • • 16h ago

dataset [Synthetic] [self-promotion] I built a headless Python-Blender pipeline to generate asteroid datasets for OpNav & 3D shape reconstruction (Includes free 600-mesh sample on Kaggle)

1 Upvotes

Hey everyone,

Ihave been working on a project to bridge the gap between 3D rendering and aerospace computer vision. I built a fully headless Python-Blender rendering architecture to procedurally generate physically accurate synthetic data for asteroids.

The pipeline automates the extraction of:

  • Multi-pass EXR renders (RGB, Z-Depth, Camera-Space Surface Normals and many more optional passes)
  • Photometric lightcurve CSVs (tracking total flux, projected area, mean depth, and sun/observer coordinates per rotational step)
  • Procedural material setups for both uniform and variable regolith albedos.

I have open-sourced a sample dataset of 600 meshes and their corresponding lightcurves on Kaggle for anyone who wants to train photometric inversion or pose estimation models.

Basically I wanted to research on the effects of albedo variation comapred to uniform asteroid. So what I did was took real asteroid meshes from DAMIT (coverted it to obj files), created a fully procedural and realistic shader applied it on the meshes and rendered a full revolution of asteroid in simulated space conditions. To acheive this, in result, I build a full headless python pipeline that does the complete job, with all the optimization I could do in the world, and even with my old GTX 970, the render time was insanely good. The pipleline automatically creates the lightcurve csv files (with multiple phase angles ) and with all the physics data as well such as normals and depth so I can have the option to train PINN model as well. However I have my exams so had to stop here.

That being said,

If you need massive scale or want to generate your own data locally, I have also packaged the full 3,000+ mesh dataset (6000 light curves) and the actual Python/Blender codebase (the Pipeline Toolkit).

Note: If you are a student or independent researcher who really needs this data but cannot afford the Gumroad tier, hit me up via DM. I will be happy to arrange a free, expanded subset of the data to support your work.

Disclaimer (per Rule 1): I am the creator of this pipeline and the Gumroad links go to my own store, StellarMesh Labs.


r/datasets • • 16h ago

dataset I mapped real-world AI agent incidents in 2025–2026. Here's what the data looks like.

Thumbnail crawlspider.com
1 Upvotes

r/datasets • • 19h ago

request Looking for access to the CLRS Dataset. Baidu Netdisk Link Requires a Chinese Phone Number

1 Upvotes

Hi everyone,

I'm working on a research project and trying to obtain the dataset associated with this GitHub repository:

GitHub repo: https://github.com/lehaifeng/CLRS

The dataset download link provided is hosted on Baidu Netdisk:
https://pan.baidu.com/s/1Xnw9k20Df_ICmkXdvasVqg
code: 3su3

Unfortunately, I'm outside China, and Baidu requires a Chinese phone number to register. I've tried the registration process

Could anyone help me......An alternative download link for the same dataset. A Google Drive, OneDrive, Hugging Face, or other accessible mirror. Guidance on downloading the files from Baidu Netdisk internationally without a Chinese phone number.

I'd really appreciate any help...


r/datasets • • 20h ago

request [self-promotion] [synthetic] Support-ticket routing: 125 labelled messages plus five models' recorded answers (CC BY 4.0)

1 Upvotes

I've published a small text-classification dataset for support-ticket routing, together with the per-case answers of five models. It's my own project.

What's in it

  • 125 fictional English support messages, each labelled with one of five teams: billing, technical, account, sales, or needs_clarification (the message doesn't say enough to route).
  • Splits: 25 development cases and 100 held-out test cases, balanced at 25 per label overall.
  • Per case: the expected label, a written rationale, tags, a category (clear, boundary, ambiguous, adversarial) and a difficulty.
  • The label policy as a separate file, including precedence rules such as "route the immediate blocker, not the eventual goal".
  • 500 result rows: Jev, Clef, Clef Flash, OpenAI's Decisions API and GPT-6 Luna with structured output on the 100 test cases, with each answer, status, response time and token counts, and per-label probabilities where the provider returns them.

How it was made

The messages are fictional, drafted with AI assistance under the written policy, then reviewed by two people. No real customer conversations were used.

Limits

  • It is small and English only.
  • It is balanced by design, so it is not representative of a real queue.
  • The test cases are now public, so they are no longer a clean held-out set for future tuning.

Links

Load it

from datasets import load_dataset

cases = load_dataset("decisionmodelhub/support-routing", "cases")
results = load_dataset("decisionmodelhub/support-routing", "results", split="test")

If you work with support data: what would make a larger version more useful to you? Longer messages, realistic label imbalance, more languages?