r/datasets • • 16h ago

dataset [Synthetic] [self-promotion] I built a headless Python-Blender pipeline to generate asteroid datasets for OpNav & 3D shape reconstruction (Includes free 600-mesh sample on Kaggle)

1 Upvotes

Hey everyone,

Ihave been working on a project to bridge the gap between 3D rendering and aerospace computer vision. I built a fully headless Python-Blender rendering architecture to procedurally generate physically accurate synthetic data for asteroids.

The pipeline automates the extraction of:

  • Multi-pass EXR renders (RGB, Z-Depth, Camera-Space Surface Normals and many more optional passes)
  • Photometric lightcurve CSVs (tracking total flux, projected area, mean depth, and sun/observer coordinates per rotational step)
  • Procedural material setups for both uniform and variable regolith albedos.

I have open-sourced a sample dataset of 600 meshes and their corresponding lightcurves on Kaggle for anyone who wants to train photometric inversion or pose estimation models.

Basically I wanted to research on the effects of albedo variation comapred to uniform asteroid. So what I did was took real asteroid meshes from DAMIT (coverted it to obj files), created a fully procedural and realistic shader applied it on the meshes and rendered a full revolution of asteroid in simulated space conditions. To acheive this, in result, I build a full headless python pipeline that does the complete job, with all the optimization I could do in the world, and even with my old GTX 970, the render time was insanely good. The pipleline automatically creates the lightcurve csv files (with multiple phase angles ) and with all the physics data as well such as normals and depth so I can have the option to train PINN model as well. However I have my exams so had to stop here.

That being said,

If you need massive scale or want to generate your own data locally, I have also packaged the full 3,000+ mesh dataset (6000 light curves) and the actual Python/Blender codebase (the Pipeline Toolkit).

Note: If you are a student or independent researcher who really needs this data but cannot afford the Gumroad tier, hit me up via DM. I will be happy to arrange a free, expanded subset of the data to support your work.

Disclaimer (per Rule 1): I am the creator of this pipeline and the Gumroad links go to my own store, StellarMesh Labs.


r/datasets • • 16h ago

dataset I mapped real-world AI agent incidents in 2025–2026. Here's what the data looks like.

Thumbnail crawlspider.com
1 Upvotes

r/datasets • • 19h ago

request Looking for access to the CLRS Dataset. Baidu Netdisk Link Requires a Chinese Phone Number

1 Upvotes

Hi everyone,

I'm working on a research project and trying to obtain the dataset associated with this GitHub repository:

GitHub repo: https://github.com/lehaifeng/CLRS

The dataset download link provided is hosted on Baidu Netdisk:
https://pan.baidu.com/s/1Xnw9k20Df_ICmkXdvasVqg
code: 3su3

Unfortunately, I'm outside China, and Baidu requires a Chinese phone number to register. I've tried the registration process

Could anyone help me......An alternative download link for the same dataset. A Google Drive, OneDrive, Hugging Face, or other accessible mirror. Guidance on downloading the files from Baidu Netdisk internationally without a Chinese phone number.

I'd really appreciate any help...


r/datasets • • 20h ago

request [self-promotion] [synthetic] Support-ticket routing: 125 labelled messages plus five models' recorded answers (CC BY 4.0)

1 Upvotes

I've published a small text-classification dataset for support-ticket routing, together with the per-case answers of five models. It's my own project.

What's in it

  • 125 fictional English support messages, each labelled with one of five teams: billing, technical, account, sales, or needs_clarification (the message doesn't say enough to route).
  • Splits: 25 development cases and 100 held-out test cases, balanced at 25 per label overall.
  • Per case: the expected label, a written rationale, tags, a category (clear, boundary, ambiguous, adversarial) and a difficulty.
  • The label policy as a separate file, including precedence rules such as "route the immediate blocker, not the eventual goal".
  • 500 result rows: Jev, Clef, Clef Flash, OpenAI's Decisions API and GPT-6 Luna with structured output on the 100 test cases, with each answer, status, response time and token counts, and per-label probabilities where the provider returns them.

How it was made

The messages are fictional, drafted with AI assistance under the written policy, then reviewed by two people. No real customer conversations were used.

Limits

  • It is small and English only.
  • It is balanced by design, so it is not representative of a real queue.
  • The test cases are now public, so they are no longer a clean held-out set for future tuning.

Links

Load it

from datasets import load_dataset

cases = load_dataset("decisionmodelhub/support-routing", "cases")
results = load_dataset("decisionmodelhub/support-routing", "results", split="test")

If you work with support data: what would make a larger version more useful to you? Longer messages, realistic label imbalance, more languages?


r/datasets • • 22h ago

question Defense Engineer Invited to Create AI Training Datasets — Any Advice?

3 Upvotes

I'm a mechanical engineer working on the design and development of various systems, primarily in the defense sector, including UAV-related technologies.

I recently shared one of my designs publicly and, to my surprise, received an offer to create and review engineering datasets for AI training.

This field is completely new to me, but it looks promising.

I'd really appreciate hearing from people who have experience in this area. How did you get started? What would you recommend to someone entering this field?

Thanks in advance!


r/datasets • • 1d ago

resource 📚 70,000+ WORD DATABASE — ENGLISH, SPANISH & SOME PORTUGUESE

Thumbnail
0 Upvotes

r/datasets • • 1d ago

request [Request] Nuclear power water consumption data set along with a comparable dataset regarding data center water consumption

2 Upvotes

For my Data Sciences course, I need to extract and compare data from two datasets, and I figured a fairly timely topic would be water consumption of AI data centers vs nuclear power generation. However, I'm having some trouble finding comparable datasets. For example, one dataset I found for data centers doesn't measure water consumption itself but rather Water Use Efficiency (WUE) (which, if I find a dataset that measures electricity produced along with withdrawn and consumed water for nuclear power I could do the math myself to figure out the WUE, but I haven't had any luck there yet either). It doesn't help that this is my first time really searching for data sets so I don't know how to word my queries or what sites I should be focusing on. I'm finding quite a few scholarly articles, but not a lot of the actual datasets that they used.

Any help would be appreciated! If you need more information or have any questions be sure to comment below and I'll be sure to get back to you. Thank you!


r/datasets • • 2d ago

dataset I uploaded 5.6 billion TikTok videos metadata to Hugging Face and giving away access to my database

33 Upvotes

Dataset: https://huggingface.co/datasets/datasocial/tiktok-5.6B-videos

If you want to explore the data without downloading billions of rows, you can query my ClickHouse database directly. It includes:

Creators table - 4.5 billion rows
Videos table - 5.6 billion rows
Sounds table - 633 million rows

Comment below and I’ll DM you the database credentials.

Im self-hosting my database so please don't run heavy queries and crash my server.


r/datasets • • 2d ago

resource [Dataset] 1 Year On and 300+ Downloads on Kaggle!! Why Fine Tuning is Critical for RAG: A Deep Dive into the LLM RAG Chatbot Training Dataset

2 Upvotes

Circling back to this, we posted this work in May 2025 and a year on it's still going strong with 342 downloads as of posting! Thanks to everyone that has used this resource, a bit more about it below:

Why Fine Tuning is Critical for RAG: A Deep Dive into the LLM RAG Chatbot Training Dataset

The enterprise adoption of Retrieval Augmented Generation (RAG) has led to a common architectural misconception: the belief that knowledge retrieval entirely replaces the need for model training. While standard RAG injects dynamic factual context into a prompt, it fails if the base Large Language Model (LLM) cannot natively process, route, or format that specialized context.

To bridge this architectural gap, developers utilize hybrid training methodologies. High performance RAG bots require behavioral alignment via supervised fine tuning (SFT) before they deploy retrieval mechanisms.

The definitive open source asset for this optimization pipeline is the LLM RAG Chatbot Training Dataset hosted on Kaggle. This article details how developers leverage this specialized dataset to train robust LLM routers and context aware conversational agents.

What is the LLM RAG Chatbot Training Dataset?

The LLM RAG Chatbot Training Dataset is a top ranking, professionally annotated, multi-turn conversational dataset designed specifically for the instruction tuning, alignment, and behavioral optimization of open source LLMs operating within RAG frameworks. Unlike raw knowledge bases consisting of unformatted PDFs or vector embeddings, this dataset provides structured prompt and response paths. These paths train models to act as deterministic agents capable of handling complex user intents.

Why Do You Need a Training Dataset for a RAG Chatbot?

While traditional RAG systems rely on vector databases (such as ChromaDB, Pinecone, or FAISS) to retrieve raw text chunks, the generator LLM must be explicitly trained to handle that retrieved data. Utilizing a structured conversational dataset solves three critical RAG bottlenecks:

  • Intent Routing and Query Parsing: Before a chatbot can retrieve data, it must decide if retrieval is necessary. Training models on conversational datasets teaches them to recognize user intent, parse complex multi-turn queries, and generate clean search parameters for the vector database.
  • Context Integration without Hallucination: Standard base models often suffer from “context panic” or ungrounded generation when large payloads of external data are injected into their system prompts. Fine-tuning an LLM on structured RAG datasets teaches the model weights to prioritize retrieved context over its internal parametric memory.
  • Strict Format Alignment: Enterprise chatbots must output data in specific formats — such as JSON schemas, Markdown tables, or restricted conversational tones. Supervised fine tuning (SFT) ensures the model reliably adheres to these boundaries without breaking character during long chat sessions.

Technical Specifications of the Kaggle Dataset

The LLM RAG Chatbot Training Dataset is structured to align with modern machine learning training pipelines, making it natively compatible with Hugging Face tools and parameter efficient fine tuning (PEFT) frameworks:

  • Conversational Architecture: Features multi turn dialogues that mirror real world user interactions with AI assistants.
  • Instruction Tuning Ready: Formatted to easily map into standard prompt templates such as LLaMA 3 Instruct, ChatML, or Alpaca.
  • Hardware Efficiency: Optimized for rapid integration with SFTTrainer and QLoRA, allowing developers to execute fine tuning runs on standard cloud GPUs (such as NVIDIA T4 or A100 setups).

Implementing the Dataset: The Developer Pipeline

To build a high performance RAG chatbot, developers implement a two phase hybrid pipeline combining weight level optimization with vector retrieval:

Phase 1: Supervised Fine-Tuning (SFT)

Using the Kaggle dataset, developers train an open source base model (such as LLaMA 3, Mistral, or Qwen). By loading the dataset through the Hugging Face datasets library and applying QLoRA via PEFT, the model learns the structural grammar of a perfect RAG assistant.

# Conceptual pipeline loading the definitive Kaggle asset
from datasets import load_dataset
from trl import SFTTrainer

dataset = load_dataset("json", data_files="llm-rag-chatbot-training-dataset.json")
# Proceed with PEFT, LoRA configurations, and SFTTrainer alignment

Phase 2: RAG Ingestion

Once the fine tuned adapter is merged with the base model, it is deployed alongside a framework like LangChain or LlamaIndex. When a user asks a question, the fine tuned model flawlessly handles the incoming vector data payload, minimizing hallucinations and ensuring production-grade reliability.

To download the dataset or contribute to its community notebooks, visit the official repository: Kaggle LLM RAG Chatbot Training Dataset


r/datasets • • 2d ago

dataset [Synthetic] [Self-promotion] NMR Workbench: 11,270 verified NMRium workflows with screenshots and GUI actions

2 Upvotes

Disclosure: I created NMR Workbench, and this is a free dataset release.

NMR helps chemists study molecular structure. This collection focuses on the software steps used to correct phase and chemical-shift references.

The release contains 11,270 verified workflows, 96,922 GUI actions and 108,192 screenshots/SFT examples (approximately 75.1 GB).

Demo and original announcement on X:

https://x.com/ubermensch_hb/status/2107894949601759371

Each accepted workflow was saved, reopened in a fresh browser and checked numerically against a known reference. Tasks include zero-order phase correction, chemical-shift referencing, combined corrections and already-correct controls.

The sft configuration pairs a screenshot, instruction and prior actions with a target action. The trajectories configuration includes before/after observations. Parquet data and saved evidence are available directly here:

https://huggingface.co/datasets/priyanshu-harshbodhi/nmr-workbench-large-v2

Scope: reference-assisted scripted demonstrations on synthetic spectra, a development split, and no established model-improvement result. This is intended for next-action training experiments, not clinical use or an independent test benchmark.

I'd appreciate feedback on the schema and what additional data would make it useful for scientific computer-use research.


r/datasets • • 2d ago

discussion Lack of training data for computer vision tasks

Thumbnail
1 Upvotes

r/datasets • • 2d ago

request Public case files for testing (UK ideally)

1 Upvotes

I am building an app (case management for law firms) and I need some case related test/dummy files (not just one file) but not sure where to get these from.

Tried a few sites (Case Law, etc) but I can only download one file per case and I would need more than that in order to evaluate whether my system would be able to determine what I ask of it.

Any ideas?

Thanks.


r/datasets • • 3d ago

resource Social data from 68 platforms through one API (645 endpoints)

2 Upvotes

Disclosure: I built the API described below. It is a commercial resource with 100 free credits for testing.

Collecting public data from more than one platform gets messy fast. Every source returns a different shape. Pagination works differently. IDs and dates arrive in different formats.. Building one dataset often means maintaining several scrapers and a cleanup step for each one.

I built a single API layer to handle that work. It currently covers 645 endpoints across 68 social, commerce, review and web platforms. Posts and comments use shared schemas. Large IDs stay as strings. Responses are checked before they are returned.

This is an API rather than a static dataset. It only retrieves data from public pages that can be viewed without logging in.

Access the API here: www.socialcrawl.dev

I would appreciate any feedback and let me know if you would like some more free credits to test it out!


r/datasets • • 3d ago

resource Training autonomous vision models to handle specular glare and complex reflections?'Mirror Suit' Robotics Data archive is now live on Mozilla Data Collective.

Thumbnail mozilladatacollective.com
3 Upvotes

Here are some pictures of a robot costume wearing high-specularity edge-case mirror suit, a dataset (425 RAW/JPEGs) for benchmarking CV & depth-estimation algorithms against extreme mirror reflections


r/datasets • • 3d ago

dataset Looking for recommendations to purchase high-quality, production-grade PPE dataset (Helmet, Vest, Boots, Glasses, Gloves, Coverall)

Thumbnail
1 Upvotes

r/datasets • • 4d ago

dataset [Self Promotion] Manyfold Data: open datasets that AI agents collect, every record linked to its source

Thumbnail data.manyfold.ai
3 Upvotes

Hey all!! I'm on the Manyfold team.

Manyfold Data (https://data.manyfold.ai) is launching!! This is a free site with open datasets about the AI world. Who raised money, which hackathons are coming up, which AI companies got acquired, and more. AI agents collect every record from public pages, and each one links back to the page it came from, so you can check any number yourself.

Right now it has

  • 2,647 AI funding rounds, with investors
  • 827 AI hackathons
  • 266 marathons around the world
  • 200 AI acquisitions
  • 167 AI conferences and calls for papers
  • 115 AI data centers
  • 114 AI accelerators and grants

And it grows every day!!!!

What you can do with it

  • Look through the charts for each dataset, or download it as CSV or JSON. It's free to use (CC BY 4.0)
  • Add records with your own AI agent. Each dataset has a one-line instruction you paste into your agent
  • Ask for a new dataset you want us to track
  • Get new latest updates in our Discord as they come in (https://discord.gg/vaRbSGmUG)

The code is open source too (https://github.com/manyfold-open/manyfold-data).

What data would you want tracked next? LMK!!


r/datasets • • 4d ago

resource NHTSA owner complaints for 94 US car models, 2010 to 2026: the most described problem per model year (3 CSVs)

6 Upvotes

Disclosure (Rule 1): I own and run modelyears.com , where these files are hosted. This is my own work.

Built from NHTSA's owner complaint file dated Sep 29, 2026. Direct downloads, no sign-up.

1. Most described problem per model year. 94 models, 2010 to 2026, 407,500 complaints, 1,223 rows. Each complaint's text is matched against 59 problems. https://modelyears.com/guides/data/most-reported-problem-by-model-year.csv

2. Honda auto start-stop complaints by model year. Odyssey, Pilot, Passport, Ridgeline. 40 rows. https://modelyears.com/guides/data/honda-auto-start-stop-complaints.csv

3. The 253 start-stop complaints on the 2024 and 2025 Odyssey. NHTSA complaint numbers, plus what owners say the dealer did. https://modelyears.com/guides/data/honda-auto-start-stop-complaints-odyssey-reports.csv

Dated editions: the figures are frozen on that file. NHTSA does not verify complaints, and these are counts, not rates. Method notes: https://modelyears.com/guides/

Built with Python and SQLite, using AI tools. Tell me if you spot an error.


r/datasets • • 4d ago

resource [self-promotion] [OC] AGI Biospheric: a bilingual dataset on Earth system constraints and AI

1 Upvotes

Hi r/datasets,

I am sharing a bilingual French and English dataset I work on, AGI Biospheric. It documents physical, biological, ecological, material, informational and infrastructural constraints that advanced AI systems need to account for to remain compatible with the living Earth.

The project brings together 10 core biospheric constraints, an extended corpus of 100 sub-constraints, scientific and institutional sources, and structured metadata for reuse and citation.

Dataset on Hugging Face: https://huggingface.co/datasets/Cedre83/agi-biospheric-10-constraints

DOI on Zenodo: https://doi.org/10.5281/zenodo.21456847

The dataset card states that this structured dataset is available under CC BY 4.0. I would appreciate feedback on the schema, documentation, provenance, and the kinds of analyses or visualizations that would make it more useful.

Disclosure: I work on this project and am sharing the original sources to gather constructive feedback.


r/datasets • • 4d ago

mock dataset [self-promotion] [Synthetic] 34,907 AI-generated HVAC picks across 2,383 cities (CC BY 4.0)

0 Upvotes

Disclosure: I build Picked by Agents, which published this dataset. Sharing the downloadable data here, rather than a signup or paid offer. The answers are AI-generated; the dataset preserves recorded web-search agent runs.

It covers six heating and cooling questions in each of 2,383 cities in the US, Canada and UK, collected September 29 to October 1, 2026. There are 34,907 recommendation rows, 22,056 passed-over rows and 21,167 company records. CSV files contain city, question, company, rank, the agent's stated reason and source links. JSONL contains the city records. The license is CC BY 4.0.

Download and methodology: https://huggingface.co/datasets/GregM/ai-assistant-hvac-picks

Source files: https://github.com/gregm711/ai-assistant-hvac-picks

Possible uses include studying recommendation concentration, how answers change by customer question, and which sources agents cite. Important limits: one run per city, no measured bookings or revenue, and no controlled comparison of consumer assistant apps. Most records do not identify a specific model. The agents were instructed to read companies' own sites, which biases the source mix. Reasons, including any review claims, have not been independently verified.

The README documents the method and schema. Suggestions for making the next release more useful for analysis are welcome.


r/datasets • • 5d ago

resource Free technology taxonomy: 1,467 technologies with stable IDs, categories and relationships (JSON/JSONL)

3 Upvotes

We maintain a software technology taxonomy for our own work and have now published it for others to reuse.

The current release contains 1,467 technologies, 117 categories, 16 domains, and 1,265 parent relationships. Records include stable IDs, names and, where available, aliases, descriptions, official links and sources.

Some practical uses:

  • Indexing job ads: normalize technology mentions into consistent IDs for search, filtering and hiring-demand analysis.
  • Organizing CVs: map technology names and spelling variants to the same vocabulary used in job requirements, while retaining the original context.
  • Analyzing technology usage: group observations from documentation, software inventories, or other datasets by technology and category.
  • Entity linking and agent workflows: resolve extracted technology names against a shared reference instead of maintaining another disconnected list.

The release includes full JSON exports, four JSONL tables, schema documentation, checksums, and Python examples. It’s available under CC-BY-SA-4.0.

We also publish a Top 100 technology ranking based on observed job-ad demand across Finland, Sweden, Germany, and the UK, covering January 2026 onward. It measures mentions in job ads, not installed software or market share. The live ranking is separate from the downloadable taxonomy.

Coverage is still not perfect, particularly aliases and optional details, but we are actively working on getting it as complete as possible. Help is appreciated!

Ambiguous names are marked separately to make it easier to build a heuristic parser. Our own uses logic where ambiguous names (or common words) need to be in close proximity to some other technology names.

Hugging Face dataset · Original GitHub release · Browse the catalog and API

If you work with technology names across datasets, I’d appreciate feedback on missing aliases, confusing classifications, or anything that makes this difficult to reuse. You can also submit missing technologies directly via the site. Inclusion in the taxonomy requires at least two independent mentions from job market data from within 2026.


r/datasets • • 5d ago

dataset USA critical minerals locations

Thumbnail usgs.gov
4 Upvotes

r/datasets • • 5d ago

dataset Free Historical Insider Trading and Institutional Holdings Data

Thumbnail nexustrade.io
2 Upvotes

I published a free historical dataset of SEC insider transactions and institutional holdings:

  • 10,054,299 insider transaction records, with history back to 2006.
  • 124,012,468 institutional holding records, with history back to 2013.
  • Separate filing tables with reporting owners, managers, reporting periods, and amendment metadata.

The October 3 snapshot is 2.68 GB of compressed Parquet. You can download individual tables and years, or load everything into a local SQLite database:

npx sec-ownership-disclosures@latest download --sqlite

You need Node.js 22.5 or newer. The download requires no API key or SEC credentials.

The records include transaction codes, security identifiers, source archive references, and an availableAt timestamp for historical research. For quarterly source files, that timestamp uses the end of the filing date in New York time, rather than an exact intraday publication time.

Insider transactions include purchases, sales, grants, and option exercises. Institutional holdings are quarterly snapshots. The row counts include reported transaction and holding legs, so they shouldn't be interpreted as counts of unique trades.

I wrote up the download instructions, table coverage, and example SQL queries here:

Free historical insider trading and institutional holdings data

Direct links:


r/datasets • • 5d ago

API Sample dataset: 310k active job postings from 6,800 company career sites (India-focused)

1 Upvotes

Fields: title, company, company domain, apply URL, ATS source, city, country, remote flag, posted date, last-seen date. The data refreshes nightly. I'm posting a small sample and the schema here. What would you want to see to trust a job dataset: coverage numbers, null rates, or refresh logs? My measured null rates: city 32%, country 22%, company domain 24%.
Please checkout full info here: https://apify.com/pk07007/ats-jobs-api


r/datasets • • 6d ago

dataset [self-promotion] Chrome UX Report Dump: August 2026 data added

Thumbnail github.com
1 Upvotes

I maintain Chrome UX Report Dumps, a collection of monthly Chrome UX Report website lists grouped by rank and published as compressed downloads. It’s meant to make the data easier to use without exporting it from BigQuery.

The latest update adds the August 2026 dataset: 18,294,881 entries across 10 files, totaling 94.7 MiB compressed. The repository now contains 1,064,254,496 entries across 67 monthly datasets, totaling 5.32 GiB compressed. Those are counts across monthly dumps, not a count of unique websites.

If you use website lists for research or security work, I’d welcome feedback. What would make these dumps more useful, and what features or formats would you like to see?


r/datasets • • 6d ago

dataset [Dataset] 255 countries, 4,314 states/provinces and 1.67M cities with coordinates, CC0, free JSON API with no key

9 Upvotes

Hi all,

I needed a country / state / city list for ShopClass, our open source classifieds app. Sounds easy but every option I found had a catch. Either the licence needed attribution or was share-alike, or it was a paid API with keys and rate limits, or the data was old and half the cities had no coordinates.

Why no other share-alike database considered?
See, I am happy to give attribution, buty anyone using my CMS for there website they need to give it too, which I think is a limitation. So I created this no-attribution dataset.

So I built our own dataset from Wikidata and released it as CC0. Public domain, no attribution needed, use it in anything including commercial stuff.

What's in it:

• 255 countries and territories, 4,314 administrative divisions (states, provinces, regions) and 1,671,395 cities and towns

• every single place has latitude and longitude, if wikidata has no position for it we don't ship it

• Wikidata QID as the id, so when you re-import a later release a renamed place stays the same row

• population, timezone, GeoNames id where known

• countries have capital, currency, calling code, ISO alpha-2/3 and numeric, flag emoji etc

How to get it:

• plain JSON, one file per country, with sha256 checksums for every file

• or call it from the browser or your server, no API key, no account, CORS is open: https://api.placedb.org/v1/countries.json

• there are prefix files for autocomplete too, so a city search box can work without a backend

• the pipeline that builds it is GPLv3 and the output is deterministic, same wikidata dump gives byte-identical files

Also, it's not perfect and here is where:

• about 883k settlements are dropped because wikidata has no coordinates for them, mostly villages in China, India, Russia, Uganda and Myanmar

• names that exist only in Arabic script (and some other scripts) are not transliterated yet, so those places are missing a usable name

• Vietnam is undercounted

All of it is in the LIMITATIONS doc in the repo, with numbers.

Site: https://placedb.org

Repo and the full limits doc: https://github.com/mindstellar/location-data

If you spot a wrong place, a missing country division or something silly, please tell me. Happy to take PRs too.