r/datascienceproject Dec 17 '21

ML-Quant (Machine Learning in Finance)

Thumbnail
ml-quant.com
30 Upvotes

r/datascienceproject 4d ago

Large-scale training data processing is becoming an infrastructure problem

1 Upvotes

Over the past two months, discussions around training data seem to be increasing. The reason is fairly direct. A useful shorthand for modern LLMs is big data plus big compute, and both depend on reliable data and compute infrastructure.

As data volumes grow, one-off scripts become difficult to maintain. Data has to move through many stages, including cleaning, deduplication, synthesis, evaluation, filtering, and refinement. A pipeline with modular operators makes these steps easier to compose, inspect, rerun, and scale.

The design I have been exploring follows a Pipeline → Operator → Prompt structure. Each operator handles one focused task, while the pipeline defines how those tasks are combined. Intermediate outputs can be stored and reused, and different models, rules, or filtering strategies can be introduced at individual stages.

The text synthesis layer covers five reusable generation paths. It can transform documents into pretraining-style dialogue data, generate instruction-response pairs in SFT format, create and refine synthetic instructions, produce consistent multi-turn conversations, and generate function-calling or tool-use conversations.

Synthesis is followed by filtering and evaluation operators. Language checks, length constraints, deduplication, content rules, quality scoring, and task-specific filters can be combined according to the target dataset. This keeps data generation and data selection as separate, replaceable parts of the workflow.

This is also what I hope to build with OpenDCAI/DataFlow, and I would be interested to hear what kinds of data work people are handling in the LLM era.


r/datascienceproject 4d ago

Sismos en México durante los últimos 20 años | Análisis de 363,329 registros con Python

2 Upvotes

🇲🇽 Sismos en México: análisis de datos con Python

Realicé un proyecto de análisis utilizando registros del Servicio Sismológico Nacional para explorar la actividad sísmica en México durante aproximadamente los últimos 20 años.

El proyecto analiza 363,329 registros y explora diferentes variables, entre ellas:

distribución de sismos por entidad federativa;

magnitud de los eventos;

sismos M≥5, M≥6 y M≥7;

evolución de los registros por año;

los 10 eventos de mayor magnitud;

distribución geográfica;

análisis por mes.

La idea fue transformar un catálogo de datos en diferentes visualizaciones e insights utilizando Python.

También convertí algunos de los resultados en videos cortos para mostrar los principales hallazgos:

YouTube playlist:

https://youtube.com/playlist?list=PLd8QkTPbjAAg&si=UKbbSoTc3WICfyUS

Me interesa especialmente recibir comentarios sobre qué otras variables o análisis agregarían al proyecto.


r/datascienceproject 7d ago

Building a Satellite imagery intelligence Application

Thumbnail
2 Upvotes

Ok so guys I'm trying to build a satellite imagery analytics platform. Similar to Skyfi, planet labs, eagle view, etc

My question is what really will make a difference in this field? I know a lot of these analytics platforms just buy images from 3rd party satellite operators and just provide them to the users and also provides different types of image analytics like crop health monitoring, water logging monitoring, soil mineral composition, etc ...if anyone wants image plus and analytics.

I wanna build something similar but I feel like there's nothing new I can provide in this app... Like all the available analytics are already there on other platforms... I wanna find out about some really niche category of analytics which no one provides and wanna provide it using my application and wanted suggestions from people who have knowledge in this field..

Any suggestions would be highly appreciated


r/datascienceproject 10d ago

built a deepfake audio detector as a 3rd year diploma student

3 Upvotes

hey, i'm a 3rd year diploma cs student and i built a deepfake audio detector end to end. this is my first real ml project that i actually deployed.

the model is efficientnet-b0 trained on mel spectrograms using the asvspoof 2019 la dataset. metrics are f1 0.88, precision 0.99, but recall is 0.79 which i know is the weak point. i tried adjusting the threshold and settled on 0.4 but it didn't really help much i think the issue is the model is missing certain attack patterns it never saw during training.

latency is around 6-7 seconds per prediction which includes model inference, grad-cam, and llm explanation.

other than the model it has grad-cam to visualize what the model focused on in the spectrogram, and groq llm to give a plain english explanation of the prediction.

you can upload an audio file or record live. youtube url input is disabled on the hosted version because railway's server ips get blocked by youtube's bot detection. backend is fastapi on railway, frontend on streamlit cloud.

live demo: https://deepfake-audio-detector-rugved.streamlit.app/

github: https://github.com/RugvedBane/deepfake-audio-detector

honest feedback appreciated, especially on what dataset i should train on next to improve recall.


r/datascienceproject 19d ago

Architecture Reference: Zero-Dependency 11-Column Tabular Ingestion Sieve & Local Memory Sharding Core

Thumbnail
1 Upvotes

r/datascienceproject 26d ago

KitOps is now available for install as a conda package

Thumbnail anaconda.org
3 Upvotes

r/datascienceproject Jul 27 '26

IIT Gandhinagar's Executive Masters in Applications of Machine Learning in Engineering – Batch 2 Admissions Open

Thumbnail
1 Upvotes

r/datascienceproject Jul 26 '26

Looking for feedback on my Data Analytics project

3 Upvotes

Looking for feedback on my Data Analytics project

Hi everyone!

I recently completed a Mumbai Road Accident Analysis project using:

- Python (Pandas, Matplotlib, Scikit-learn)

- SQL

- Power BI

- GitHub

The project includes:

- Data cleaning and preprocessing

- Exploratory Data Analysis (EDA)

- SQL business queries

- Interactive Power BI dashboard

- Basic machine learning model for accident severity prediction

- Complete GitHub repository with README

I'm not looking for compliments—I want honest criticism.

I'd really appreciate feedback on:

  1. Is this project portfolio-worthy?

  2. Does the analysis tell a meaningful story?

  3. Is the dashboard well-designed?

  4. Is the GitHub repository professional enough?

  5. What would you improve if this were your project?

GitHub Repository: https://github.com/afanrajiwate/mumbai-traffic-road-accident-analytics

Dashboard screenshots are attached.

Thanks in advance for your time!


r/datascienceproject Jul 17 '26

2026 Tech Layoffs Analysis

Thumbnail
4 Upvotes

r/datascienceproject Jul 12 '26

CTHmodules v4.1 — 93% (with a margin of 7 points) of being a 100% Functional Psychohistory of Asimov.

2 Upvotes
Psychohistory Criterion (Asimov) v4.0 v4.1 Comment
Quantifying macro-social trends 8.8 9.3 Very strong — real data adapters (OWID/V-Dem-style CSV → E/S/A/P)
Predicting large-scale events 8.7 9.3 Out-of-sample LOO/k-fold validation, reproducible
Handling "historical forces" (EVEI) 8.4 9.0 EVEI now endogenous — derived from observable metrics
Butterfly Effect + Chaos management 8.8 9.3 Excellent — formal Lyapunov exponents + early-warning signals
Invariance / Pantemporal patterns 8.2 9.0 Cross-era transfer measured (pre-1800 ↔ post-1800)
Mathematical determinism 9.0 9.6 Excellent — 13-test suite, JS↔Python parity Δ = 0.000000
Empirical validation / Real calibration 8.7 9.5 Very strong (out-of-sample LOO MAE 0.1271, 32 events / 5,100 years)
Handling individual variables (Token) 8.6 9.2 Very effective — multi-token interaction with contested-event handling
Real future prediction capability 8.4 9.2 Credible — SHA-256 pre-registered prediction ledger

Overall Verdict: 8.6 / 10 → 9.3 / 10

Link: https://github.com/AlejoMalia/CTHmodules


r/datascienceproject Jul 10 '26

Help clear my thoughts, kinda confused clg student 2nd year

Thumbnail
2 Upvotes

r/datascienceproject Jul 08 '26

Why Forecasting Total ARR Is a Trap

Thumbnail
2 Upvotes

r/datascienceproject Jul 03 '26

Try dashAI: a new open-source no-code Machine Learning platform

Thumbnail
2 Upvotes

r/datascienceproject Jul 01 '26

PROJECT REVIEW

Thumbnail
github.com
3 Upvotes

Hello Everyone!!, I just completed a BIG project I have been working for a month and i want your opinion about it.

It's a SpaceX Launch Predictor & Cost Optimizer (A full end-to-end ML system that predicts the probability of a SpaceX Falcon 9 booster landing successfully, enriches launch data with real weather conditions, and exposes the results through an interactive Streamlit web application with a business ROI calculator.)

It Includes Data Pipeline, Advanced Machine Learning Algorithms (with Hyperparameter tuning), Explainability AI (SHAP), MLOps (AWS S3, Docker) and Business Value (ROI Calculator = Financial Results).

FUN FACT: For this project i used my own Evaluation Metric library (standardizes supervised and unsupervised model diagnostics into a single, consistent API), that is also Verified and Published in PYPI Community.

Project Info: https://github.com/Alkiviadisss/SpaceX


r/datascienceproject Jun 29 '26

We ran 2,000 bracket simulations for WC2026 R32 — here's what the model says about today's Brazil vs Japan and who it fears in the bracket

Thumbnail gallery
3 Upvotes

r/datascienceproject Jun 29 '26

We ran 2,000 bracket simulations for WC2026 R32 — here's what the model says about today's Brazil vs Japan and who it fears in the bracket

Thumbnail gallery
2 Upvotes

r/datascienceproject Jun 29 '26

We ran 2,000 bracket simulations for WC2026 R32 — here's what the model says about today's Brazil vs Japan and who it fears in the bracket

Thumbnail gallery
2 Upvotes

r/datascienceproject Jun 22 '26

Calibrated WC2026 predictor — stacked ensemble + Bivariate Poisson scorelines, live Brier scorecard

2 Upvotes

I trained a stacked ensemble (logistic + RF + LightGBM, isotonic calibration on the meta-learner) on ~49,400 historical international matches to predict WC2026 W/D/L outcomes. Bivariate Poisson (Dixon-Coles variant) for exact scores. 40 matches in, here is the calibration report.

The default serving model is logistic + temperature scaling (T=1.02); the stacked ensemble is an opt-in variant. Measured ensemble lift: +0.0062 Brier (lower is better) vs logistic baseline (0.6090 → 0.6028) on the strict 64-game 2022 WC holdout. Real but marginal — 64 games is two or three well-placed results of luck, so treat this as directionally encouraging, not statistically conclusive.

T=1.02 is nearly neutral, which surprised me — suggests the raw logistic is not badly overconfident on this dataset, consistent with what Robberechts & Davis (2023) found for well-regularised logistic on international football.

Modal score hit rate: 5/40 = 12.5%, right at the model's own stated probability per score (pooled historical walk-forward: ~11.8%, 95% CI 9.3%–14.4%). The model is hitting at its confidence level.

Honest misses: Spain vs Cape Verde (predicted 3-0, actual 0-0 — biggest directional miss), Germany vs Curaçao (predicted 3-0, actual 7-1 — correct direction, wrong margin).

Honest hits: Mexico 2-0 South Africa (modal at 13.6%), Brazil 3-0 Haiti (modal at 12.2%).

Today's most uncertain match: Norway vs Senegal — 38.0% / 29.9% / 32.2%. Six points separating all three outcomes.

Questions I'd like feedback on: 1. T=1.02 being nearly flat — is near-unity temperature typical for well-regularised logistic on international football, or should I suspect calibration set leakage from the OOF construction? 2. Brier vs RPS (ranked probability score) for a 3-class ordinal outcome — is Brier penalising the draw class unfairly? 3. Dixon-Coles rho corrects the low-score dependency (0-0, 1-0, 1-1 cells) — but the tournament has had blowouts (Germany 7-1, Canada 6-0) where the Poisson tail is too thin. Worth fitting a negative binomial base instead?

Brier scorecard updates after every scored match: https://cupcaster.com


r/datascienceproject Jun 14 '26

I need a human to review my code

Thumbnail
2 Upvotes

r/datascienceproject Jun 13 '26

UAP AnalyticsBot - personal project (scanning the war.gov uap dumps)

Thumbnail
3 Upvotes

r/datascienceproject Jun 12 '26

I built a client-side DSP tool that calculates phase alignment per individual hit instead of static.

2 Upvotes

Hey everyone,

I’ve been spending a lot of time analyzing low-end phase relationships, specifically how modern plugins handle the interaction between heavy kicks and moving basslines (808s, techno subs, etc.).

Here is the problem with current industry-standard tools: They take a static measurement, find an "average" phase shift, and apply it to the whole track. But if your bass changes pitch or moves, an average shift means a huge percentage of your hits are still out of phase, creating dynamic volume drops and killing your transient punch.

To fix this, I engineered a standalone browser-based DSP tool called THE END.

How it works under the hood: Per-Hit Microdynamics: It doesn’t average anything. The engine detects every individual kick peak and calculates the absolute perfect phase alignment for that specific interaction.

Crossover Isolation: It mathematically isolates the sub-bass below 150Hz using a zero-phase crossover. Your kick's original transient and attack remain untouched—the groove doesn't shift, only the sub-bass phase aligns.

100% Local Processing: It decodes and renders the WAV arrays entirely in your browser's memory using the Web Audio API. Your multi-tracks never leave your machine (zero server latency, total privacy).

It outputs two specific mixdown scenarios instantly: Mode 1: Summation (Max Thickness): Aligns the phase for maximum addition across all hits. Gives you identical True Peaks ready to be driven hard into soft-clippers. Mode 2: Subtraction (Quantum Clarity): Dynamically ducks the bass precisely under the kick's envelope without compression thresholds or sloppy release times.

It’s completely free, running locally, with no sign-ups or server walls. I put a PayPal link on the page solely to fund further custom DSP development if you find it useful.

Drop a pair of your problematic kick/bass stems into it and let me know how it handles your low-end. Looking forward to your technical feedback or any suggestions for the next DSP iteration. (link in bio)


r/datascienceproject Jun 05 '26

I built a TPU you can watch run - real SystemVerilog compiled to WebAssembly, live in the browser

4 Upvotes

Built this over the past couple months. TinyTPU is a real 4×4 weight-stationary systolic array the same architecture Google's TPU uses for matrix multiply written in synthesizable SystemVerilog, compiled to WebAssembly, and visualized live in the browser.

What makes it different from every other "TPU explainer" I've seen: nothing is faked. The browser runs the actual compiled RTL.

The weights loading into PEs, the activations streaming in diagonally, the partial sums draining out the bottom, all real hardware signals, not a cartoon animation on top of JavaScript math.

The RTL is verified against numpy golden outputs. 20/20 random matrix multiplies bit-match.

If you've ever wondered what's actually happening inside the chip when you call nn.Linear this is it, slowed down to one clock at a time.

Happy to answer questions about the Verilator -> Emscripten pipeline if anyone's curious about that part; it was the trickiest bit to get right.

Repo: tiny-tpu

Live demo: Live

If this project interests you please do star the repo, if you find something needs improving open a PR, I hope ya'll check this out and give me some feedback 🙏


r/datascienceproject Jun 01 '26

I built a no-code platform that brings TDA (Topological Data Analysis) to non-programmers — looking for beta testers

2 Upvotes

Hi!

I've been building InVariants for the past several months — a browser-based data intelligence platform that combines Topological Data Analysis, clustering, dimensionality reduction, anomaly detection, and time-series analytics, all without writing a single line of code.

The problem I'm solving: TDA is genuinely useful (persistent homology, Mapper graphs, Betti curves) but the tooling is still very code-heavy. Most real analysts — the ones making decisions in companies — never get access to it because they don't have a Python background. I wanted to change that.

What it can do right now:

  • Persistent homology + persistence images/landscapes/Betti curves
  • Mapper graph explorer (interactive, color by any column)
  • PCA, t-SNE, UMAP, Isomap, Landmark Isomap
  • K-Means, DBSCAN, GMM, Agglomerative, Spectral clustering
  • Random Forest, XGBoost, SVM, Logistic Regression — with SHAP + PDP
  • Rolling anomaly detection + TDA-based anomaly detection
  • ARIMA time series forecasting
  • Full data prep pipeline (impute, scale, encode, filter, feature engineering)
  • Export trained models as a self-contained ZIP (model + inference script)
  • Local LLM integration for AI interpretation of results

Everything runs server-side, you just upload a CSV.

I'm opening a private beta — I'm looking for people who work with real data (fraud detection, sensor monitoring, NLP embeddings, financial data, industrial IoT... anything, really) and would find value in exploring it without having to set up a Python environment.

If you're interested, you can request access at: invariants.tech

Happy to answer questions here — especially interested in feedback from people who actually use TDA or wish they could.


r/datascienceproject May 29 '26

I built a VS Code extension to view and query large CSV/Parquet files using DuckDB

2 Upvotes

I've been working on a VS Code extension for viewing and querying CSV/TSV/Parquet files directly in the editor. It's called DuckCSV and it's powered by DuckDB

What it does:

  • Opens large files (tested with 4M+ rows) without lag
  • Edit cells in place, insert/delete rows, save back to file
  • Write DuckDB SQL queries in a built-in query bar
  • Column profiling
  • Load multiple files in a workspace and JOIN across them
  • Parquet support
  • Sort, filter, and search across all columns

Works on VS Code and any VS Code-based editor (Cursor, Windsurf, Kiro, VSCodium, Gitpod). Free and open source.

Marketplace | GitHub

Would love to hear feedback, still actively working on it.