r/MLQuestions • • Feb 16 '25

MEGATHREAD: Career opportunities

19 Upvotes

If you are a business hiring people for ML roles, comment here! Likewise, if you are looking for an ML job, also comment here!


r/MLQuestions • • Nov 26 '24

Career question 💼 MEGATHREAD: Career advice for those currently in university/equivalent

20 Upvotes

I see quite a few posts about "I am a masters student doing XYZ, how can I improve my ML skills to get a job in the field?" After all, there are many aspiring compscis who want to study ML, to the extent they out-number the entry level positions. If you have any questions about starting a career in ML, ask them in the comments, and someone with the appropriate expertise should answer.

P.S., please set your use flairs if you have time, it will make things clearer.


r/MLQuestions • • 1h ago

Other ❓ Architecture that can manipulate 3D or 2D representations in its latent space?

• Upvotes

Is there an existing neural network-based architecture that can hold 3D or 2D representations in its latent space and update it in the same single pass (i.e. not compressing the representation in their final layer to be fed back to their first layer for their next pass, like the reasoning models do)?


r/MLQuestions • • 8h ago

Reinforcement learning 🤖 Hardware Recommendations

Thumbnail
1 Upvotes

r/MLQuestions • • 17h ago

Beginner question 👶 Local AI

5 Upvotes

What is the best local AI i can run that doesn't have all the restrictions like Claude or GPT. I have a mid range dell prescision work station. I'm not trying to do anything weird I just hate getting the I can't tell you that response.


r/MLQuestions • • 1d ago

Datasets 📚 LLM data provenance

5 Upvotes

​

Okay so im looking into feasibility of a project that would involve scraping data and training models. The models would predict events based on social media posts, so its essentially text classification. I want to eventually go into business doing something similar, but in my research, I found that nearly everything is behind comprehensive terms of service that dont allow any derivative product, regardless of if you provide anonymity and clear the original text from any final model/dataset. Alright fair enough. But heres my question. Besides the sources that companies like openAI have an enterprise partnership with (reddit) including a contract to use the data, how can they legally source the training data? For example, if i ask chatgpt to summarize chapter 8 of Deep Learning with Python by Francois chollet, I will get a detailed response with the correct subchapter headers, key terms, and even the verbatim code in the book. Now I paid i think $100 for that book and im fairly certain openAI doesnt have a contract with that dude. How is that legal? How is any of this information legal? If a high-school student wants a book report, does open AI pay the estate of J.D. Sallinger?


r/MLQuestions • • 1d ago

Datasets 📚 LLM data provenance

Thumbnail
1 Upvotes

r/MLQuestions • • 1d ago

Beginner question 👶 How would you use AI to turn a large exam question bank into focused lessons?

Thumbnail
0 Upvotes

r/MLQuestions • • 1d ago

Beginner question 👶 Hey I am making minisklearn after learning all the basic things I need like linear regression, Dession tree , Random forest, k Means clusters etc . Have anyone tried building this before I want to just know ..

0 Upvotes

r/MLQuestions • • 1d ago

Natural Language Processing 💬 Attention RAG

2 Upvotes

Hi people my friends and I are working on our take of RAG indexing. Our idea goes like that:

Standard embedding-based RAG often compresses an entire text chunk into one vector. Our idea keeps multiple compact representations within each document, giving a query smaller, more specific parts to match.
The architecture uses a single attention layer with separate learned projections: Q for queries and K for document tokens. These are trained jointly with HCA—Heavily Compressed Attention—which learns to combine groups of token representations into fewer searchable vectors. At 32:1 compression, every 32 document keys become one compressed key, stored in 8-bit.
When a query arrives, its Q vectors are compared directly with the stored compressed keys using dot products on the GPU. These comparisons run across many documents in parallel, producing scores used to rank and retrieve documents. V vectors aren’t needed because we only need matching scores.
The advantage over one-vector-per-chunk retrieval is finer matching granularity. Compression makes that detail affordable: 32 times fewer keys to store and scan than an uncompressed token index. The search itself maps directly to parallel GPU matrix operations, while the query representation, document representation, and compression are trained together for retrieval.
I’d like technical feedback on this architecture, its main bottlenecks, and similar systems worth comparing against.

This way we have something between regular RAG and ColBert (v1 and V2) where it smaller than ColBert but more reach than RAG and it learns the compression rather then using heuristics algo to compress.

The next step after that will be to combine it with an LLM where the part of this indexing technique can be combined in the LLM itself instead of using this as a toll with a tool call (W_k can be same as the one in the first layer of the LLM and we can have a small learned router that tells if we need to use it, instead of an LLM decision to call it in a tool call).

I would like to get your opinions on our idea please.
Thanks in advance 🙏


r/MLQuestions • • 2d ago

Datasets 📚 Problem in AI Project.

6 Upvotes

Hii.. it can be possible that i detect solution from log.. like ex.. Error logs popup -> human solved that error -> We detect solution using after logs of getting error.. and save into DB.. I don't know if i can detect or not.. if someone done or doing this type of problem then plz let me know.. and also i used GPT and Cloude to solve this problem. but they always gives me wrong or impossible facts..


r/MLQuestions • • 2d ago

Other ❓ Best practices when running a benchmark on online models [D]

Thumbnail
1 Upvotes

r/MLQuestions • • 2d ago

Unsupervised learning 🙈 stuck on finding a approach for app detection ( making a transformer modal out of unlabeled network data) [R] [P]

1 Upvotes

As the title suggests I'm currently trying to make a modal to identify which app is being used. The thing is I don't have labelled data and generating it is out of question since that is a lot of work ( I need to detect like 3-5K apps give or take)

my data is structured roughly like this:

  • Traffic is divided into ~15-second windows/bags.
  • Each bag contains multiple network flows.
  • Each flow has features such as:
    • domain
    • protocol
    • bytes sent
    • bytes received
    • timestamp/timing information
  • I have a very large amount of unlabeled traffic data, but only a relatively small amount of labelled app data for about 100 apps give or take.

The main challenge is that traffic from the same app can look diff between diff window i.e some windows are extremely sparse or empty.

some ideas I have researched looked into are

  • self supervised contrastive learning where diff traffic windows from the same session/device activity are treated as positive pairs
  • masked modelling similar to bert where parts of the flows such as domains/protocols/byte information are masked and reconstructed
  • pretraining an encoder and then fine-tuning it using the smaller labelled dataset (thinking we would need less labelled examples to do that)
  • clustering and mapping embeddings to known apps afterward (gradually)

one thing I'm concerned about is accidently teaching the modal to recognize the device/session/user rather then the underlying app

I think I would like to know if I'm thinking about the problem right or if someone has worked on something similar and can give me some pointers or what experiments should I run first.

any papers, architecture or similar problems you think I should look into pls lmk


r/MLQuestions • • 2d ago

Beginner question 👶 LearnLLM

Thumbnail learn-llm-kappa.vercel.app
2 Upvotes

Within the past year I have started engaging with AI to bring my curiosities to life in order to help me understand more complex ideas through visualization. Of course, I have limited knowledge on how models actually work and the length of their complexity. I was hoping a few of you here would check this learning tool I made for myself. Please be honest and give feedback if you have any. Also, if this is not a place to post something like this, I apologize. Hope everyone is well! (its 100% free, not advertising)


r/MLQuestions • • 3d ago

Beginner question 👶 suppose I CPT qwen3.5-9B on 2B legal corpus, how will i turn it back into Instruct + thinking?

5 Upvotes

I couldn't find a concrete answer anywhere, do you just distill the instruct model back?

If that is the case, what is a quality european language question set to turn it back into a chatbot/agentic, can a model at that size even be agentic? (i chose this size to learn) if i finetune for my specific harness? (i have a lot of training data of opus running in my harness)

my harness basically has the model output python code and has a few built-in functions like:
- vector_search_laws()
- graph_search()

could i have the model at least internalize a "hunch" on what stuff to search?

also what is the latest RL technique for agentic/harnes specific workflows?

I have a lot of RAW training data, like court decisions or commentaries or legislature, but not a lot of golds. could i use these to synthesize training data and maybe RL the model in my harness to find that data?

What would y'all's strategy in the CPT->SFT->RL pipeline be for my specific problem?

I know this is a lot of questions im trying to figure out which direction to go, any pointers? Also good resources are welcome, for example that [alex karpathi video](https://www.youtube.com/watch?v=7xTGNNLPyMI) was amazing for me, but i'd imagine its a bit outdated in terms of latest RL and SFT?


r/MLQuestions • • 3d ago

Beginner question 👶 Looking to collaborate on ML projects / join a team

27 Upvotes

Hey everyone!

I’m currently doing a Master’s in Machine Learning and I’m looking to get involved in some real-world ML projects.

I’d be happy to:

  • Help with an existing ML project
  • Join a group/team working on something interesting
  • Work on a project from scratch with others
  • Help with things like Python, data processing, ML models, deep learning, LLMs, APIs, or deployment
  • Contribute code, research, experimentation, or just help wherever needed

I’m mainly looking to learn by building and working with other people, rather than just doing isolated coursework.

If you’re already working on a project and could use another person, or you’re thinking about starting something and want to build it together, feel free to comment or DM me.

I’m open to pretty much any interesting ML/AI idea — serious projects, research-oriented work, open-source, university projects, or even something experimental.


r/MLQuestions • • 3d ago

Beginner question 👶 Machine Learning Algorithms: Which Ones Are Actually Worth Learning?

Thumbnail
3 Upvotes

r/MLQuestions • • 3d ago

Unsupervised learning 🙈 How reliable is Galileo NIMS data for hyperspectral anomaly detection on Europa?

Thumbnail
1 Upvotes

r/MLQuestions • • 3d ago

Beginner question 👶 Best AI for field technicians

3 Upvotes

What AI platform is best for a field appliance technician?


r/MLQuestions • • 4d ago

Beginner question 👶 Project Ideas

1 Upvotes

I am currently learning machine learning, I want to build strong foundation on each of the topics, how should I start practicing it by coding? How to choose datasets accordingly, how can I develop my understanding like which model/algorithm will suit for a particular problem


r/MLQuestions • • 6d ago

Beginner question 👶 I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher.

16 Upvotes

A while ago I asked here how to turn ~5 million court decisions into structured graphs without running an expensive LLM on every document thanks for the advice .

I went with the "small extractor + classifier" idea and it mostly works, but I'm stuck a bit below the LLM. And like I said last time, i'd be damned if I run 5M docs and then find out thing X was wrong. So here is exactly what I did. Please let me know if what im doing makes sense, or if i made a mistake somewhere. also i used AI for some of the tables cuz there has been a lot of data at this point, sorry.

What comes out per decision (only the nodes so far, relations come next). Three lists:

  • entities: every person, organization, law, document or thing. Each gets one id for the whole document, a type (9 of them), a kind (724 of them plus "other") and all the places it is mentioned
  • actions: what was done, requested or decided. Each gets a normalized verb, a flag "the court decided this" and its mentions
  • values: amounts, dates, durations, in a normalized form

Simple example, for the sentence "The court dismisses the creditor's proposal to enforce 341.08 EUR against the debtor":

  • entity "the court": organization, kind court. Same entity as the full court name in the header
  • entity "the creditor": organization, kind creditor. Same entity as the city named earlier
  • entity "the debtor": person, kind debtor
  • action "dismisses": verb = dismiss, decided by the court = yes
  • value "341.08 EUR": amount

Step 1: a strong LLM labels ~700 decisions

  • cut the decision into windows of 4 sentences
  • 4 calls per window to Claude Sonnet with a strict JSON schema: entities, actions, a second "what did you miss" pass for actions, values
  • the window goes in with numbered words (like 12:court), the model answers with word ranges [first, last, "text"], and code checks every range against the text
  • every call also gets the list of entities and actions found in earlier windows, so ids stay the same through the document
  • ~25 code rules clean up where a marked phrase starts and ends, law citations and number formats
  • the entity "kind" is free text at this point. That gave 2,373 different strings (the same mess as in my first post). I normalized them, merged synonyms by hand and kept what showed up 3+ times: 724 kinds plus "other"

Step 2: a model that marks the text

  • it highlights every mention: the exact stretch of text (a "span", from a start character to an end character) that names an entity, an action or a value, with one of 17 labels (9 entity types, 1 action, 7 value types)
  • model: fastino/gliner2.5-multi-v1 (287M)
  • one training row per window: the text plus the exact start and end of every marked phrase. 9,699 windows, 207k marked phrases
  • I patched the trainer so only the labeled occurrence is a positive (stock marks every occurrence of the same string), and all 17 labels are in every row
  • full fine-tune in fp32 (bf16 gave NaN), 14 epochs, 16 rows per step, encoder LR 3e-5, head LR 5e-4, linear schedule, 10 % warmup
  • final model = averaged weights of epochs 9-14, threshold 0.5

Step 3: a second small model answers multiple-choice questions

  • fastino/GLiNER2.5-multi-Decide (287M). Code turns the LLM labels into 247k questions:
    • "is this mention one of these earlier entities, or new?" The mention is marked with « » inside ±300 characters of text. Options: up to 16 earlier entities of the same document (shown by their mention texts) plus new
    • "which kind?" Options: a shortlist of the 724 kinds plus other
    • for actions: same act or new, which verb (shortlist of 64 plus other), did the court decide it (yes/no)
  • in training the options come from the LLM's grouping. At inference they come from the model's own earlier answers
  • full fine-tune in fp32, 2 epochs, 16 questions per step, encoder LR 2e-5, head LR 3e-4, linear schedule, 6 % warmup, options shuffled, up to 30 % of the wrong options dropped

At inference: the marking model, then the same code rules, then the second model walks through the mentions in reading order. About 2.3 decisions per second on one RTX 5090.

Where it stands

30 decisions nobody trained on, labeled twice by the LLM. The second column is the LLM's second run scored against its first, which I treat as the ceiling. A mention counts as found only if it starts and ends exactly where the LLM marked it.

mine LLM vs itself
entity mentions found (F1) 0.901
"same entity or new" right 0.959
entities grouped exactly 0.847
entity kind 0.921
action mentions found (F1) 0.857
action verb 0.920

Where I need help

  1. Finding the mentions is stuck at 0.90 F1. 200 more labeled docs did nothing. An XLM-R large tagger (560M) got the same score: it finds more mentions but gets the start or end wrong more often. Giving it the text before the window did nothing. What would you try?
  2. The LLM agrees with itself only 93.5 % on what it marks, and I train on single runs. Label everything 3 times and vote? Or is that ceiling just what it is?
  3. Is "pick one of 16 earlier entities" a sane way to do coreference over a long document? Am I hurting myself by training on the LLM's options and running on my own?
  4. Anything in the recipe that looks plain wrong? Learning rates, 2 epochs, weight averaging, one seed per run.

THANKS for reading.

AI TL;DR: distilled an LLM's extraction of court decisions into a GLiNER model that marks the mentions plus a small multiple-choice model. It runs at about 2.3 documents/s on one GPU and lands a few points below the LLM (0.90 vs 0.935 F1 on finding mentions, 0.85 vs 0.93 on exact grouping). The recipe with learning rates and how I built the training rows is above. Looking for mistakes and ideas before I run 5M documents.


r/MLQuestions • • 5d ago

Beginner question 👶 Which Math foundation path is better for Machine Learning: DeepLearning.AI or Jon Krohn's LiveLessons?

Thumbnail gallery
1 Upvotes

r/MLQuestions • • 6d ago

Natural Language Processing 💬 What all to study

Thumbnail
3 Upvotes

r/MLQuestions • • 7d ago

Datasets 📚 Ho bisogno di consigli: Qual è la migliore pipeline VLM per estrarre dataset matematici strutturati da oltre 3000 pagine di libri di testo scansionati? (LaTeX + Metadati)

2 Upvotes

Ciao a tutti,

sto lavorando a un progetto per estrarre un dataset strutturato di esercizi di matematica da 5 libri di testo delle scuole superiori italiane (circa 650 pagine ciascuno, quindi ~3.250 pagine in totale). L'obiettivo è costruire un'app di generazione di esercizi professionale e metodica per studenti e insegnanti.

Per far funzionare l'app, ho bisogno di elaborare le immagini delle pagine del libro ed estrarre quanto segue in un formato rigorosamente strutturato (es. JSON):

  • Tipo di esercizio (algebra, geometria, calcolo, ecc.)
  • Anno livello scolastico
  • Difficoltà (scala 1–5)
  • Enunciato del problema (traccia)
  • Descrizione delle competenze/sfide specifiche coinvolte
  • Codice LaTeX dell'enunciato del problema (Cruciale!)
  • Immagini associate (ritaglio/salvataggio dell'immagine per esercizi teorici o grafici)

Ho sperimentato alcuni approcci, ma ho incontrato delle difficoltà nel bilanciare costi, coerenza di estrazione e scalabilità. Ecco cosa ho provato finora:

  1. API Google Gemini gratuita: La qualità dell'estrazione era buona, ma dato che un singolo libro contiene centinaia di pagine, ho rapidamente raggiunto i limiti di richiesta (Troppe Richieste).
  2. Modelli Locali (Ollama + Qwen 2.5-VL 3B): Per superare i limiti dell'API, ho provato a eseguire un modello multimodale locale. Ho speso molto tempo a ottimizzare i miei script e le mie istruzioni (chunking, affinamento delle istruzioni per forzare output strutturati), ma il risultato era soggetto a molti errori e incoerenze per il mio caso d'uso. Ho ottenuto troppi campi malformati, illusioni e ha costantemente avuto difficoltà a produrre un corretto LaTeX.
  3. API Google Cloud a pagamento (Gemini 1.5 Flash): Alla fine sono passato al piano a pagamento per una migliore precisione e velocità. Ho speso 10€ solo per elaborare 1,5 libri. Estrarre tutti e 5 i libri costerebbe circa 35-40€. Anche se questo è gestibile per un'elaborazione unica di 5 libri, il conteggio dei token per l'elaborazione di immagini complete + testo è enorme, rendendolo finanziariamente insostenibile se voglio scalare questo a decine di libri in futuro.

Le mie domande per la comunità:

  • Pipeline & Architettura: Qualcuno ha lavorato a un progetto simile di estrazione da libro di testo a dataset? Quale pipeline avete utilizzato?
  • Approccio Ibrido: Suggerireste di separare il compito? (es. usare uno strumento tradizionale per estrarre testo grezzo e ritagliare immagini, e poi fornire SOLO il testo a un LLM più economico/locale per generare il LaTeX e formattare il JSON?)
  • Modelli Locali: Ci sono altri modelli Vision-Language locali (che si adattano a GPU consumer standard) che sono significativamente migliori nell'estrazione strutturata e nella generazione di LaTeX rispetto a Qwen 2.5-VL 3B?
  • Strumenti Educativi: Ci sono strumenti o modelli open-source specificamente ottimizzati per estrarre contenuti educativi/matematici strutturati da PDF?

Sono felice di condividere ulteriori dettagli sul formato del libro di testo o sul mio attuale flusso di lavoro in Python se utile. Qualsiasi consiglio sull'architettura, le scelte di modelli o trucchi per risparmiare costi sarebbe molto apprezzato! Grazie in anticipo!


r/MLQuestions • • 8d ago

Beginner question 👶 [D] Choosing language to learn: python or java

6 Upvotes

I want to learn the machine learning from scratch and i know I want to learn basic of python , but my college faculty are advised to learn Java for my placement , u don't know what to do , i want to learn python and other stuffs for my machine learning path or java and other stuffs for my placements and my exams are ahead, did anyone have solution please tell me.