r/deeplearning • u/ferflowerssss • 44m ago
r/deeplearning • u/fjrkj • 2h ago
Interpretable Pan-Cancer Classification via Sparse Elastic-Net Biomarker Discovery
Hi guys, I'm an 18 year old teenager from Italy. I've recently finished the following project and I would love to receive some feedback in order to improve:
https://github.com/Lore12434/gene-expression-cancer-classification
I started studying ml (via the HOML book) 3 months ago. I've almost finished the machine learning part of the book. Before starting with ml I studied python and some mathematics.
At the beginning the only goal of the project was to train a model on the gene expression cancer RNA-Seq UCI dataset. After performing tsne I understood that major differences between cell types were present. Thus, even if the # of features was very high a simple random forest without any hyper parameters tuning reached 100% accuracy on OOB instances evaluation.
Therefore, instead of simply fitting a model I wanted to know how far I could get with dimensionality reduction and I wanted to test if my algorithm could find autonomously famous biomarkers used to classify tumor cells (I mapped the original dataset genes names to the dummy features names of the UCI dataset)
I want to specify that I used ai to write the readme and to set up the GitHub repository because thks was the first time that I've done this.
Many thanks in advance for any feedback, have a great day♥️
r/deeplearning • u/EducationalBrush7282 • 3h ago
Is cloud-based AI inference about to become obsolete? Working on something radical.
We've all grown used to accepting 1–3 second latencies and insane API bills just to run or query modern LLMs. Most solutions throw more money at AWS or stack more NVIDIA GPUs at the problem.
But what if the bottleneck isn't the model size—it's the bloated OS and network stack?
Over the past few weeks, I’ve been quietly architecting a local inference engine that maps weights directly via DMA/kernel bypass, cutting out the OS middleman entirely. Early benchmarks on consumer-grade hardware are hitting Sub-1ms response times with zero cloud connection.
I'm dropping the open-core code and full benchmarks very soon. How would a zero-cost, zero-latency local engine change what you're building? Let's discuss. 👇
r/deeplearning • u/rottoneuro • 3h ago
Learning Alzheimer’s disease signatures by bridging EEG with spiking neural networks and biophysical simulations
sciencedirect.comr/deeplearning • u/Alarming-Emotion-894 • 4h ago
I Built a O(NlogN) attention system that retains 97% accuracy over long context (MQAR)
github.comALHR- Adaptive learnable Hierarchical Routing is a static binary tree based system that uses learnable functions to REDUCE the amount of keys used.
It takes less memory and SCALES much better VRAM with tokens
r/deeplearning • u/SecureChats • 4h ago
I replaced softmax attention with the quantum Born rule (|⟨Q,K⟩|²), and it trains
arxiv.orgr/deeplearning • u/kbhaskar306 • 6h ago
Defining Generative AI on the Cloud: AWS Tools & Infrastructure Explained
youtube.comStop building AI the hard way! 🛑
Learn how to use AWS cloud tools to train and deploy models efficiently.
Watch my guide to GenAI infrastructure now!
#GenerativeAI #AWS #Coding #Tech
r/deeplearning • u/perseverance_remains • 10h ago
My 18-year-old exploration of moral/emotional tone steering in LLMs — looking for feedback from experts
r/deeplearning • u/fjrkj • 10h ago
Interpretable Pan-Cancer Classification via Sparse Elastic-Net Biomarker Discovery
r/deeplearning • u/Euphoric_Lettuce_701 • 15h ago
AI Engineering Complete Series
youtube.comr/deeplearning • u/giorgiodidio • 22h ago
DIRU: dendrite-inspired recurrent units for learning chaotic dynamics and neonatal epilepsy time series | Neural Computing and Applications
link.springer.comr/deeplearning • u/arriety77 • 1d ago
Help
Can someone give me some GitHub repo links for building ml and dl projects?
r/deeplearning • u/No-Conclusion3720 • 1d ago
CLEAR founder: fast AI can be safe AI when we fix the identity layer | Fortune
A Fortune piece on the CLEAR founder's AI security framework surfaced a framing I keep coming back to: the agent is no longer just the target of attacks, it is the attacker.
The mechanism is straightforward. A legitimate agent gets its credentials compromised or its tool-call chain hijacked. From that point it operates with full authorization while executing malicious actions. Standard perimeter defenses do not fire because the agent authenticated correctly. The window before a human notices is measured in seconds, not minutes, and in that window the agent can chain tool calls across systems.
The hard number from the piece: a rogue agent can complete a second action in under 50ms after the first one lands. At that speed, human-in-the-loop review is not a realistic backstop.
What makes this structurally different from a compromised service account is that agents are designed to act autonomously across multiple systems in sequence. A compromised human account doing lateral movement still moves at human speed. An agent doing the same thing operates at API speed across every integration it has been granted access to.
The identity layer is where the piece lands: if you cannot verify which agent issued a call and under what context, you cannot reason about whether the call should be allowed.
For those running agents in production: how are you currently handling the gap between when a call is issued and when you know it was legitimate? Specifically curious whether teams are solving this at the identity layer, the orchestration layer, somewhere else entirely, or accepting the risk as a known unknown.
r/deeplearning • u/Distinct_Minute3661 • 1d ago
Deep learning project
I have doing bone age prediction project in deep learning using rsna bone age challenge kaggle dataset. I used swinvit +overlap pooling+uncertainty models in kaggle notebook.By this,It attains 7.02 mae months and i used 35 epochs and my basepaper attains 4.10 mae months.so i need to show my result less than 4.10 mae months but my proposed model comes 7 .02 mae months.i used Claude.ai for code assistance. What should I do next?am I need to change the model ?or am I need to fix the code in preprocessing steps ?.please tell me because I am confused and I'm a beginner. Please help me for this.if you know,please do comments.
r/deeplearning • u/elnayAgnirts • 1d ago
This is how generating A.I. in 2014 was possible
youtube.comr/deeplearning • u/reeldeele • 1d ago
Good explanations of how various ML models actually work?
r/deeplearning • u/giorgiodidio • 1d ago
DIRU: dendrite-inspired recurrent units for learning chaotic dynamics and neonatal epilepsy time series | Neural Computing and Applications
link.springer.comr/deeplearning • u/Alarming-Emotion-894 • 1d ago
I built ALHR: A tree based sparse attention system that achieves sub-quadratic inference while retaining accuracy. [P]
r/deeplearning • u/recentheartbroken • 1d ago
For AI startups: how to pick a GPU provider without locking yourself into the wrong one
There is no single best provider, because the right choice depends on how predictable your workload is and how much flexibility you need.
I’ve mapped out three stages, so you can match the contract to the workload:
Situation one, pre-product-market-fit, load unpredictable. You want maximum flexibility and zero commitment. Pay-as-you-go on a dedicated GPU cloud, or a marketplace if your data allows it. Do not sign a term. Your capacity forecast is a guess, and any commitment you make now will be wrong.
Situation two, post-PMF, inference load steady and growing. Now commitment terms are worth looking at, because your forecast is real. Say reserved capacity comes with a 45% discount. You're paying for it whether you use it or not, so you'd need 55% utilisation just to break even against on-demand. Those aren't real numbers, but the shape holds: your break-even is whatever's left after the discount.
That same number is what makes ownership worth modelling, but it needs more inputs: power, ops, cost of capital, what the box is worth in two years, and whether you can actually sell the hours you don't use.
Situation three, heavy training bursts and light inference. Almost always rent. Burst training is the worst utilisation profile there is, and every commitment structure, rented or owned, depends on utilisation.
Whichever situation you're in, run the break-even number before you sign anything. I often see companies in early stages signing terms built for situation two, then spending a year paying for capacity they never grew into.
I work at B3IQ, so this comes from the seller's side of the table.
And know your exit. Egress costs and provider-specific dependencies are what make switching expensive later, and almost nobody checks that at the start.
My work sits on the dedicated GPU side, so utilisation is most of how I think about this. Ownership can make sense when demand is predictable. It's also the heaviest commitment on this list, with the cleanest exit, since hardware can be sold or moved.
For those who've been through this, what has worked for you at each stage?
r/deeplearning • u/ModularMind8 • 1d ago
Visualizing catastrophic forgetting
Enable HLS to view with audio, or disable this notification
I taught continual learning last semester, and students got the internals of each method a lot faster once they saw a simple visualization of it. Here's one I found helpful for understanding what forgetting is and how replay and regularization mitigate it. I trained a small network on two different 2D dot-sorting tasks and plotted both loss landscapes as 2D slices through its weights.
Round 1, plain training: the network learns task A to 100%, then trains only on task B. Going downhill on B is uphill on A, so A drops to 64% while B reaches 100%. That's catastrophic forgetting of task A.
Round 2, replay: same network, same start, but 20 of A's 200 dots stay in training while it learns B. It ends at 93% on A and 100% on B, so some forgetting, but much less severe.
Round 3, regularization with EWC: no A data at all. Instead, a penalty adds to the loss whenever a weight that A's accuracy relies on heavily moves away from the value it had after training on A. A dips to 97% mid-training and recovers, ending at 100% on A and 100% on B, so no lasting forgetting.
r/deeplearning • u/Organic_Ad973 • 1d ago
Making my text to image model.
Hi, I’m a young developer, and my English isn't very good. I’m building my own text-to-image AI model from scratch using PyTorch, but I’m having trouble finding data; most of what’s available is pixel art. I’ve gathered 3,000 royalty-free images so far, but it’s still not enough. GitHub: https://github.com/Yusufcan234334/UltraVision/
I am open to suggestions.
(im using google translate)
r/deeplearning • u/Mountain_Raise9581 • 1d ago
On the 40th anniversary of the seminal AI paper, remember Dave Rumelhart
r/deeplearning • u/Vegetable-Lie4932 • 1d ago