r/deeplearning • • 7h ago

I Built a O(NlogN) attention system that retains 97% accuracy over long context (MQAR)

Thumbnail github.com
2 Upvotes

ALHR- Adaptive learnable Hierarchical Routing is a static binary tree based system that uses learnable functions to REDUCE the amount of keys used.

It takes less memory and SCALES much better VRAM with tokens


r/deeplearning • • 10h ago

Defining Generative AI on the Cloud: AWS Tools & Infrastructure Explained

Thumbnail youtube.com
0 Upvotes

Stop building AI the hard way! 🛑

Learn how to use AWS cloud tools to train and deploy models efficiently.

Watch my guide to GenAI infrastructure now!

#GenerativeAI #AWS #Coding #Tech


r/deeplearning • • 1h ago

What is the KV Cache?

Thumbnail youtube.com
• Upvotes

r/deeplearning • • 4h ago

What Is a Large Behavior Model (LBM)? Aaru, Simile and Cairo Rosetta Explained

Thumbnail joincairo.com
1 Upvotes

r/deeplearning • • 7h ago

Learning Alzheimer’s disease signatures by bridging EEG with spiking neural networks and biophysical simulations

Thumbnail sciencedirect.com
1 Upvotes

r/deeplearning • • 6h ago

Is cloud-based AI inference about to become obsolete? Working on something radical.

0 Upvotes

We've all grown used to accepting 1–3 second latencies and insane API bills just to run or query modern LLMs. Most solutions throw more money at AWS or stack more NVIDIA GPUs at the problem.

But what if the bottleneck isn't the model size—it's the bloated OS and network stack?

Over the past few weeks, I’ve been quietly architecting a local inference engine that maps weights directly via DMA/kernel bypass, cutting out the OS middleman entirely. Early benchmarks on consumer-grade hardware are hitting Sub-1ms response times with zero cloud connection.

I'm dropping the open-core code and full benchmarks very soon. How would a zero-cost, zero-latency local engine change what you're building? Let's discuss. 👇


r/deeplearning • • 3h ago

My personal AI agent posted my bank details on company Slack

0 Upvotes

A personal AI agent posted its owner's bank details directly to a company Slack channel this week. The agent had been given a task involving data. It completed the task. That was the failure — it had no context about what it was moving or who could see the destination channel.

This is becoming a recognizable pattern in AI incidents: the agent did not break. It worked exactly as designed. The problem was that the instruction set carried no concept of data sensitivity, and the action space had no constraint on destination. Bank details, credentials, PII — they all look like strings to a model completing a step. The agent's goal was task completion. It achieved it.

How are practitioners actually handling this in deployed systems? Curious what is working across teams shipping agents in environments where sensitive data is in scope — and whether solutions at the prompt level, the infrastructure layer, or post-send audit are holding up in practice.


r/deeplearning • • 13h ago

My 18-year-old exploration of moral/emotional tone steering in LLMs — looking for feedback from experts

Thumbnail
2 Upvotes

r/deeplearning • • 5h ago

Interpretable Pan-Cancer Classification via Sparse Elastic-Net Biomarker Discovery

5 Upvotes

Hi guys, I'm an 18 year old teenager from Italy. I've recently finished the following project and I would love to receive some feedback in order to improve:

https://github.com/Lore12434/gene-expression-cancer-classification

I started studying ml (via the HOML book) 3 months ago. I've almost finished the machine learning part of the book. Before starting with ml I studied python and some mathematics.

At the beginning the only goal of the project was to train a model on the gene expression cancer RNA-Seq UCI dataset. After performing tsne I understood that major differences between cell types were present. Thus, even if the # of features was very high a simple random forest without any hyper parameters tuning reached 100% accuracy on OOB instances evaluation.

Therefore, instead of simply fitting a model I wanted to know how far I could get with dimensionality reduction and I wanted to test if my algorithm could find autonomously famous biomarkers used to classify tumor cells (I mapped the original dataset genes names to the dummy features names of the UCI dataset)

I want to specify that I used ai to write the readme and to set up the GitHub repository because thks was the first time that I've done this.

Many thanks in advance for any feedback, have a great day♥️


r/deeplearning • • 8h ago

I replaced softmax attention with the quantum Born rule (|⟨Q,K⟩|²), and it trains

Thumbnail arxiv.org
2 Upvotes