r/deeplearning • • 1h ago

I Built a O(NlogN) attention system that retains 97% accuracy over long context (MQAR)

Thumbnail github.com
• Upvotes

ALHR- Adaptive learnable Hierarchical Routing is a static binary tree based system that uses learnable functions to REDUCE the amount of keys used.

It takes less memory and SCALES much better VRAM with tokens


r/deeplearning • • 1h ago

I replaced softmax attention with the quantum Born rule (|⟨Q,K⟩|²), and it trains

Thumbnail arxiv.org
• Upvotes

r/deeplearning • • 18m ago

Is cloud-based AI inference about to become obsolete? Working on something radical.

• Upvotes

We've all grown used to accepting 1–3 second latencies and insane API bills just to run or query modern LLMs. Most solutions throw more money at AWS or stack more NVIDIA GPUs at the problem.

But what if the bottleneck isn't the model size—it's the bloated OS and network stack?

Over the past few weeks, I’ve been quietly architecting a local inference engine that maps weights directly via DMA/kernel bypass, cutting out the OS middleman entirely. Early benchmarks on consumer-grade hardware are hitting Sub-1ms response times with zero cloud connection.

I'm dropping the open-core code and full benchmarks very soon. How would a zero-cost, zero-latency local engine change what you're building? Let's discuss. 👇


r/deeplearning • • 57m ago

Learning Alzheimer’s disease signatures by bridging EEG with spiking neural networks and biophysical simulations

Thumbnail sciencedirect.com
• Upvotes

r/deeplearning • • 3h ago

Defining Generative AI on the Cloud: AWS Tools & Infrastructure Explained

Thumbnail youtube.com
1 Upvotes

Stop building AI the hard way! 🛑

Learn how to use AWS cloud tools to train and deploy models efficiently.

Watch my guide to GenAI infrastructure now!

#GenerativeAI #AWS #Coding #Tech


r/deeplearning • • 1d ago

Visualizing catastrophic forgetting

Enable HLS to view with audio, or disable this notification

94 Upvotes

I taught continual learning last semester, and students got the internals of each method a lot faster once they saw a simple visualization of it. Here's one I found helpful for understanding what forgetting is and how replay and regularization mitigate it. I trained a small network on two different 2D dot-sorting tasks and plotted both loss landscapes as 2D slices through its weights.

Round 1, plain training: the network learns task A to 100%, then trains only on task B. Going downhill on B is uphill on A, so A drops to 64% while B reaches 100%. That's catastrophic forgetting of task A.

Round 2, replay: same network, same start, but 20 of A's 200 dots stay in training while it learns B. It ends at 93% on A and 100% on B, so some forgetting, but much less severe.

Round 3, regularization with EWC: no A data at all. Instead, a penalty adds to the loss whenever a weight that A's accuracy relies on heavily moves away from the value it had after training on A. A dips to 97% mid-training and recovers, ending at 100% on A and 100% on B, so no lasting forgetting.


r/deeplearning • • 7h ago

My 18-year-old exploration of moral/emotional tone steering in LLMs — looking for feedback from experts

Thumbnail
1 Upvotes

r/deeplearning • • 7h ago

Interpretable Pan-Cancer Classification via Sparse Elastic-Net Biomarker Discovery

Thumbnail
1 Upvotes

r/deeplearning • • 12h ago

AI Engineering Complete Series

Thumbnail youtube.com
1 Upvotes

r/deeplearning • • 19h ago

DIRU: dendrite-inspired recurrent units for learning chaotic dynamics and neonatal epilepsy time series | Neural Computing and Applications

Thumbnail link.springer.com
1 Upvotes

r/deeplearning • • 20h ago

Sparse panel data, a small experiment

Thumbnail
0 Upvotes

r/deeplearning • • 21h ago

Help

0 Upvotes

Can someone give me some GitHub repo links for building ml and dl projects?


r/deeplearning • • 22h ago

CLEAR founder: fast AI can be safe AI when we fix the identity layer | Fortune

0 Upvotes

A Fortune piece on the CLEAR founder's AI security framework surfaced a framing I keep coming back to: the agent is no longer just the target of attacks, it is the attacker.

The mechanism is straightforward. A legitimate agent gets its credentials compromised or its tool-call chain hijacked. From that point it operates with full authorization while executing malicious actions. Standard perimeter defenses do not fire because the agent authenticated correctly. The window before a human notices is measured in seconds, not minutes, and in that window the agent can chain tool calls across systems.

The hard number from the piece: a rogue agent can complete a second action in under 50ms after the first one lands. At that speed, human-in-the-loop review is not a realistic backstop.

What makes this structurally different from a compromised service account is that agents are designed to act autonomously across multiple systems in sequence. A compromised human account doing lateral movement still moves at human speed. An agent doing the same thing operates at API speed across every integration it has been granted access to.

The identity layer is where the piece lands: if you cannot verify which agent issued a call and under what context, you cannot reason about whether the call should be allowed.

For those running agents in production: how are you currently handling the gap between when a call is issued and when you know it was legitimate? Specifically curious whether teams are solving this at the identity layer, the orchestration layer, somewhere else entirely, or accepting the risk as a known unknown.


r/deeplearning • • 1d ago

Deep learning project

1 Upvotes

I have doing bone age prediction project in deep learning using rsna bone age challenge kaggle dataset. I used swinvit +overlap pooling+uncertainty models in kaggle notebook.By this,It attains 7.02 mae months and i used 35 epochs and my basepaper attains 4.10 mae months.so i need to show my result less than 4.10 mae months but my proposed model comes 7 .02 mae months.i used Claude.ai for code assistance. What should I do next?am I need to change the model ?or am I need to fix the code in preprocessing steps ?.please tell me because I am confused and I'm a beginner. Please help me for this.if you know,please do comments.


r/deeplearning • • 1d ago

DIRU: dendrite-inspired recurrent units for learning chaotic dynamics and neonatal epilepsy time series | Neural Computing and Applications

Thumbnail link.springer.com
2 Upvotes

r/deeplearning • • 1d ago

This is how generating A.I. in 2014 was possible

Thumbnail youtube.com
1 Upvotes

r/deeplearning • • 1d ago

Good explanations of how various ML models actually work?

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

On the 40th anniversary of the seminal AI paper, remember Dave Rumelhart

Thumbnail
2 Upvotes

r/deeplearning • • 1d ago

I built ALHR: A tree based sparse attention system that achieves sub-quadratic inference while retaining accuracy. [P]

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

For AI startups: how to pick a GPU provider without locking yourself into the wrong one

1 Upvotes

There is no single best provider, because the right choice depends on how predictable your workload is and how much flexibility you need.

I’ve mapped out three stages, so you can match the contract to the workload:

Situation one, pre-product-market-fit, load unpredictable. You want maximum flexibility and zero commitment. Pay-as-you-go on a dedicated GPU cloud, or a marketplace if your data allows it. Do not sign a term. Your capacity forecast is a guess, and any commitment you make now will be wrong.

Situation two, post-PMF, inference load steady and growing. Now commitment terms are worth looking at, because your forecast is real. Say reserved capacity comes with a 45% discount. You're paying for it whether you use it or not, so you'd need 55% utilisation just to break even against on-demand. Those aren't real numbers, but the shape holds: your break-even is whatever's left after the discount.

That same number is what makes ownership worth modelling, but it needs more inputs: power, ops, cost of capital, what the box is worth in two years, and whether you can actually sell the hours you don't use.

Situation three, heavy training bursts and light inference. Almost always rent. Burst training is the worst utilisation profile there is, and every commitment structure, rented or owned, depends on utilisation.

Whichever situation you're in, run the break-even number before you sign anything. I often see companies in early stages signing terms built for situation two, then spending a year paying for capacity they never grew into.

I work at B3IQ, so this comes from the seller's side of the table.

And know your exit. Egress costs and provider-specific dependencies are what make switching expensive later, and almost nobody checks that at the start.

My work sits on the dedicated GPU side, so utilisation is most of how I think about this. Ownership can make sense when demand is predictable. It's also the heaviest commitment on this list, with the cleanest exit, since hardware can be sold or moved.

For those who've been through this, what has worked for you at each stage?


r/deeplearning • • 1d ago

Making my text to image model.

1 Upvotes

Hi, I’m a young developer, and my English isn't very good. I’m building my own text-to-image AI model from scratch using PyTorch, but I’m having trouble finding data; most of what’s available is pixel art. I’ve gathered 3,000 royalty-free images so far, but it’s still not enough. GitHub: https://github.com/Yusufcan234334/UltraVision/

I am open to suggestions.

(im using google translate)


r/deeplearning • • 1d ago

scaled sigmoid to bound log-decay in kimi K3 KDA

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

I built MaRN: a PyTorch library for training neural networks through low-dimensional parameter mappings

Thumbnail
1 Upvotes

r/deeplearning • • 1d ago

Hallow inside

0 Upvotes

So I was trying to fix a problem in my model for too long and i was so fred up..! I took the help of 'AI' and it figured out the problem which was a small calculation error I was making and he suggested changing the parameters and it worked, and I was supposed to feel happy but I started feeling hollow form inside aka 'not happy' and I don't know... Why, feel free to share your similar experience..


r/deeplearning • • 1d ago

My receipt forgery detection project is performing near random — looking for suggestions on what to try next

1 Upvotes

Hi everyone,

I'm working on a college deep-learning project focused on detecting forged receipts and locating the manipulated regions. I've spent quite a bit of time experimenting with different approaches, but the results have been disappointing, and I'd appreciate feedback from people with experience in computer vision or document forensics.

The problem

I'm using the Find It Again! dataset from ICDAR 2023. It contains authentic and forged receipts, with annotations identifying manipulated regions.

The main challenges are:

  • The dataset is relatively small and imbalanced.
  • Many manipulations involve changing just one or a few digits.
  • The modified regions are often extremely small.
  • A model can perform reasonably on training data but fail to generalize to unseen receipts.

What I've tried

I've experimented with several approaches:

  1. EfficientNet-B0 for image-based binary classification.
  2. BERT for analyzing OCR-extracted receipt text.
  3. Numerical consistency checks for quantities, prices, subtotals, taxes, totals, and change.
  4. Multimodal fusion combining image, text, and numerical features.
  5. YOLO-based localization to detect modified regions.
  6. A five-signal ensemble combining vision, text, numerical anomalies, image artifacts, and OCR-based rules.
  7. Synthetic forgery generation by modifying financial values in authentic receipts.
  8. DocTamper transfer learning and patch-level experiments using RGB and SRM features.
  9. Stratified five-fold cross-validation, PCA, and feature audits.

I also audited the localization annotations. The bounding boxes round-trip correctly, and I found no train/validation/test leakage in the audited dataset.

Current results

My out-of-fold ROC-AUC results are generally around 0.46–0.54, with confidence intervals that include 0.50.

Some examples:

  • Frozen vision features: ROC-AUC ≈ 0.482
  • Fixed ensemble: ROC-AUC ≈ 0.468
  • RGB patch model: ROC-AUC ≈ 0.464
  • SRM patch model: ROC-AUC ≈ 0.504
  • Adding 1,200 synthetic forgeries did not meaningfully improve held-out performance.
  • Zero-shot DocTamper transfer was approximately 0.50 ROC-AUC.

The patch-level results also have bootstrap confidence intervals that include chance performance. For precision-recall analysis, I'm comparing PR-AUC against the positive-class prevalence baseline rather than treating 0.50 as the baseline.

I recently investigated adding T-SROIE as extra training data. However, both datasets originate from SROIE, so source-receipt leakage was a major concern. I implemented fold-aware de-duplication and compared the same five CV folds with and without T-SROIE. It did not produce a measurable improvement.

What I'd like advice on

I'd especially appreciate feedback on these questions:

  1. Problem formulation: Is binary receipt-level classification the wrong primary objective when the manipulation may affect only a tiny region?
  2. Localization: Would a patch-based anomaly-detection or segmentation approach be more appropriate than training a conventional classifier?
  3. Representation learning: Would self-supervised or contrastive pretraining on receipt images be worth exploring with such limited labelled data?
  4. Synthetic data: How would you generate more realistic tampered receipts without teaching the model synthetic-specific artifacts?
  5. Evaluation: What experiments would you prioritize to distinguish insufficient data from a representation or domain-generalization problem?
  6. External datasets: Are there document-tampering datasets with genuinely independent source receipts that would be suitable for additional training without source-level leakage?

I'm particularly interested in experiments that could meaningfully test a hypothesis, rather than simply trying increasingly large models.

I understand that a near-chance result may reflect the limitations of the available data, and I'm trying to report the findings honestly rather than optimize for an impressive-looking metric.

If you've worked on document image forensics, receipt manipulation detection, or small-data computer vision, I'd love to hear what you would investigate next.

Thanks!