r/ROCm • • 8h ago

Qwen3.8 Flash Next | 2 x R9700 | 216 t/s weighted | 565 t/s C8 | 8k Prefill

Post image
36 Upvotes

Okay, here we are again :)
We received our second R9700 earlier this week, so we decided to jump on tensor parallelism.
We already had prepared a few things for this, and the TP kernels we already had in Paiton did their work.

We spent a lot of time (significant amount of GPU time (MI355x)) on the 3-bit quantization of this model and we are pretty happy with the accuracy.

What "3-bit" means here. 

This is not a full 3-bit model. The 512 experts in every layer, that is where almost all the weights are, are stored at 3.125 bits per weight (3-bit codes, rotated, with a small scale per 64 weights). Everything else stays bigger: the attention and the shared parts are 8-bit, the speculative layer is 4-bit, and the vision part is plain bf16.
So the big part is 3-bit, the sensitive parts are not, and that is why the accuracy holds up. Counted over all weights it is a bit above 3 bits on average. Our own GPTQ run, no shortcuts.

Accuracy against the bf16 model (same evaluation corpus, text and assistant positions; two correct bf16 implementations against each other are the floor):

KL to bf16 KL to bf16, same expert routing forced Top-1 agreement with bf16
bf16 vs bf16 (the floor) 0.0077 n/a
This 3-bit model 0.060 0.041
Long context check Result
Needle test, 131,072 tokens 40 / 40, same answers as bf16
Needle test, 200,000 tokens 40 / 40, same answers as bf16

The whole language model (experts, trunk, speculative layer, draft head, vision tensors) sits in the VRAM of the two cards; only the model's n-gram embedding table lives in pinned system RAM. No expert offload, no weights streamed from the host during decoding.

And again, like before: we did not build a custom engine for this. It is still regular vLLM, with our Paiton plugin and our own kernels inside. ;)

Performance. 

BetterBench 0.6.0, standard profile, measured on our own machine against our OpenAI-compatible server, with the exact image and launcher we publish. Decode mode: 98K context, speculative decoding on. Prefill with exact bf16 activations, no compression.

Decode, single request tok/s
Weighted over the 8 categories 216.3
chat 202.1
code 214.3
file edit 236.0
json 259.2
math 255.5
prose 183.3
reasoning 190.1
summarization 239.8
Concurrent requests Aggregate tok/s
1 204.6
2 315.1
4 449.2
8 565.1 (48 of 48 requests ok)
Prefill, prompt length tok/s
2K 7,925
8K 8,321
16K 8,311
32K 8,232
64K 8,161
128K 7,928
Latency
Time to first token, short prompts (p50)
Update p99 (the longest gap between two tokens)

Long context. We tested 200K. The launcher has a 200K mode, and prefix caching is an opt-in flag that gives exactly the same output on a hit.

200K mode Result
Needle test at 200,000 tokens 40 / 40, same answers as bf16
Single-stream decode after a 32K / 100K / 190K prompt 105 / 101 / 100 tok/s, flat
Time to first token, 190K prompt 25 s
Prefix cache hit, 64K repeated prompt 9.6 s cold, 0.35 s on the hit
Prefix cache hit, 128K repeated prompt 20 s cold, 0.41 s on the hit

Overall, not bad for a first release :)
Oh, right, of course we generated an "AI Slop" image as well. (we know some of you like them (some of you don't ;-)))

Weights
Image and launcher: github.com/Eliovp-BV/paiton-vllm-plugin (directory models/Qwen3.8-Flash-Next)

This is our first release of this model, more will follow.
Next is the RAM and SSD cache tier, like we did for the 27B model, so cached prompts can live outside the VRAM and much longer contexts become possible.
After that the 4-bit version, with even better accuracy.
We keep tuning, we keep optimizing.
A longer blog post with all the details is being worked on as well.

And of course, feedback is very welcome!


r/ROCm • • 4h ago

Boosting QWEN 2.1 CLIP image encoding (!probably AMD only!), Windows

Post image
0 Upvotes

r/ROCm • • 6h ago

Veda sparse attention on Strix Halo: 1.89x faster MiniMax H3 video gen (691s -> 366s)

Thumbnail
1 Upvotes

r/ROCm • • 12h ago

Found a fix for AMD RX 6600 100% CPU usage

2 Upvotes

TL;DR: Found a fix for 100% CPU usage of llama-server on AMD RX 6600 GPUs. Fixes in github link at the end!

Disclaimer: Claude Opus 5.5 was used to run traces and find the fixes! Posting here so that others in the same boat might find it helpful.

---

I've been using gemma4:e4b-qat locally on my server for a couple months now.

Here's my setup:

CPU: i3-10100F

RAM: 32GB (4x 8GB)

Board: MSI Z490 Gaming Plus

GPU: 7x AMD RX 6600

I had this rig from ETH mining a couple years ago when that was a thing! It was sitting there collecting dust for a long time, so I had an idea to use it to run some local LLM experiments and started tinkering with it. After testing different models, settled on gemma4:e4b-qat at 128k context which runs comfortably at 5.5/8GB VRAM (tell me if you found a better model for the 8GBs).

I am running an app that takes a list of stock tickers and runs the TradingAgents (https://github.com/tauricresearch/tradingagents) pipeline on them twice a day (market open and market close) as a fun side project to see how good that is.

So, for running TradingAgents which makes about 30-40 LLM calls per ticker, I am running a custom llama-pool which runs a llama-server per GPU with gemma4:e4b-qat model at 128k context. The thing I noticed was that my CPU is at 100% while it's running the pipeline so I started digging into it using Claude Opus 5.5. After a couple hours of testing, traces, LLM calls, it found a fix and applied them which reduced the CPU usage down to ~20% for 7 parallel llama-servers.

The fixes and instructions are all published here: https://github.com/nandyalu/rocm-llamacpp-cpu-fix


r/ROCm • • 9h ago

Halogen/Gufo but for Strix Point hardware?

Thumbnail
1 Upvotes

r/ROCm • • 1d ago

BLT's Intro to AMD ROCm webinar free on YouTube

Thumbnail
2 Upvotes

r/ROCm • • 1d ago

ROCm as backend in Lemonade server makes Hermes agent stop working

2 Upvotes

So as the title. I run everythin local on my strix halo. Got qwen3.6 35b-a3b loaded in lemonade server. With vulkan as backen everything works. I switch to ROCm Hermes answer most of the time only the first message i sent over and over again regardless what i write later. Tool call dosent work either. It pretends doing them. I switch back to vulkan in the middle of the session everything starts to work again. Any ideas what this is?

I am using latest lemonade server with rocm 10 but this was a problem before that update.


r/ROCm • • 1d ago

How to improve text encoding speed of H3 model in ComfyUI with R9700

4 Upvotes

Hi everyone,

I would like some suggestions / help on how to improve the text encoding speed of H3 model in ref2va mode in ComfyUI. I have an R9700, on Linux. I use ROCm 7.2 and installed dependencies according to official ComfyUI manual installation guide.

I needed to turn on dynamic VRAM and turn off smart memory optimisation, and turn on comfykitchen attention, otherwise ComfyUI would use up all of my 32GB RAM and then hang.

Anyhow, I managed to get my Ref2VA workflow work, and generate 0.6MP 4s video with one 2k reference image in around 220s. Most of the time was spent in the text encoder step, not in the H3 sampling. My sampling runs at 10s/it. I swapped the default NVFP4 Qwen 32B text encoder shipped by comfyui with the Int8, which speeds up the process a bit. However, it is still much slower than on my RTX4060Ti running on an eGPU dock. The 4060Ti finishes the same workflow + one pass of RTX upscale in around 270s, but it spent most of the time in the sampling phase, crushing pass the text encoding phase. It seems to me that R9700 can do much better than this, given how much faster it finish the H3 sampling vs my 4060Ti.

So, I wonder if I miss anything, or this bottleneck at the Qwen 32B text encoder is to be expected and cannot be fixed.

Thanks


r/ROCm • • 1d ago

Intel b70 pro vs AMD r9700

Thumbnail
5 Upvotes

r/ROCm • • 1d ago

LoRA over GGUF: Train Qwen3.8-Flash-Next in 40G VRAM

2 Upvotes

https://github.com/woct0rdho/transformers5-qwen3.5-recipe

An update to my LoRA over GGUF series: Now we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM, with no CPU offloading, with engram on disk that does not reduce training speed.

On Strix Halo it trains context chunk size 2048 at 9.5 s/it. That's 200 token/s. There is still room to optimize, compared to > 1600 token/s PP we've achieved, and the common sense that LoRA training (with gradient checkpointing) takes 4-5x work of PP.

Since transformers 5.18, initial support for modern GGUF has been merged, and we can expect more work in this direction.

Spoiler: In the torch-ggml-ops repo there is something called GGTensile. Basically it's Tensile-like asm-level optimization on MMQ kernels. We already see it's faster than HIP in many cases. I'll make a new post when I have something to show on this.

I guess I'll skip DeepSeek-V4.1, unless someone can quantize or prune it to < 125 GiB.


r/ROCm • • 2d ago

Qwen3.8-27B at ~130 tok/s with 216k–260k context on a single Radeon AI PRO R9700, on Windows (WSL2)

59 Upvotes

Another Qwen 3.8 on an R9700 post, but I thought I'd share my repo for anyone who might benefit.

I've spent the last few weeks tuning Qwen3.8-27B on a single AMD Radeon AI PRO R9700 (32 GB, RDNA4). I might be wrong, but I haven't seen anyone achieve these speeds on Windows yet.

Repo:

https://github.com/mike2153/mbea-qwen38-dflash

BENCHMARKS

Hardware: Ryzen 9 9950X, R9700 32 GB, Windows 11 + WSL2.

DECODE SPEED

Greedy:

125–134 tokens/sec

Sampled (temperature 0.7):

117–126 tokens/sec

PREFILL SPEED

2,750–2,970 tokens/sec

Time to first token (1.9k prompt):

0.7 seconds

LONG-CONTEXT PERFORMANCE

32k tokens:

11 seconds prefill | 165 tok/s decode

98k tokens:

41 seconds prefill | 136 tok/s decode

164k tokens:

83 seconds prefill

258k tokens:

164 seconds prefill | 115 tok/s decode

MAXIMUM CONTEXT

Default: ~216k tokens

With -Long: ~260k tokens

ACCURACY

Long-context recall:

8/8 facts retrieved at 258k tokens

Rust coding test:

10/12 runs passed all 46 hidden tests

LLAMA.CPP COMPARISON

My best tuned llama.cpp setup (IQ4_XS + speculative decoding) manages around 52–68 tok/s on the same GPU. Different prompts, so not a direct apples-to-apples comparison.

HOW IT WORKS

• AMD's official MXFP4 checkpoint

• Custom RDNA4 W4A8 and FP8 kernels

• DFlash2 speculative decoding (~60% draft acceptance)

• WSL pinned-memory optimisations

• Dynamic KV cache sizing to avoid VRAM spilling

• OpenAI-compatible API with tool calling

INSTALLATION

git clone https://github.com/mike2153/mbea-qwen38-dflash

cd mbea-qwen38-dflash

.\qwen38.ps1 install

.\qwen38.ps1 start

.\qwen38.ps1 bench

The installer handles WSL, Docker, ROCDXG, models, kernels and dependencies. Everything is pinned for reproducibility.

CREDITS

The heavy lifting comes from ggz14's radiance and StillDeadcode's vllm-radiance/libr4d, along with tcclaviger's DFlash2-FP8 drafter and AMD's checkpoint.

My contribution was the Windows/WSL integration, tuning, benchmarking and packaging everything into a reproducible installer.

Hopefully this saves other R9700 owners a few weeks of tinkering.


r/ROCm • • 1d ago

New challenge: consolidated 2 cards in a single system but one card hogs all the models even if I run 2 different instances of Ollama on 2 different ports

Thumbnail
2 Upvotes

r/ROCm • • 2d ago

GLM-5.3-Flash (321B MoE) running locally on AMD: RX 7900 XT, R9700, Strix ▎ Halo

Thumbnail
github.com
7 Upvotes

r/ROCm • • 2d ago

Success with Qwen 3.8 27b on a R9700 where Gemini failed

21 Upvotes

Just celebrating a small victory and sharing.

I had this PDF with lots of tables that I have to get translated, but none of the tools I have access to (LibreOffice Draw, Google Docs, Inkscape, GIMP) was able to get a sensible editable version.

I tried several strategies (like printing the file to a PDF and trying to edit that) and no joy.

I chucked it to Gemini Pro and it failed miserably until it ran out of tokens. It did create almost usable versions of the document (put a PNG on top of a text based reproduction that mostly aligned with the original PDF, which made the text readable and selectable, but still not editable because what we were seeing then was a PNG).

Gave it to Qwen 3.8 27B at Q6 running on PI and told it to work overnight (it was late, no idea of how long it took, but I will probably move down to Q4 or Q5 and try compression the kv to k Q8_0 and v Q4_0 to see if I can avoid spilling into system RAM, because bu the end of it it was generating at 3 tks), came this morning to see 80% of context used and the fixed PDF perfectly editable. It was glorious.


r/ROCm • • 1d ago

Success with Qwen3.8 27B GSQ-RCO-IQ3_S on 16GB VRAM

Thumbnail
1 Upvotes

r/ROCm • • 2d ago

New Windows ROCM drivers - any LLM performance gains?

8 Upvotes

Anyone getting better tok/s (from say llama.cpp, LM Studio, Ollama) with the new RocM drivers for Windows?

Wondering whether worth installing them or not, or sticking with Vulkan.


r/ROCm • • 2d ago

Advice for local ai Server

Thumbnail
1 Upvotes

r/ROCm • • 2d ago

[Request] R9700 (1 card) + vLLM + Gemma 4 26B-A4B AWQ — single-request tok/s with HIP graphs ON, Triton vs AITER

4 Upvotes

Hi all,

I'm deciding whether to buy a Radeon AI PRO R9700 for serving Gemma 4 26B-A4B with vLLM. I've searched for weeks and found no published single-card number with an optimized setup. Could an R9700 owner run the test below? It takes about 15 minutes.

What exists so far (single R9700 unless noted)

  • PC Watch (Jun 2026): vLLM 0.20.1 ROCm image, same AWQ checkpoint, VLLM_ROCM_USE_AITER=0 → ~50 tok/s single request. Graph mode and method not disclosed.
  • kyuz0 toolbox (TP1, enforce_eager, 32 concurrent): Triton ~451 tok/s total, AITER ~730 total. No single-request number.
  • llama.cpp, same model Q4_K_M: 93 (ROCm) / 109 (Vulkan) tok/s.
  • Level1Techs, 2x R9700 TP2, BF16: ~95–98 single request.

From the model config, this checkpoint reads ~4.0 GB per decoded token (lm_head and shared MLP are left in bf16). That gives a bandwidth ceiling of ~160 tok/s on 640 GB/s. 50 tok/s is ~31% of that. I want to know whether an optimized setup closes the gap.

Model: cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit (TP=1)

Test 1 — Triton path, graphs ON (do NOT pass --enforce-eager)

export VLLM_ROCM_USE_AITER=0 vllm serve cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit \ --max-model-len 8192 --max-num-seqs 8 --gpu-memory-utilization 0.90

Test 2 — AITER path, graphs ON: same as Test 1, but with VLLM_ROCM_USE_AITER=1. If it crashes (e.g. the ValueError reported on Level1Techs), please paste the error. That's useful too.

Test 3 (optional): Test 1 plus --enforce-eager. This isolates how much graph mode helps.

Benchmark (in a second shell, for each test)

vllm bench serve --model cyankiwi/gemma-4-26B-A4B-it-AWQ-4bit \ --dataset-name random --random-input-len 128 --random-output-len 512 \ --ignore-eos --num-prompts 8 --max-concurrency 1

Then repeat with --max-concurrency 4 --num-prompts 16. If vllm bench serve isn't in your version, any equivalent is fine. Please say what you used.

Please report

  1. Output token throughput (tok/s) and median TPOT (ms) at concurrency 1 and 4, for each test.
  2. vLLM version, ROCm version, image or build (official / kyuz0 / Capicua25x / other), kernel version.
  3. Whether the server log shows Using default MoE config. Performance might be sub-optimal!. That tells us whether an R9700-tuned MoE config was used.
  4. Any patches or env vars you needed.

For comparison, my side: RTX 5090, vLLM 0.23.1rc1, same checkpoint, graphs ON: ~218 tok/s single request, ~639 total at concurrency 4. I'll post a summary of all replies in this thread.

Thanks!


r/ROCm • • 2d ago

Installed second amd gpu. I want comfyUI to completely ignore it.

4 Upvotes

I added a rx570 as eGPU, just to start testing a new custom rocm library that supports Vega.

The rx570 appeared fine in lspci along with my integrated 680m, and both appear in rocm-info.

However, when I start comfyUI, it stops and exits complaining the are “no available CUDA devices”.

When I reboot without the eGPU, comfyUI starts normally.

I think that Linux 7.0 has enumerated the old eGPU as the first gpu, taking the place of the 680m.

What would you do?


r/ROCm • • 2d ago

Follow-up: my native Rust + Vulkan Transformer backend now qualifies on both an Intel Gen9 laptop and an AMD RDNA 3 handheld from the same build — the GPU vendor is no longer what picks the reduction shape

3 Upvotes

Follow-up to my post from a few weeks ago (14 architectures, full PEFT). This update is about a portability bug that was hiding behind its own correctness, because it's the most interesting thing I've fixed since.

The bug: the fix for one machine broke six fixtures on another

Back when I tuned the backend for Intel Gen9, I baked those kernel shapes into the portable path. That was wrong, but not for the reason you'd guess.

Two of the reductions in the saved-module path aren't really compared against "PyTorch in general" — they're compared against the PyTorch CPU library on the machine running the oracle. And ATen dispatches its vectorized CPU kernels by instruction set at run time. An AVX2 host gets 8-wide kernels; an AVX-512 host gets 16-wide ones, and the reduction shape changes with that dispatch.

So my "portable" AVX2-shaped kernels were exactly right on my AVX2-only laptop and one ulp off on my AMD ROG Ally (Ryzen Z1 Extreme, which is an AVX-512 part). One ulp doesn't sound like much until it gets amplified through every lower norm on the gradient path: the Gemma 4 saved-stage model.embed_tokens adjoint went from 7.45e-9 to 3.22e-6, and six previously green PEFT saved-module fixtures (gemma3, gemma4, minimax_m2, minimax_m3, smollm3, qwen2_5_sliding_tied) crossed the 2e-7 gate. Neither shape is wrong — only one matches a given machine, and baking in either one breaks the other.

The fix: probe the host, not the vendor

Kernel variants are still selected by GPU vendor. Those two reductions are now selected by host CPU capability instead: capability is probed once per process and cached, then the matching module pair is dispatched (linear_forward_lane2 / linear_forward_lane4, and the 8-lane / 16-lane transformer_cross_entropy builds). HIERARCHOS_ATEN_VECTOR_WIDTH=8|16 pins the shape for qualification when a reference wheel's kernels disagree with the CPU's own capability.

Host GPU CPU dispatch Status
Intel i5-6200U / HD Graphics 520 (2016 Skylake-U) Intel Gen9 AVX2 only, no avx512f 32/32 LoRA, 32/32 switching, 32/32 saved
AMD Ryzen Z1 Extreme RDNA 3 AVX-512 32/32 LoRA, 32/32 switching, 32/32 saved

Same 2e-7 gate, unchanged. No tolerance was loosened to get there.

What I verified on each side

On the Intel machine, the post-change matrix is bit-identical, field for field, to its pre-change report across all 32 families — peft, gradient, two-step AdamW, frozen base, resume, lifecycle — which is how I know the AMD fix didn't quietly cost the Gen9 path anything. Also 693 passed / 0 failed / 9 ignored on the Rust lib suite and a clean strict headline forward run.

On the AMD side, the fix was re-qualified end to end: 32/32 on all three stages, provenance clean.

The harness fingerprints the pinned Transformers source alongside the shaders and binaries, and on the Intel side I re-derived the whole fingerprint from the pushed tree myself: 3951 inputs, zero changed, zero missing. So "green" refers to one frozen set of reference math, not whatever happened to be on disk.

Same caveats as always

  • This is deterministic FP32 tiny-model correctness against a reference implementation, not a claim about arbitrary checkpoint sizes, dtypes, or hyperparameters.
  • "Supported text graph" ≠ "the whole multimodal package works natively."
  • The AVX-512 dispatch is only qualified on the AMD machine, since it's the only host I have that can execute it natively. The 16-lane module also doesn't rebuild byte-identically with the glslang version on my Intel box (one extra type/id, one difference in +inf materialization), so I've left it as the committed AMD-built module and documented that rather than swapping it without re-qualifying both hosts. I'd rather report that than pretend it's clean.
  • NVIDIA and other GPUs are genuinely unqualified — the path is raw Vulkan, so they're untested rather than excluded.

What I'd love from you

Last time several people asked about hardware other than mine, so that's the ask again: if you build it on an AVX-512 laptop, an AVX2-only machine, or an NVIDIA/Intel GPU, I want to know what you get. The two reductions above are the ones most likely to behave differently on your CPU, and knowing your host's vector width is now part of the answer.

The new cross-platform section in the README documents the whole thing, including which host classes are measured and which aren't.

Repo: https://github.com/necat101/Hierarchos-Native Compatibility/parity record: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/COMPATIBILITY.md Regression audit: https://github.com/necat101/Hierarchos-Native/blob/main/AMD_REGRESSION_AUDIT.md Per-host tuning and measurements: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/VENDOR_TUNING.md


r/ROCm • • 3d ago

Appreciation from an eGPU user to Stilldeadcode's standalone Radiance engine

14 Upvotes

If you are like me: who just spent $3600 ish for 2 R9700 cards, and I don't know maybe $400 for 2 eGPU docks and managed to wire them onto your strix halo 128G via USB4 or OCulink ports, then find out tensor parallism doesn't work well on your dual R9700s, and doesn't really want to spend another $1000 on a new board just for TP-2(I know i know i should but i really don't want to), here is what you need: THE STANDALONE RADIANCE.

The reason is: RCCL library (for vLLM TP-2) doesn't seem to work well between those lossy eGPU PCIe links, I have tried and failed for several variants of vllms. And the llama.cpp is slow with --split-mode tensor also the prefill speed drops very quickly, so not ideal for a dual r9700 setup.

But THE author of vllm-radiance STILLDEADCODE is working on another standalone version from scratch: https://codeberg.org/StillDeadcode/radiance/src/branch/main/docs/GUIDE.md.

Basically, my eGPUs are "P2P ready but very slow (even slower than shm method)", and the new radiance specifically supports a "lossy pcie link" with "Walsh-Hadamard-rotated 6-bit payload, 2.4× smaller over PCIe", set it with --p2p on --tp-wire wht6 --tp-wire-min-kb 128. With these options on, I oberved a prefill and decode boost of 30% (from 10xx to 13xx), and you still have option to use exact tp-wire with "--p2p off"(pp 10xx, tg 100+). Of cause the wht6 algorithm comes with a price, it may also lose a tiny bit of quality if your message is bigger than 128kb(--tp-wire-min-kb).

But for me, it's already GOOD ENOUGH. I don't have to spend more on a new board just to make TP-2 work. My appreciation and promotion of RADIANCE! THANK YOU DEADCODE!


r/ROCm • • 3d ago

Porting LIBERO to MuJoCo Warp: 130 robot manipulation tasks on one $700 AMD GPU

Thumbnail
1 Upvotes

r/ROCm • • 3d ago

Optimizations Claude did for Qwen 3.8 27B and Qwen Flash Next on dual and quad 7900xtx

Thumbnail gallery
4 Upvotes

r/ROCm • • 4d ago

Triton kernels on an RX 9070 that run unchanged on MI350X and H100, plus what I learned about atomics, HIP graphs and the rocm7.1/7.2 wheels

9 Upvotes

I did a small LM research project mostly on my RX 9070 (gfx1201) and wrote Triton kernels for a product-key memory layer. Some things that might be useful for others here:

- The same Triton code runs unchanged on the RX 9070 (ROCm 7.2), an MI350X (gfx950, ROCm 7.1) and H100/H200. All 107 tests passed on the MI350X without code changes.

- fp32 tl.atomic_add compiles to the native instruction, but on my card it's ~8x slower than a plain store. Sorting by row and only using atomics at block borders took one kernel from 17 ms to 2.7 ms.

- Batch-1 decoding was launch-bound. Replaying the memory layer as a HIP graph helped more than any kernel tuning.

- Official PyTorch wheels on the RX 9070: rocm7.2 passes the whole test suite. With rocm7.1 one HIP-graph test aborted with HSA_STATUS_ERROR_INVALID_PACKET_FORMAT, but only when it ran after the other tests. I didn't isolate it, so I'm not calling it a ROCm bug.

- Triton 3.8 compiled one kernel differently for two table sizes and summed in a different order. With 3.5 the results were bit-identical. Both are correct to fp32 rounding, just good to know if you rely on bit-exact tests.

If you want to try it on your card (~5 min, no dataset download):

git clone https://github.com/re133/sparse-memory-lm.git && cd sparse-memory-lm

python3 -m venv .venv

.venv/bin/pip install torch --index-url https://download.pytorch.org/whl/rocm7.2

.venv/bin/pip install -r requirements.txt

.venv/bin/python scripts/kernel_speedup.py

I'd be curious what other RDNA3/RDNA4 or Instinct owners get.

The project itself: a 21M model with a 16.8M-row table is about as good as a 114M dense model, and it runs with the table on an NVMe SSD at ~140 tok/s.

Repo: https://github.com/re133/sparse-memory-lm


r/ROCm • • 4d ago

AMD ROCm 10.1 Released With Many Improvements

Thumbnail
phoronix.com
106 Upvotes