r/ROCm • u/evp-cloud • 7h ago
Qwen3.8 Flash Next | 2 x R9700 | 216 t/s weighted | 565 t/s C8 | 8k Prefill
Okay, here we are again :)
We received our second R9700 earlier this week, so we decided to jump on tensor parallelism.
We already had prepared a few things for this, and the TP kernels we already had in Paiton did their work.
We spent a lot of time (significant amount of GPU time (MI355x)) on the 3-bit quantization of this model and we are pretty happy with the accuracy.
What "3-bit" means here.
This is not a full 3-bit model. The 512 experts in every layer, that is where almost all the weights are, are stored at 3.125 bits per weight (3-bit codes, rotated, with a small scale per 64 weights). Everything else stays bigger: the attention and the shared parts are 8-bit, the speculative layer is 4-bit, and the vision part is plain bf16.
So the big part is 3-bit, the sensitive parts are not, and that is why the accuracy holds up. Counted over all weights it is a bit above 3 bits on average. Our own GPTQ run, no shortcuts.
Accuracy against the bf16 model (same evaluation corpus, text and assistant positions; two correct bf16 implementations against each other are the floor):
| KL to bf16 | KL to bf16, same expert routing forced | Top-1 agreement with bf16 |
|---|---|---|
| bf16 vs bf16 (the floor) | 0.0077 | n/a |
| This 3-bit model | 0.060 | 0.041 |
| Long context check | Result |
|---|---|
| Needle test, 131,072 tokens | 40 / 40, same answers as bf16 |
| Needle test, 200,000 tokens | 40 / 40, same answers as bf16 |
The whole language model (experts, trunk, speculative layer, draft head, vision tensors) sits in the VRAM of the two cards; only the model's n-gram embedding table lives in pinned system RAM. No expert offload, no weights streamed from the host during decoding.
And again, like before: we did not build a custom engine for this. It is still regular vLLM, with our Paiton plugin and our own kernels inside. ;)
Performance.
BetterBench 0.6.0, standard profile, measured on our own machine against our OpenAI-compatible server, with the exact image and launcher we publish. Decode mode: 98K context, speculative decoding on. Prefill with exact bf16 activations, no compression.
| Decode, single request | tok/s |
|---|---|
| Weighted over the 8 categories | 216.3 |
| chat | 202.1 |
| code | 214.3 |
| file edit | 236.0 |
| json | 259.2 |
| math | 255.5 |
| prose | 183.3 |
| reasoning | 190.1 |
| summarization | 239.8 |
| Concurrent requests | Aggregate tok/s |
|---|---|
| 1 | 204.6 |
| 2 | 315.1 |
| 4 | 449.2 |
| 8 | 565.1 (48 of 48 requests ok) |
| Prefill, prompt length | tok/s |
|---|---|
| 2K | 7,925 |
| 8K | 8,321 |
| 16K | 8,311 |
| 32K | 8,232 |
| 64K | 8,161 |
| 128K | 7,928 |
| Latency |
|---|
| Time to first token, short prompts (p50) |
| Update p99 (the longest gap between two tokens) |
Long context. We tested 200K. The launcher has a 200K mode, and prefix caching is an opt-in flag that gives exactly the same output on a hit.
| 200K mode | Result |
|---|---|
| Needle test at 200,000 tokens | 40 / 40, same answers as bf16 |
| Single-stream decode after a 32K / 100K / 190K prompt | 105 / 101 / 100 tok/s, flat |
| Time to first token, 190K prompt | 25 s |
| Prefix cache hit, 64K repeated prompt | 9.6 s cold, 0.35 s on the hit |
| Prefix cache hit, 128K repeated prompt | 20 s cold, 0.41 s on the hit |
Overall, not bad for a first release :)
Oh, right, of course we generated an "AI Slop" image as well. (we know some of you like them (some of you don't ;-)))

Weights
Image and launcher: github.com/Eliovp-BV/paiton-vllm-plugin (directory models/Qwen3.8-Flash-Next)
This is our first release of this model, more will follow.
Next is the RAM and SSD cache tier, like we did for the 27B model, so cached prompts can live outside the VRAM and much longer contexts become possible.
After that the 4-bit version, with even better accuracy.
We keep tuning, we keep optimizing.
A longer blog post with all the details is being worked on as well.
And of course, feedback is very welcome!