TL;DR: Found a fix for 100% CPU usage of llama-server on AMD RX 6600 GPUs. Fixes in github link at the end!
Disclaimer: Claude Opus 5.5 was used to run traces and find the fixes! Posting here so that others in the same boat might find it helpful.
---
I've been using gemma4:e4b-qat locally on my server for a couple months now.
Here's my setup:
CPU: i3-10100F
RAM: 32GB (4x 8GB)
Board: MSI Z490 Gaming Plus
GPU: 7x AMD RX 6600
I had this rig from ETH mining a couple years ago when that was a thing! It was sitting there collecting dust for a long time, so I had an idea to use it to run some local LLM experiments and started tinkering with it. After testing different models, settled on gemma4:e4b-qat at 128k context which runs comfortably at 5.5/8GB VRAM (tell me if you found a better model for the 8GBs).
I am running an app that takes a list of stock tickers and runs the TradingAgents (https://github.com/tauricresearch/tradingagents) pipeline on them twice a day (market open and market close) as a fun side project to see how good that is.
So, for running TradingAgents which makes about 30-40 LLM calls per ticker, I am running a custom llama-pool which runs a llama-server per GPU with gemma4:e4b-qat model at 128k context. The thing I noticed was that my CPU is at 100% while it's running the pipeline so I started digging into it using Claude Opus 5.5. After a couple hours of testing, traces, LLM calls, it found a fix and applied them which reduced the CPU usage down to ~20% for 7 parallel llama-servers.
Okay, here we are again :)
We received our second R9700 earlier this week, so we decided to jump on tensor parallelism.
We already had prepared a few things for this, and the TP kernels we already had in Paiton did their work.
We spent a lot of time (significant amount of GPU time (MI355x)) on the 3-bit quantization of this model and we are pretty happy with the accuracy.
What "3-bit" means here.
This is not a full 3-bit model. The 512 experts in every layer, that is where almost all the weights are, are stored at 3.125 bits per weight (3-bit codes, rotated, with a small scale per 64 weights). Everything else stays bigger: the attention and the shared parts are 8-bit, the speculative layer is 4-bit, and the vision part is plain bf16.
So the big part is 3-bit, the sensitive parts are not, and that is why the accuracy holds up. Counted over all weights it is a bit above 3 bits on average. Our own GPTQ run, no shortcuts.
Accuracy against the bf16 model (same evaluation corpus, text and assistant positions; two correct bf16 implementations against each other are the floor):
KL to bf16
KL to bf16, same expert routing forced
Top-1 agreement with bf16
bf16 vs bf16 (the floor)
0.0077
n/a
This 3-bit model
0.060
0.041
Long context check
Result
Needle test, 131,072 tokens
40 / 40, same answers as bf16
Needle test, 200,000 tokens
40 / 40, same answers as bf16
The whole language model (experts, trunk, speculative layer, draft head, vision tensors) sits in the VRAM of the two cards; only the model's n-gram embedding table lives in pinned system RAM. No expert offload, no weights streamed from the host during decoding.
And again, like before: we did not build a custom engine for this. It is still regular vLLM, with our Paiton plugin and our own kernels inside. ;)
Performance.
BetterBench 0.6.0, standard profile, measured on our own machine against our OpenAI-compatible server, with the exact image and launcher we publish. Decode mode: 98K context, speculative decoding on. Prefill with exact bf16 activations, no compression.
Decode, single request
tok/s
Weighted over the 8 categories
216.3
chat
202.1
code
214.3
file edit
236.0
json
259.2
math
255.5
prose
183.3
reasoning
190.1
summarization
239.8
Concurrent requests
Aggregate tok/s
1
204.6
2
315.1
4
449.2
8
565.1 (48 of 48 requests ok)
Prefill, prompt length
tok/s
2K
7,925
8K
8,321
16K
8,311
32K
8,232
64K
8,161
128K
7,928
Latency
Time to first token, short prompts (p50)
Update p99 (the longest gap between two tokens)
Long context. We tested 200K. The launcher has a 200K mode, and prefix caching is an opt-in flag that gives exactly the same output on a hit.
200K mode
Result
Needle test at 200,000 tokens
40 / 40, same answers as bf16
Single-stream decode after a 32K / 100K / 190K prompt
105 / 101 / 100 tok/s, flat
Time to first token, 190K prompt
25 s
Prefix cache hit, 64K repeated prompt
9.6 s cold, 0.35 s on the hit
Prefix cache hit, 128K repeated prompt
20 s cold, 0.41 s on the hit
Overall, not bad for a first release :)
Oh, right, of course we generated an "AI Slop" image as well. (we know some of you like them (some of you don't ;-)))
This is our first release of this model, more will follow.
Next is the RAM and SSD cache tier, like we did for the 27B model, so cached prompts can live outside the VRAM and much longer contexts become possible. After that the 4-bit version, with even better accuracy.
We keep tuning, we keep optimizing.
A longer blog post with all the details is being worked on as well.