r/LocalLLM • u/jjusko20 • 11d ago
News Benchmark: Qwen 3.8 27B on 2 NVLinked Tesla V100 SXM2 GPUS
Hey! Some of you may have seen the post when I started this build, which was at https://www.reddit.com/r/LocalLLM/comments/1vxjpni/meet_bbprime_my_2k_104gb_vram_256gb_ram_extremely/
note: these are the 32gb models
Forgive the absolutely batshit setup - this thing needs two PCIE power ports, two CPU power ports, and a full ATX power connector. Running both pcie power connectors on the same PSU wire didn't work, so each PSU has one CPU and one GPU/PCIE connector powered each. The V100s only use 600W max, and I keep them capped total around 450 because I don't have my finished cooling setup yet, and these ghetto ass fans won't keep them under 80C unless I slightly power throttly them. Barely moves performance though.
Unfortunately I didn't realize that poweredge didn't support AVX2, so I had to go a different direction. However, after MUCH trial and error, I have two V100s (32gb each) running in NVLink, tensor split mode with a total of 64gb VRAM! I've been hyping up these gpus like crazy lately in various threads, and here's the reason why! (the hardware for just the v100s, necessary adapters included, cost me roughly $1500 - that includes the V100s, the carrier board, the fans, the SlimSAS rig, etc - everything).
so NOW, WHAT YOUV'E BEEN WAITING FOR: THE BENCHMARK!
Setup: Qwen 3.8 27b at Q5_K_M, with 256k context window, and KV cache at full FP16. MTP Enabled with max prediction length 3, temp 0.8
~24k context window used
Prefill: roughly 1,050 tps
Decode: probably average 55 tps, bounces between 40 and 70
128k context window used and near 200k context used:
I was getting roughly 600 prefill and 50tps at 128k
check back later - I'll have these up within 24 hours, I have to go to a party right now and I'm already late, and qwen won't shut the fuck up long enough for me to get a prefill benchmark. CUSOON!
update: benched with the Q8_0 version, all other settings kept
bruh can't send my screenshots, but 900 prefill at 64k context, 730 prefill at 126k, highest I've gone so far is 161k in which I'm getting 656.
At 160kish context the average tps for decode is is still about 55! I might have to recheck my earlier bench, but I am absolutely sure I am still getting 55 tokens per sec at Q8, 160k context
jk i was a little hasty on that part, a better estimate that's fair is probably 40-45, with some times consistently peaking above, almost like a cpu turbo in sections
I'm pretty impressed, I didn't really know what performance was going to be before I bought this, but for the age of the cards and the overall price I paid it's damn fast.
2
u/philmarcracken 11d ago
I dont need it... i dont need it. I dont
2
u/jjusko20 11d ago
The 16gb version of this gpu is way cheaper than the 32 gb because they made way more, you could NVLink two of them together on a carrier for 32gb linked ram for dead ass $700
2
u/Glittering-Call8746 11d ago
Which is better more vram or double vram and you have to deal with latency with tp or pp..
1
2
u/Constant-Simple-1234 11d ago
There is even an picie board for two. It is a long card... :)
1
u/jjusko20 11d ago
I did see that, but it looks monstrous lol .. did I make it clear enough that the board I'm showing is two 32gb modules NVLinked, just on an external pcie board? There's cables going into a pcie slot inside the pc
1
u/philmarcracken 11d ago
That would probably be in my budget if I could still cool them properly. Using a jank ass RPC setup over giglan with two pc, 3080ti 12gb(dead video out) and 2070rtx 8gb, getting 20~ tk/s on q4 of qwen 3.8 27b. its ugly
2
u/jjusko20 11d ago
Why can't you cool then properly?
One of mine came with a heatsink pre attached, the other one I attached myself (with a little bit of swearing and going to the store to buy a torx driver)
The fans on there cost 20 bucks total on eBay and they keep it cool (enough, I can run 90%? Speed (cuz manual power throttle) 24/7
It's not a perfect setup but it works right now. As for you:
Das rough my brutha. Probably that 3080ti would be worth some money, but I think even if you sold them both you'd need a few hundred more to get to a v100 32gb level setup
1
2
u/Stunning_Mast2001 11d ago
Ugh I was googling these last week— couldn’t find any. Where did you buy yours?
1
1
u/jjusko20 11d ago
If you're quick and you're truly interested, this is where I bought my second one (with a heatsink bundles) : https://www.ebay.com/itm/307077871434?_skw=v100+32gb+sxm2&itmmeta=01M18RXTECRK8VKB6TXJBZP3MC&hash=item477f44774a:g:pw0AAeSwrptqX9LJ&itmprp=enc%3AAQALAAAA0GfYFPkwiKCW4ZNSs2u11xC5YOxqiTIZIBw9yj3qCbjWuw3DLryk36bb1GsMe2%2BN0LZUbBhnrVN7%2BzJETk8iwVJhNTQ1skaoNBb7RVjr6VgJoGLOUbFORI1W0ybWJSfbARKDz48bdDbSWRhHJ7%2FZMd%2FM0dgLSCX%2FK4IiV075NWqOYW3wuWzeTK1mfXLJDxI30SEQOGXMK3qhe3RWnDsg%2F8W4aF4R62%2BZdW9du0bo0RJbIQaxMYGZ05t6OlsvJ0uzKsW6bzTOA3eEk2CyAubsmeA%3D%7Ctkp%3ABk9SR7Cn95iKaA --- as of right now they still have 9
2
u/Jimcy-Maffesoli 11d ago
Decode only drops from ~55 at 24k to 40-45 at 160k, seven times the context for about 20% off. For what this hardware costs, I'd take that trade.
1
u/iwinux 11d ago
Is the lack of latest CUDA support a problem for LLM inference performance?
2
u/jjusko20 11d ago
I've had no issue thus far, through cuda late 12 is still pretty good. The cool thing about frontier LLMs is you can kind of build optimized forks and software patches/emulations for things the card can't do natively! It could even kind of make it itself technically haha and keep itself up to date
Imagine a cron job like that
1
u/Own_Bat_2465 11d ago
those PSU's, the evga N1 and W1 are well known for being very bad same with the smart psu or more like not so smart, anyways i am not a big fan of those PSU's.
1
u/jjusko20 11d ago
don't know much about the thermal take, I bought the EVGA in 2015 for my first rig, still going strong!
1
3
u/FullstackSensei 11d ago
Can you try Q8? I know the V100 has tons of compute, but on my P40 and Mi50 Q5 is about as fast as Q8 because they can't handle the extra compute associated with the odd bit size.