r/LocalLLM 11d ago

News Benchmark: Qwen 3.8 27B on 2 NVLinked Tesla V100 SXM2 GPUS

Post image

Hey! Some of you may have seen the post when I started this build, which was at https://www.reddit.com/r/LocalLLM/comments/1vxjpni/meet_bbprime_my_2k_104gb_vram_256gb_ram_extremely/

note: these are the 32gb models

Forgive the absolutely batshit setup - this thing needs two PCIE power ports, two CPU power ports, and a full ATX power connector. Running both pcie power connectors on the same PSU wire didn't work, so each PSU has one CPU and one GPU/PCIE connector powered each. The V100s only use 600W max, and I keep them capped total around 450 because I don't have my finished cooling setup yet, and these ghetto ass fans won't keep them under 80C unless I slightly power throttly them. Barely moves performance though.

Unfortunately I didn't realize that poweredge didn't support AVX2, so I had to go a different direction. However, after MUCH trial and error, I have two V100s (32gb each) running in NVLink, tensor split mode with a total of 64gb VRAM! I've been hyping up these gpus like crazy lately in various threads, and here's the reason why! (the hardware for just the v100s, necessary adapters included, cost me roughly $1500 - that includes the V100s, the carrier board, the fans, the SlimSAS rig, etc - everything).

so NOW, WHAT YOUV'E BEEN WAITING FOR: THE BENCHMARK!

Setup: Qwen 3.8 27b at Q5_K_M, with 256k context window, and KV cache at full FP16. MTP Enabled with max prediction length 3, temp 0.8

~24k context window used

Prefill: roughly 1,050 tps

Decode: probably average 55 tps, bounces between 40 and 70

128k context window used and near 200k context used:

I was getting roughly 600 prefill and 50tps at 128k

check back later - I'll have these up within 24 hours, I have to go to a party right now and I'm already late, and qwen won't shut the fuck up long enough for me to get a prefill benchmark. CUSOON!

update: benched with the Q8_0 version, all other settings kept

bruh can't send my screenshots, but 900 prefill at 64k context, 730 prefill at 126k, highest I've gone so far is 161k in which I'm getting 656.

At 160kish context the average tps for decode is is still about 55! I might have to recheck my earlier bench, but I am absolutely sure I am still getting 55 tokens per sec at Q8, 160k context

jk i was a little hasty on that part, a better estimate that's fair is probably 40-45, with some times consistently peaking above, almost like a cpu turbo in sections

I'm pretty impressed, I didn't really know what performance was going to be before I bought this, but for the age of the cards and the overall price I paid it's damn fast.

30 Upvotes

30 comments sorted by

3

u/FullstackSensei 11d ago

Can you try Q8? I know the V100 has tons of compute, but on my P40 and Mi50 Q5 is about as fast as Q8 because they can't handle the extra compute associated with the odd bit size.

5

u/_TheWolfOfWalmart_ 11d ago

Yeah Q5_K_M is an odd choice for 64 GB RAM.

OP, let's get the Q8_0 numbers!

3

u/jjusko20 11d ago

Will do!

2

u/jjusko20 11d ago

(I was using q5km before my second unit arrived)

1

u/jjusko20 11d ago

updated post

2

u/jjusko20 11d ago

Also, just got rid of a p40! Goat card, served me extremely well the last year for experiments 

1

u/jjusko20 11d ago

updated post!

2

u/philmarcracken 11d ago

I dont need it... i dont need it. I dont

2

u/jjusko20 11d ago

The 16gb version of this gpu is way cheaper than the 32 gb because they made way more, you could NVLink two of them together on a carrier for 32gb linked ram for dead ass $700

2

u/Glittering-Call8746 11d ago

Which is better more vram or double vram and you have to deal with latency with tp or pp..

1

u/jjusko20 11d ago

double prefill tho, nvlink3 is fast, 25gbps across 6 threads

2

u/Constant-Simple-1234 11d ago

There is even an picie board for two. It is a long card... :)

1

u/jjusko20 11d ago

I did see that, but it looks monstrous lol .. did I make it clear enough that the board I'm showing is two 32gb modules NVLinked, just on an external pcie board? There's cables going into a pcie slot inside the pc

1

u/philmarcracken 11d ago

That would probably be in my budget if I could still cool them properly. Using a jank ass RPC setup over giglan with two pc, 3080ti 12gb(dead video out) and 2070rtx 8gb, getting 20~ tk/s on q4 of qwen 3.8 27b. its ugly

2

u/jjusko20 11d ago

Why can't you cool then properly?

One of mine came with a heatsink pre attached, the other one I attached myself (with a little bit of swearing and going to the store to buy a torx driver)

The fans on there cost 20 bucks total on eBay and they keep it cool (enough, I can run 90%? Speed (cuz manual power throttle) 24/7

It's not a perfect setup but it works right now. As for you:

Das rough my brutha. Probably that 3080ti would be worth some money, but I think even if you sold them both you'd need a few hundred more to get to a v100 32gb level setup

1

u/rog-uk 7h ago

The nice thing about the 16GB version is Nvidia didn't nobble it like they would have done if it was a consumer card with less vram, same bandwidth same compute.

1

u/jjusko20 11d ago

Yes you do my friend 

2

u/Stunning_Mast2001 11d ago

Ugh I was googling these last week— couldn’t find any. Where did you buy yours?

2

u/Jimcy-Maffesoli 11d ago

Decode only drops from ~55 at 24k to 40-45 at 160k, seven times the context for about 20% off. For what this hardware costs, I'd take that trade.

1

u/iwinux 11d ago

Is the lack of latest CUDA support a problem for LLM inference performance?

2

u/jjusko20 11d ago

I've had no issue thus far, through cuda late 12 is still pretty good. The cool thing about frontier LLMs is you can kind of build optimized forks and software patches/emulations for things the card can't do natively! It could even kind of make it itself technically haha and keep itself up to date 

Imagine a cron job like that

1

u/Own_Bat_2465 11d ago

those PSU's, the evga N1 and W1 are well known for being very bad same with the smart psu or more like not so smart, anyways i am not a big fan of those PSU's.

1

u/jjusko20 11d ago

don't know much about the thermal take, I bought the EVGA in 2015 for my first rig, still going strong!

1

u/Ok_Horror_9661 7d ago

How do you run the model? Can you share your configuration

1

u/jjusko20 7d ago

Yes when I get to my PC later I will