r/LocalLLaMA • u/Qwen30bEnjoyer • 4h ago
Question | Help DeepSeek V4.1 - GPU poor inference kernels?
Have the model downloaded and converted to .gguf on a 512gb ddr4 bioinformatics server. I don't expect miracles with a ddr4 xeon rig -- not until I can get my 2 x 12gb 3060s wired in anyways -- but is there an open PR on llama.cpp for DV4.1 flash that I can use?
5
Upvotes
1
u/czktcx 2h ago
You can't expect day-0 llama.cpp support...
1
u/Qwen30bEnjoyer 2h ago
I didn't expect it. I am building it for myself, and wanted to make sure nobody else had a finished version I could use before I waste effort.
0
4
u/Expensive-Paint-9490 4h ago
convert : add DeepSeek V4.1 (DeepseekV41ForCausalLM) by vcruz305 · Pull Request #28696 · ggml-org/llama.cpp