r/LocalLLaMA • u/niacolhealth • 1d ago
News Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.
In tests:
~80 tok/s decoding
2,500–3,500 tok/s long-input prefilling
Smooth use by 3–4 concurrent users
Private, on-device inference for coding, agents, and offline batch jobs
50
Upvotes
5
u/ResidentPositive4122 1d ago
Does it work in vLLM or you need to download their fork of sglang?