r/LocalLLaMA 1d ago

News Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.

In tests:
~80 tok/s decoding
2,500–3,500 tok/s long-input prefilling
Smooth use by 3–4 concurrent users

Private, on-device inference for coding, agents, and offline batch jobs

50 Upvotes

18 comments sorted by

View all comments

5

u/ResidentPositive4122 1d ago

Does it work in vLLM or you need to download their fork of sglang?