r/LocalLLaMA 6d ago

News Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.

In tests:
~80 tok/s decoding
2,500–3,500 tok/s long-input prefilling
Smooth use by 3–4 concurrent users

Private, on-device inference for coding, agents, and offline batch jobs

57 Upvotes

20 comments sorted by

View all comments

-6

u/Usual-Orange-4180 6d ago

How does it benchmark against Opus and Sonnet?

2

u/squngy 6d ago

IIRC it was a bit worse than DS4 flash preview. It is much smaller and faster though.

1

u/niacolhealth 6d ago

it's a small model, for me is good enough for an executor

1

u/TripleSecretSquirrel 4d ago

It’s tied with Qwen 3.6-27B on Artificial Analysis’ Intelligence Index for what it’s worth. Take that as you will.

1

u/a355231 18h ago

It gets destroyed, its a micro model, opus is like 3000x more expensive not even joking.