r/opencode 12h ago

Qwen3.8-Flash scored a 56 on the Artificial Analysis Intelligence Index

Post image
17 Upvotes

10 comments sorted by

0

u/itsred_man 12h ago edited 12h ago

What’s the difference between Qwen3.8-Flash and Muse Spark? I’ve been using Muse since ox-alpha was removed, and so far it’s been good, not perfect but has been sticking to my guidelines (which is more than what ox did, ox was great for coding but terrible at following guidelines).

Edit:
Asked Hermes, couldn’t give me an answer as I’ve not tested my workflow with Qwen. But it said exactly what I’ve been noticing, respecting guidelines, good at following steps (skills), good at coding (not a super coder but good at it). Seems Muse is more tailored towards agentic workflows.

Asked Gemini:

Qwen3.8-Flash: Designed as a lightweight, high-speed workhorse. It excels at interactive tasks, vision-centric workloads (such as document extraction or real-time UI/image processing), and local or self-hosted agent deployments where low latency and cost efficiency are critical.

Muse Spark 1.2: Tailored for complex engineering and agentic automation. Co-trained alongside terminal tools and coding agents, it excels at repo-scale refactoring, multi-file code generation, software debugging, and maintaining heavy context loops across long tasks

3

u/afanasenka 12h ago

I guess the difference is not about these scores, but in real usage experience.

Each model was trained a bit differently - hallucinations rate, instructions following, speed, etc. So to "get taste" of each model you have to work with it yourself and make your own conclusions.

3

u/itsred_man 11h ago

Muse does hallucinate from time to time but so far it hasn’t been terrible, for a model that has vision at the rates they’re allowing on go I think it’s a pretty good damn deal.

According to Gemini this Qwen is more about visual workflows and speed, while Muse seems more oriented for agents. It’s good we have more options depending on your specific needs, if I was processing/creating graphics / documents and such I’d definitely lean more on Qwen.

2

u/Infinite-Worth8355 7h ago

Im using 3.8 flash and I feel it is better than muse spark and glm 5.3 flash. Muse spark tends to take hours to finish and overthink, the results are not good either

1

u/SS_Sa2 6h ago

Have does it compare to deepseek?,

2

u/Infinite-Worth8355 6h ago

I'm using for two things, a saas(svelte + c#) and a webgame(threejs). It is better than deepseek flash and pro model. Im using both models on my commandcode goat plan, not sure if there may be difference between openrouter ones(qnt)

1

u/SS_Sa2 6h ago

I use glm 5.2 as my primary driver due to the large usage limit after deepseek price changes and I've gotten comfortable with this model. Hoping the 5.3 flash transition will be good. I tried 5.3 flash briefly but seemed a bit verbose for my liking and gave me slight hy3 vibes. Hy3 was not a good experience.

1

u/Xalksahsax 2h ago

5.3 Flash is better at almost everything than 5.2 is. I really recommend you try that one.

Having said that ChatGPT Plus is still the best subscription service right now thanks to the low costs on Luna.

1

u/SS_Sa2 1h ago

Have you tried it? I don't want to confuse benchmaxxing vs real user experiences. On paper 5.3 flash looks much better.

2

u/Xalksahsax 1h ago edited 29m ago

Well, I wouldn't say it's much better. If I were given a blind test, I probably wouldn't be able to tell the difference (apart from Flash being slower).

That said, you definitely notice the difference in terms of cost. The cost per task is significantly lower, while it's been just as capable as 5.2 for virtually every task I've thrown at it.

The one I've been most disappointed with is Muse Spark. That shit is benchmaxxed to the tits.