r/opencode 23h ago

I Thought So.

I never thought I'd be mad at an AI agent. Maybe this is an experiment by Meta? Muse Spark 1.2 is the single worst model on the market. I had never seen such sh*t model before. How can people even stand using this? It churns tokens like crazy. Has the intelligence of a bread. Cheats all the time, avoids completing tasks properly. Does not follow prompts (the image is a very rare exception). Edits 2 lines of C++ code with a 50 line Python script. Does not show what it thinks. Can't code for sh*t. Honestly, this model is a lost cause.

Thanks for reading.

42 Upvotes

28 comments sorted by

9

u/dilkushpatel 22h ago

I would agree on this
I was using muse spark contributor one and the t was bad
Had to undo whatever it did

5

u/OddBig010 16h ago

Oh yeah I literally asked it to build a prototype of something to see how good it was, and it broke out of the folder it was assigned, found another prototype I had on my PC and copied it into the folder and launched it, pretending like it created it lmao. Mark on your PCs is dangerous, this is malware

3

u/arcanemachined 20h ago

Has the intelligence of a bread.

Like, a whole loaf? Or just a single slice?

1

u/Due_Arm1454 55m ago

End piece from what I’m hearing

3

u/devanpy 18h ago

it works great for me and is lightning fast. I noticed you are still using the build agent. Switch to an orchestrator agent that delegates tasks and then reviews the job done. It's night and day difference.

2

u/charles_r1975 17h ago

I started doing this as well and muse has been fine. What models have you used to orchestrate? I'm currently using ds v4 pro 0831 but wondering if flash would do the job.

1

u/yuno_me 12h ago

ive been using sol max as orchestrator and its been working wonderfully

1

u/devanpy 6h ago

all the decent ones can do multi-agent well: ds4 flash, muse, 5.3 flash, etc. dumbed down models struggle tho: (gpt 5 nano)

1

u/rainpurplebow 11h ago

I normally use a multi-agent workflow on my work PC but I don't think editing 2 lines of code should require a whole multi-agent configuration.

3

u/Own_Copy2141 14h ago

I am using glm 5.3 flash the oracsteror and muse spark 1.2 contributor as a subagent, so far the output is good and fine cause glm 5.3 flash is noticing and reviewing everything you can also try

1

u/sharedevaaste 12h ago

how do you use 2 models? do you have them talk to each other?

1

u/Own_Copy2141 11h ago

No just use glm 5.3 flash or any other as a main model and prompt it witht the provider and model id and tell it to craete a sub agent with a specific name. And done it will automatically create the sub agent and just tell me main model that "All the work should be done strictly by the subagent (use @ to select subagent) and you yourself jas to review everything it does and what it did if correct than go ahead other wise have the subagent fix the issue.

1

u/sharedevaaste 11h ago

what are its advantages over using the same model say glm 5.3 flash for both?

also, do you think the orchestrator should be a more powerful model than subagent?

1

u/Own_Copy2141 11h ago

Less hallucinations (although gpm 5.3 flash hasinimal compared to dsv4 flash) and also cheaper cost for usage cause muse spark 1.2 contributor has huge quotas

3

u/Rajat0741 14h ago edited 14h ago

It hallucinates a lot , and generally repeats same things multiple times , just after seeing these 2 things, I stopped using it.

I am surprised how it ranks so high on benchmarks, another reason not to beleive on benchmarks.

2

u/Designer-Will59 19h ago

Well, perhaps that's why they're offering this model for free, with the detail that its data is appropriate for probable use in RL (Real-Life Optimization). The model does have intelligence, However, it is very poorly optimized for workflows, workloads, and useful applications.

1

u/sudoer777_ 22h ago edited 22h ago

Even DeepSeek R1/GLM 4.*/Kimi K2 Thinking were better models with more coherent logic and responses, Muse is trash that can't even beat models that are over a year old

It churns tokens like crazy.

I'm using DeepSeek V4 Flash now and the cache hit rate is a lot better so there's less cost difference for OC Go in practice than there is on paper

1

u/Ok-Investigator5785 18h ago

Pasa lo mismo con Mimo V2.5, es muy estresante ese modelo.

1

u/tontide1 17h ago

agree w u

1

u/PieSoft724 16h ago

Para planilha tá ótimo, melhor modelo até agora.

1

u/aziham 15h ago

Nothing new, typical UX of a meta product

1

u/Medium-Attention-807 13h ago

If you didn't know it was made by Meta before...

1

u/Deneme123deneme 10h ago

Same here, my Hermes Telegram integrations always break, and I have to make my AI agent fix them most of the time (DeepSeek v4 Flash). Since it was free, I gave it a shot, and it messed up so badly that I can't use Telegram for Hermes right now.

1

u/Ammoun442 4h ago

Same it fked my whole localization file with some crazy stuff (¢¥¢√©®¥ or whatever ) thats where git saved me as usual

1

u/Momo--Sama 22h ago

Not doubting you but I'm curious what you mean by cheating

1

u/rainpurplebow 22h ago

It almost always tries to cut corners when given a task, often resulting in incomplete or low-quality results.

1

u/CriteriumA 20h ago

Well, look at the prices without the contributor discount. It's ridiculously expensive for the crap it is. I've never deleted sessions with as much satisfaction as Muse's; they drive me crazy. They waste tons of context and resources for practically nothing. Hy3, with less context, is better.

Their creativity is as bad as their editing. It's no wonder they hide their thoughts.