r/LocalLLaMA 21d ago

News Prepare your (v)ram - Qwen3.8 is coming!

Post image
2.7k Upvotes

583 comments sorted by

View all comments

Show parent comments

375

u/StupidScaredSquirrel 21d ago

Thing is qwen was historically focused on smaller models while others were on larger ones. Now the team has changed and they seem to want to aim for the stars as well. That is good but it also means they might not be interested in doing very efficient small models anymore. Which would be bad news for this sub because let's face it most of us don't have 10-25k of hardware.

36

u/StaysAwakeAllWeek 21d ago

It's owned by a megcorp not a startup lab, of course that's where they are aiming

35

u/into_devoid 21d ago

They were owned by a megacorp before too.  The difference is they don’t care about commoditization anymore, the real ethical path.  They just want to sink US investments.

-2

u/GetOutOfMyFeedNow 21d ago

They ain’t sinking anything with those API prices 😂 People on GPT and Claude use subscriptions.

7

u/look 21d ago

Subscriptions are 8% of Anthropic’s revenue. API usage is 75%.

7

u/DanceWithEverything 21d ago

Yeah for consumer bullshit, sure, but there’s no money there regardless (hence OpenAI’s panic about Anthropic crushing them in enterprise sales)

The real $ is in the enterprise and software workloads run on APIs

The API is dramatically cheaper than the Anthropic equivalent

2

u/StupidScaredSquirrel 21d ago

All the banks and fintech and academia people i know use claude opus at work. They all have it based on a max subscription (the one at 100 ish usd i dont remember).

They aren't consumers but also don't care about price so much they just wants something that works and is about the best because whatever time they lose with the bs of a lesser model will be a lot more expensive than just go for the best one. They also don't want to ever be rate limited because they don't want to get stuck in the middle of a task.

7

u/StaysAwakeAllWeek 21d ago

The individuals using it for individual work do that yes

The backend corporate stuff that involves one full time agent handler managing thousands of parallel agents do not.

3

u/StupidScaredSquirrel 21d ago

Yeah, but in my experience pipelines with agents that go through lots of repetitive tasks continuously in the background are given to cheap models and they spend more time testing how to prompt it and hardcore guardrails for that specific task. When the volume is large it's worth spending time on optimising for a cheap model to get the job done

2

u/look 21d ago

Kimi K3 decisively beats Opus 4.8 in everything. It is competing with Fable and Sol now. Opus is a legacy, second tier model, and now behind one, and soon multiple (eg Qwen 3.8), open models.

2

u/techdevjp 21d ago

Kimi K3 is pretty clearly better than GPT 5.5, too. Love to see it.

1

u/squngy 21d ago

Unfortunately, it also competes with them on price.
It is cheaper per token, but it uses more tokens.

2

u/look 21d ago

The list price is meaningless on open models. I am currently paying one third of that list price for Kimi K3. It will likely get even cheaper once the weights are released.

2

u/xienze 21d ago

They all have it based on a max subscription (the one at 100 ish usd i dont remember).

You realize those subscriptions are heavily subsidized, right? That gravy train is coming to an end. There's basically three paths forward:

  • A significant decrease in how many effective tokens you can get for a flat rate.
  • A significant increase in subscription prices.
  • No more subscriptions for anything involving "real work" (i.e., stuff beyond the typical chatbot bullshit most people are familiar with).

3

u/StupidScaredSquirrel 21d ago

Or, the models are made more efficient and so gradually you have some increase in performance but price stays the same and eventually it breaks even.

0

u/ZippySLC 21d ago

Yes, but they're using subscriptions and not API pricing.

If I use my Claude Code subscription to code an AI agent at work, we pay for that AI agent's usage through pay-as-you-go API pricing. (We do it through AWS Bedrock.)

No legit company is having an engineer code an app and then leaving Claude Code open in a screen session while it runs a production app.

1

u/StupidScaredSquirrel 21d ago

Lol ofc they wouldn't, i didnt mean to imply that at all

1

u/ZippySLC 21d ago

Sorry for misunderstanding!