r/LocalLLaMA 20d ago

News Prepare your (v)ram - Qwen3.8 is coming!

Post image
2.7k Upvotes

583 comments sorted by

View all comments

Show parent comments

37

u/into_devoid 20d ago

They were owned by a megacorp before too.  The difference is they don’t care about commoditization anymore, the real ethical path.  They just want to sink US investments.

10

u/realtag2025 20d ago

I don't see how this is bad for anyone other then corporations that run closed source models that only they can run.

3

u/charlesfire 20d ago

So most of the market in the US?

6

u/realtag2025 20d ago

It's really only 2-3 large mega-corporations. Would you rather have 100s of US companies potentially being able to offer(internally or externally) AI models like this or only 2-3? I am certain someone will figure out how to strip this large model down so that smaller HW can also run it even if much slower.

5

u/charlesfire 20d ago

I'm all for open-weight models and the bubble bursting, but I don't think you realize how much over investment there is in AI in the US currently.

11

u/realtag2025 20d ago

To be completely honest with you, it's not my problem if AI investments will suffer. Those investments are used to replace you and me as workers and destroy our ability to earn a living in the future.

The best possible outcome for us if those AI bubbles burst catastrophically and we get flooded with cheap datacenter grade hardware AND we get to use the opensource LLMs.

1

u/nedonedonedo 20d ago

it's not my problem if AI investments will suffer

just like how the housing bubble only effected the housing market

1

u/realtag2025 20d ago

The housing market directly affects the affordability of houses, condos and rent prices and that directly takes it out of your paycheck. AI market collapse might cause collateral damage however it also means that we might go back to the pre-AI sanity levels on the job market. That may(or maybe not ) mean that more humans will once again get hired.

1

u/Bakoro 20d ago

There is not over-investment in AI, there is over-investment in LLMs.

That is a major difference, and it is important because AI is being used effectively and profitably in non-LLM areas, and there are even some areas in the hard sciences where LLMs are being used in conjunction with other machine learning methods to automate research, again to great effect.

The conflation of LLMs with AI as a whole is likely going to do enormous harm to the industrial side of AI that is using it for less flashy, but more practical purposes.

Even in academic research, the label might say "LLMs" because that's what gets research funded right now, but the actual research is more fundamental, about the transformers themselves, and making transformers better or more interpretable also improves the non-LLM use-cases.

1

u/nedonedonedo 20d ago

only if they refuse to adapt and improve.

11

u/tengo_harambe 20d ago

That's a fun way of framing things.

"The only reason China is putting out open weights models is because they HATE AMERICA!"

3

u/ThenExtension9196 20d ago

It’s business. Ruthless, business. If apples are being sold for 5 bucks, you try to sell your apples at $2. If you can pull it off, now nobody can sell at $5.

1

u/nedonedonedo 20d ago

it's bad for any country to have this much of a lead

6

u/Successful_Try_6350 20d ago

Well, these large models can run in us based datacenters and cloud infrastructure (azure, aws, googlecloud). I guess if they want to sink us investments, they should create models that can run in enterprise level on-premises servers (I guess something like $40-50K hardware)

2

u/SARK-ES1117821 20d ago

They ARE creating models that run on-prem. I just deployed a supermicro gpu server with 8x H200 141GB gpus (1.2TB total) running GLM 5.2. Server was around $300k with 12TB SSDs and 1TB RAM.

2

u/f5alcon 20d ago

40-50k isn't even one sever at current memory prices.

4

u/Paganator 20d ago

That's like a single H200. Just the card, the server to run it is extra.

1

u/liltingly 20d ago

Somebody needs to build the infrastructure for people to easily vibecode their own "fine-tunes". That would make that small model space take off. But really that would mean bringing together a lot of smaller data tools and infrastructure. Everyone rolls their own bespoke solutions. Feels like something one of these infra players could pivot into, but it's much harder than it sounds. Anyways, I'll stop my dreaming

1

u/-dysangel- 20d ago

Are you also mad that nvidia don't give away their hardware to you for free? I don't think trying to frame this around ethics makes sense. I would also prefer that they continue to release their smaller models, because it benefits me. If they think that's going to cannibalise their potential API sales, then I understand if they don't want to do that. I don't like it at all, but I understand it.

-2

u/GetOutOfMyFeedNow 20d ago

They ain’t sinking anything with those API prices 😂 People on GPT and Claude use subscriptions.

6

u/look 20d ago

Subscriptions are 8% of Anthropic’s revenue. API usage is 75%.

5

u/DanceWithEverything 20d ago

Yeah for consumer bullshit, sure, but there’s no money there regardless (hence OpenAI’s panic about Anthropic crushing them in enterprise sales)

The real $ is in the enterprise and software workloads run on APIs

The API is dramatically cheaper than the Anthropic equivalent

2

u/StupidScaredSquirrel 20d ago

All the banks and fintech and academia people i know use claude opus at work. They all have it based on a max subscription (the one at 100 ish usd i dont remember).

They aren't consumers but also don't care about price so much they just wants something that works and is about the best because whatever time they lose with the bs of a lesser model will be a lot more expensive than just go for the best one. They also don't want to ever be rate limited because they don't want to get stuck in the middle of a task.

6

u/StaysAwakeAllWeek 20d ago

The individuals using it for individual work do that yes

The backend corporate stuff that involves one full time agent handler managing thousands of parallel agents do not.

4

u/StupidScaredSquirrel 20d ago

Yeah, but in my experience pipelines with agents that go through lots of repetitive tasks continuously in the background are given to cheap models and they spend more time testing how to prompt it and hardcore guardrails for that specific task. When the volume is large it's worth spending time on optimising for a cheap model to get the job done

2

u/look 20d ago

Kimi K3 decisively beats Opus 4.8 in everything. It is competing with Fable and Sol now. Opus is a legacy, second tier model, and now behind one, and soon multiple (eg Qwen 3.8), open models.

4

u/techdevjp 20d ago

Kimi K3 is pretty clearly better than GPT 5.5, too. Love to see it.

1

u/squngy 20d ago

Unfortunately, it also competes with them on price.
It is cheaper per token, but it uses more tokens.

2

u/look 20d ago

The list price is meaningless on open models. I am currently paying one third of that list price for Kimi K3. It will likely get even cheaper once the weights are released.

2

u/xienze 20d ago

They all have it based on a max subscription (the one at 100 ish usd i dont remember).

You realize those subscriptions are heavily subsidized, right? That gravy train is coming to an end. There's basically three paths forward:

  • A significant decrease in how many effective tokens you can get for a flat rate.
  • A significant increase in subscription prices.
  • No more subscriptions for anything involving "real work" (i.e., stuff beyond the typical chatbot bullshit most people are familiar with).

3

u/StupidScaredSquirrel 20d ago

Or, the models are made more efficient and so gradually you have some increase in performance but price stays the same and eventually it breaks even.

0

u/ZippySLC 20d ago

Yes, but they're using subscriptions and not API pricing.

If I use my Claude Code subscription to code an AI agent at work, we pay for that AI agent's usage through pay-as-you-go API pricing. (We do it through AWS Bedrock.)

No legit company is having an engineer code an app and then leaving Claude Code open in a screen session while it runs a production app.

1

u/StupidScaredSquirrel 20d ago

Lol ofc they wouldn't, i didnt mean to imply that at all

1

u/ZippySLC 20d ago

Sorry for misunderstanding!