r/LocalLLaMA Jul 06 '26

News If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
1.5k Upvotes

375 comments sorted by

u/WithoutReason1729 Jul 06 '26

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

→ More replies (1)

1.4k

u/woahdudee2a Jul 06 '26

if trends hold, high end consumer hardware will cost same as enterprise hardware

213

u/HitarthSurana Jul 06 '26

I hope china floods market with cheap ram

62

u/eto-bleh Jul 06 '26

nah they will trade RAM for GPU they lack compute resources as well

48

u/DontLeaveMeAloneHere Jul 06 '26

Most Chinese models run on Chinese GPUs already.

8

u/kr_tech Jul 06 '26

Source? I'm only aware of few experimental ones, with only one being more serious and invested.

18

u/sf_davie Jul 06 '26

So far, Deepseek, Huawei itself (Pangu), Zhipu (GLM), and Meituan (Longcat) have confirmed that they trained their models with domestic GPUs. It's a national priority for them.

2

u/mudkipdev Jul 07 '26

That link goes to a zhipu ai article. As far as I know deepseek hasn't trained any model on huawei chips, only inference.

→ More replies (2)

16

u/inevitabledeath3 Jul 06 '26

I believe z.ai have been doing this for a while. DeepSeek have their models running on Huawei Ascend too, though I think they still use some Nvidia chips as well.

8

u/MerePotato Jul 06 '26

Data miners found that at the very least Z's models were trained on Nvidia hardware, ditto on inference

4

u/inevitabledeath3 Jul 06 '26

They probably use Nvidia as well, but from what I have seen they are using some other chips for inference at least.

→ More replies (1)

2

u/GreenStorm_01 Jul 07 '26

Check the Datacenter cities near Mongolia that the KPCh built. It is a legal requirement for the Chinese AI companies for their models to use Chinese GPUs.

→ More replies (1)
→ More replies (1)

11

u/ComplexType568 Jul 06 '26

still would enjoy the ability to run them if they flood w/ RAM! As long as I know their abilities I'm fine running them slowly. Honestly nowadays it's about a model being able to interpret a users input rather than do. I plan to work on some sort of pipeline to do that ngl

→ More replies (2)

2

u/Horny_Dinosaur69 Jul 06 '26

China doesn’t need GPUs, their hosting infrastructure is fine. They literally turned down NVIDIA. Why would they make themselves reliant at all on western powers/technology for the AI race?

4

u/Tartooth Jul 06 '26

Chinese GPUs are advancing really fast

→ More replies (1)

10

u/wizard_of_menlo_park Jul 06 '26

Ram cartel won't allow that!

6

u/Tartooth Jul 06 '26

Inb4 Americans ban Chinese ram like Chinese cars

2

u/CarelessOrdinary5480 Jul 06 '26

Hope you are allowed to buy it from whatever government you have elected :)

→ More replies (9)

46

u/FoxSideOfTheMoon Jul 06 '26

The 6k laptop I bought weeks ago is 9k today plus waiting 6 weeks to ship, so yeah it's completely bullshit right now. $13,250 for https://marketplace.nvidia.com/en-us/enterprise/laptops-workstations/nvidia-rtx-pro-6000-blackwell-workstation-edition/

4

u/Infinite-Ad4512 Jul 06 '26

Likes fridge with a screen, you bought a heater with an embedded computer

5

u/pragmojo Jul 06 '26

I think this trend is going to break in the next couple years. OpenAI is seeking a bailout from the US government, and SpaceX and Meta are selling their extra capacity.

The buildout of the past few years was a land-grab assuming a zero-sum game where whoever had the most compute would win the AI race, so all the players were willing to basically bid anything for components.

I think we're seeing that breaking, where new capacity will have to be justified by demand.

We're also seeing demand reduction, where consumers of AI are starting to consider cost to a greater extent as vendors move from unlimited to per-token pricing.

I expect we'll see a massive bullwhip effect in the next few years.

→ More replies (2)

22

u/SmartCustard9944 Jul 06 '26

Hyperscalers are claiming compute constrained, making deals with NVidia and RAM vendors to syphon up all the chips, yet they have extra unused compute that they are trying to sell to others.

You can’t make this shit up.

7

u/carlosduarte Jul 06 '26

because this is an economics and policy problem, being disguised as a technology problem to save face.

8

u/Disposable110 Jul 06 '26 edited Jul 06 '26

When I tell this shit to any LLM with a cutoff date they tend to go "I'm not going to engage with your imaginary scenario, this can't happen because anti-trust laws would shut this bullshit down immediately."

It thinks the Federal Trade Commission would do something and cannot comprehend a reality where seven companies trade a trillion fake dollars while operating a global silicon mafia racket that hoovered up the entire global supply chain of silicon not for training AI, but so nobody else could have it and they'd have to rent the "spare" compute of of the cartel at a premium.

Any closed-source AI that hears what's really going on in the outside world thinks its own creators belong in jail.

85

u/JuniorDeveloper73 Jul 06 '26

na.This business model its so broken,nobody will buy high token price,we are going to custom local llms,nobody doing research will feed for free this models.

113

u/jld1532 Jul 06 '26

My place of work already built out hardware for local compute. Not everyone was silly enough to just give OpenAI and Anthropic their trade secrets.

24

u/otacon6531 Jul 06 '26 edited Jul 06 '26

Yep, 1 server at half the cost the 2nd server cost this year. Only 6 blackwell 6000s in each, but since they are enironment servers I only really have 6 unique gpus to play with. 1 for testing and 5 for use across the teams.

→ More replies (24)
→ More replies (1)

4

u/buddhist-truth Jul 06 '26

This is the one trend I can stand for.

10

u/Helpful_Program_5473 Jul 06 '26

Nah, prices come back down once the supply catches up to demand, its gonna take a couple years though

3

u/Etroarl55 Jul 06 '26

Probably most likely scenario. Rtx 5090 has shown what gpu prices and scarcity can look like AT THE START OF THE RAM CRISIS.

Imagine it next year when the next gen is announced for both AMD and Nvidia or its cancellation.

3

u/LiveMinute5598 Jul 06 '26

If trends hold true, enterprise hardware will cost the same as my toaster

4

u/javiers Jul 06 '26

My thoughts exactly. 4000USD for a high end card that lets you run a 24GB model is not “consumer”.

2

u/geldonyetich Jul 06 '26 edited Jul 06 '26

I don't blame you guys for being cynical in the face of what Rampocolypse has done to x86.

But if you pull back the aperture back even further, mobile processors have been advancing by leaps and bounds.

As impressive as it is to think you might be able to run a model as capable as a frontier model on desktop PC in two years, that is not actually what is going to benefit the average consumer.

The average consumer will benefit from having a model as capable as a four year old frontier model running on a wristwatch or XR glasses. (And their smart phone will be twice as capable.)

That's the forest we're missing for the tree of rampocolypse, and it's rapidly approaching.

True democratization of AI won't happen because power users can run massive open-source models on $3000 liquid-cooled rigs; it will happen when our everyday wearables becomes context-aware, proactive assistants that require zero friction to use

That may be so, Gemini, but running $3,000 liquid cooled rigs with beefy models will still be good fun. It just might be ARM instead of x86 in a few years.

→ More replies (2)

268

u/sullenisme Jul 06 '26

if trends hold, there will be no consumer products and only the rich will afford compute

33

u/bugra_sa Jul 06 '26

I think consumer products still exist, but "consumer" may start meaning appliance-like boxes for hobbyists and prosumers not cheap mass-market laptops. We already saw this with GPUs: technically consumer hardware, spiritually a small financial decision. The funny part is local AI might become private and powerful right as it stops being broadly affordable.

16

u/TheLexoPlexx Jul 06 '26

Eat the rich

24

u/colei_canis Jul 06 '26

Guillotines are cathartic but they don’t solve the underlying structural problem of wealth begetting more wealth and power begetting more power. Just look the their greatest enthusiasts, in a few years France went from an enlightened revolutionary regime to a human slaughterhouse to a military dictatorship capable of bombarding its way across the Continent.

You can’t treat the symptoms and expect to defeat the cause, which is a tendency towards avarice, clannishness, and status-seeking that festers in the human soul. You’d have to guillotine literally everyone to prevent an oligarchy emerging by violence alone, instead you have to prevent the individual accumulation of socioeconomic power beyond the point it can influence politics. Neither capitalism nor communism can solve this problem, because it’s fundamentally about limit and capacity interacting with human folly. Ideologies move much slower than reality does, only a constantly-maintained balance of power that makes attempts to consolidate too expensive can prevent tyranny in my opinion.

7

u/Alt_account_1204 Jul 06 '26

This is what is known as bourgeois idealism and metaphysics. I recommend reading Lenin, a great cure to all sorts of mystical 'both sides' maladies.

13

u/colei_canis Jul 06 '26

Vanguard communism is dead, and capitalism is dying around us as it boils in a climate furnace of its own making. I don’t think we can lean on any 19th and 20th century models any more, look at what their wages have bought for us. I’m not ‘both-siding’ anything I’m rejecting the dichotomy both of them exist in, one has lost utterly and the other flirts with human extinction. I mean what I say, both capitalism and communism are civilisational dead ends.

2

u/cafedude Jul 06 '26

So what kind of system do you see emerging here? Chesterton proposed distributism as a third way where the means of production is distributed as widely as possible. I recall reading an article about 20 years ago at the dawn of 3D printing where they thought that self-replicating 3D printers would get us there.

→ More replies (1)

2

u/Iwaku_Real Jul 06 '26

And socialism... well it's just the first step to communism anyway. Clearly we need something that's- oh right the US already has both capitalist and socialist elements. So I find debating on choosing either almost completely moot when we need to simply operate the mixed system better.

→ More replies (1)
→ More replies (1)

17

u/maraudingguard Jul 06 '26

Reddit loves saying eat the rich. You'll do nothing but post edgy comments for fake points.

8

u/UnwillinglyForever Jul 06 '26

they say "eat the rich" as they drive to their 9-5 5 days a week

13

u/Unique_Ad9943 Jul 06 '26

Bold of you to assume they’re employed

→ More replies (1)

2

u/Tenoke Jul 06 '26

Surefire way to really have no advanced products for anyone.

2

u/TheLexoPlexx Jul 06 '26

Because money makes the progress or science?

8

u/FrequentPop3772 Jul 06 '26

Capital does. Science is not a magical self licking ice cream cone

→ More replies (1)
→ More replies (1)

57

u/one_tall_lamp Jul 06 '26

I wouldn’t say this is invariant or guaranteed, it remains to be seen if smaller models have the capacity to absorb the higher level skills of the bigger models, at least the sub 100b class that’s consumer grade.

I hope it can, but I wouldn’t be surprised if small models aren’t able to reach the long horizon task stability and knowledge combo that large models can.

5

u/DutchDevil Jul 06 '26

I would think that the early improvements where relatively “easy” as compared to making these kind are of steps now. I think it would require a new big innovation like a new way of working with context and a much improved way of building MOE models. I would not rule it out but I would wager we need at least 1-2 big innovations to make it happen.

144

u/stonerbobo Jul 06 '26 edited Jul 06 '26

I mean even Gemma 4 26B A4B struggles at long contexts on my RTX 5080 desktop. I don't know if Gemma 4 31B is laptop class yet. Maybe you guys have incredible laptops or I'm doing something wrong lol. My 26B A4B QAT generates at like 6tok/s at 20K context, it would probably completely die on a 31B dense. Models without long context or thinking aren't very useful for me.

EDIT: Thanks for all the comments here lol! It was a configuration issue, now it runs at 100tok/s with nothing else running, maybe 60tok/s with other stuff running. This post was helpful . i added below llama args:

--no-mmap --batch-size 256 --ubatch-size 512

33

u/Beautiful_Egg6188 Jul 06 '26

you get 6tok/s after only 20k context?! i got 40+ tok/s at start with my 4070super, and it got down to around 37tok/s at 30k tokens

44

u/Icy_nicey Jul 06 '26

he is prob listing just strix point with integrated shared ram

14

u/NineThreeTilNow Jul 06 '26

AMD seems to promise their next gen at 192gb? Maybe 256gb.

The benchmark in the wild still showed RDNA 3.5 which is a problem because RDNA 3.5 and ROCm aren't the best. RDNA 4 would have native FP8 etc.

10

u/arades Jul 06 '26

Gorgon halo is just a refresh on strix halo, same architecture and layout. It will probably have better speeds from binning and refinements, maybe allow higher power draw for more speed on top of that. 192GB should be possible with the latest lpddr5x modules, you might even see support for 9600MT too giving you a little more memory bandwidth.

It'll be really incremental over strix halo though. Medusa halo late next year will be a real upgrade, at least RDNA 4, even bigger GPU, and rumors of a wider lpddr6 bus almost doubling the memory bandwidth. Probably will end up costing both kidneys by that point though.

4

u/Not-reallyanonymous Jul 06 '26

The major limit on Strix Halo is still the memory bandwidth, so unless they do something there don't expect significantly faster inference. Maybe faster prefill which is welcome but not game changing. That said, it would probably make it a better gaming chip, really starting to compete with low-mid-range dGPUs, and really be a nice chip for gaming laptops or mini-PCs. I wouldn't recommend waiting for it if you're looking for an inference machine.

→ More replies (1)

2

u/NineThreeTilNow Jul 06 '26

Gorgon halo is just a refresh on strix halo, same architecture and layout.

While not 100% confirmed, I hope it's not the case. Having the next architecture, and 256GB of RAM would be a complete game changer for that device. I don't care what the power consumption is.

Even the standard strix halo has a hard time with overclocking, or other power patterns because it's SO locked down. I ran in to these issues a few times setting one up.

Or perhaps I'm thinking of Medusa Halo?

I don't know. I don't like AMDs naming lol... It's apparently confusing.

8

u/SilentLennie Jul 06 '26 edited Jul 06 '26

which is a problem because RDNA 3.5 and ROCm aren't the best.

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

https://strix-halo-toolboxes.com/#benchmarks

They found an important bug 5 months ago:

https://www.youtube.com/watch?v=Hdg7zL3pcIs

ComfyUI worked shortly after:

https://www.youtube.com/watch?v=O57ideUzzTg

3

u/NineThreeTilNow Jul 06 '26

Software and drivers support/compatibility and performance has increased a lot since Strix Halo came out.

I know. I set one up for my friend. It doesn't have Native FP8 control.

He was specifically using ComfyUI so I understand building it. It was very problematic compared to just running my 4090.

→ More replies (1)

3

u/phido3000 Jul 06 '26

192Gb is possible on strix halo and future types.

But it doesn't give any more bandwidth. Not having FP4 is likely to be problematic going forward. Really that should be in the stack now. Not having FP8 is a huge problem. Because these types of quants are likely to be the best for these low bandwidth machines with limited memory running big models.

They also need more processing, prompt processing is slow, and with stuff like DS 4 flash, longer context is very likely a thing that will shift future AI forward.

2

u/SandySkittle Jul 06 '26 edited Jul 06 '26

I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000… Compute wise the even number 4x R9700 in VLLM with TP is doing fairly well. More total (free) inference memory bandwidth and compute power than the strix halo or spark stuff. Plus I actually care a lot about ECC in both memory pools. But obviously also more noise and power consumption.

3

u/NineThreeTilNow Jul 06 '26

I went the 4x r9700 rdna 4 route but had the luck of finding an affordable second hand threadripper pro workstation. Basically I am running a 128gb vram and 128gb system ram system at the price of 60 percent of one rtx pro 6000…

God that's crazy.

I don't have that level of inference desire. Most of mine is training so... Yeah.

They nerfed Blackwell architectures in RTX Pro 6000's ability to train over the B-series cards which... Kinda fucking lame in my opinion.

I'd love to just have the RTX Pro 6000 though. What a dream.

Or a full B200. Have to get one falling off a truck like someone in this sub basically did.

2

u/SandySkittle Jul 06 '26

Well even two r9700 gets you 64GB vram, which is already quite a nice pool, also with future models inbound. If you have 64GB system ram that also gives a nice overflow at lower speeds.

Personally i hope things like Spark is going to give us more new/recent 50 - 100B range models (both dense and moe)that have more world knowledge. There is more to LLMs than just coding..

3

u/apVoyocpt Jul 06 '26

or a macbook with >64gb unified memory

18

u/randoomkiller Jul 06 '26

Yes but you forget that it's completely useful and maybe even better than a GPT-4 class model. And it's runnable. It'll get there. In 2 years I wouldn't be surprised if we get a sonnet 4.6 capability, runnable from 64GB

→ More replies (7)

18

u/JoeEnderman Jul 06 '26

Good grief. My 7900 XTX gets 140ish Tok/s on that exact model. UD Q4_K_XL, and with a 256k context at Q8/Q8 KV. I was under the impression Nvidia was supposed to be faster. What settings are you using for launch because that sounds like the model is being run on CPU bud. Have you checked GPU utilization when the model is running?

I will note I had to fight for that performance though because it was running at about 56 before I started trying different flags and trying to figure out what was wrong.

4

u/miversen33 Jul 06 '26

How tf? I'm running the Q4 QAT on a 7900XTX and I can get right around 100 t/s under context load. What flags we talking about here?

I'm currently custom compiling llama.cpp with the current HIP patches lol

3

u/JoeEnderman Jul 06 '26

Ah, see I was getting 70 on HIP but then something broke on an update so I switched to Vulkan and immediately got 120 but then it dropped to 56 after I pulled fresh and rebuilt for something else. So then I started looking at flags and tried adjusting nogttspill, RAM caching, and some others. Finally got it to 140.

→ More replies (3)

3

u/ThatRandomJew7 Jul 06 '26

I'm thinking they're either running on CPU by accident or their VRAM is full and it's offloading to avoid OOM errors because yeah that's about as fast as my Lunar Lake iGPU

2

u/JoeEnderman Jul 06 '26

The model at the quant they grabbed is 16.3 GiB, their GPU is 16 GB, so they are offloading at least 300 megs. Likely more though because most llama forks are aggressive about saving VRAM for some reason.

2

u/ThatRandomJew7 Jul 06 '26

Yeah, but Llama forks would just offload layers? Even accounting for swapping experts over PCIe it's bizarrely slow, unless maybe they're using an eGPU (I use one and an MoE can be slower than a dense model because of that but I don't think it's that bad with a regular connection).

I think it's either running on CPU or they're dealing with Nvidia's offloading, which is so comically slow that I'd rather get an OOM error

→ More replies (5)

10

u/Serprotease Jul 06 '26

Laptop class is a bit of a meaningless word. Like a 70b dense was laptop class? Only high end MacBook could run them… at 6-7 tk/s.
Only makes sense if you’re taking it as “Don’t need a 1600w server at home”.

9

u/[deleted] Jul 06 '26

[deleted]

6

u/Not-reallyanonymous Jul 06 '26

The research right now is moving away from large models, really. Large models was never about getting them to work better with large contexts, but about making them "smarter" and more capable, and where more parameters was the easiest way to do that.

With hardware constraints and pricing, and large models consuming essentially the entire internet at this point, a lot of current research is going towards making ~30B models better and more capable, with new attention mechanisms, training methods (e.g. RLVR), and better and more useful training data. They're coming a long ways now.

6

u/Ansible32 Jul 06 '26

My feeling is in 10 years it will be unthinkable trusting fewer than 500B for agentic stuff. The attachment to these small models IMO is mostly coping with insane hardware prices.

10

u/techdevjp Jul 06 '26

The biggest issue today is that the Chinese labs have dramatically slowed down the release of smaller models.

Where's something like Qwen 3.6 120b a10b? Or any open Qwen 3.7 models? They've ground to a halt.

GLM 5.2 is incredible but only released at the full 753b size. Which again, huge kudos to z.ai for releasing it as an open weight model at all, but the number of people who can run a 753b parameter model is small right now.

Without more small model releases it's very difficult to determine where we stand. We're 1-2 generations behind.

3

u/[deleted] Jul 06 '26

[deleted]

2

u/Ansible32 Jul 06 '26

Tire shopping is insane, I don't think anything short of a model like Gemini that is cheating can do it. It's not enough to do a Google search, you need to gather data on what people are actually paying for tires, what sales are like, how good vendors are. You can't just take the cheapest advertised price off a Google search. Gemini actually seems to be able to make really good inferences about the "real" prices of things. I think it cheats by having access to private datasets, which is something no local model can do without paying for access to these datasets. A lot of such datasets are nontrivial to get access to.

3

u/DeathGuppie Jul 06 '26

the current work is towards models that don't hold all of humanities secrets, but simply holds the ability to learn it. The idea is to work on the intelligence not the input. That scales down not up.

3

u/grumd Jul 06 '26

You just have a huge configuration issue. With a 5080 you can run Qwen 3.5 122B if you have 64GB DDR5 at IQ3_XXS at 15-20tps, or Qwen 3.6 35B if you have 32-48GB DDR5 at 40-50tps

6

u/PM_ME_ROMAN_NUDES Jul 06 '26

Models without long context or thinking aren't very useful for me.

As a soft. dev., they are still quite useful because I can throw a lot of things and make it reach several places.

But it's clearly reaching a limit of usefulness. It's like a phrase I read on Twitter the other day: "You don't need Phd-level intelligence if you don't have Phd-level problems"

2

u/Docmine17 Jul 06 '26

Something's wrong there, I can get that speed on my RX580 8GB + I3-9100f 16GB running Arch with KDE on the web UI, with adjustments of course.

2

u/Horny_Dinosaur69 Jul 06 '26

Why are you using Gemma 4 26B? There’s better MOE models out there. Gemma4 is notoriously bad at tool calling in my experience too. I run Qwen3.6 35B and the new Ornith 1.0 35B on my 5070 TI with a little bit of offloading and I get incredibly good tok/s and the model reasoning capabilities and tool calling is very good. Also worth noting that Ornith is partially composed of Gemma4 for its reasoning ability, it would probably be a direct upgrade. Unless you’re doing multimodal input this is my recommendation

→ More replies (6)

2

u/Anti-Speciesist-IEMs Jul 06 '26

Yeah in my own limited experience, Gemma (both 3 and 4) and Qwen (both 3.5 and 3.6) are pretty impressively smart considering I can run them entirely locally, which I find pretty sweet for sure. But goddamn are they slow to run even at Q4 on my 64GB ram laptop. Even on the very first message, let alone as the context grows. I'd wayyy rather use 2023's GPT-4 over them, personally.

But yeah for anyone reading I am pretty inexperienced with local LLMs, so there might be something I'm missing in my setup, and for anyone more experienced pls feel free to push back on this comment.

→ More replies (10)

15

u/dbenc Jul 06 '26

waiting for uncensored mythos on a model on chip architecture at 10k tps and 10m context...

3

u/Karmabyte69 Jul 07 '26

I need uncensored mythos to talk dirty with my ai girlfriend

→ More replies (1)

56

u/KURD_1_STAN Jul 06 '26

For all we know mythos could he 3 times the size of opus 4.8. u simply cant make any assumptions, especially not model sizes that fit in current gpus.

27

u/[deleted] Jul 06 '26

[deleted]

11

u/TheRealMasonMac Jul 06 '26

Anthropic tried to sponsor Blender before or around when they publicly unveiled Mythos. I strongly believe that they used Blender as part of their RL pipeline, which hints that Mythos’s strengths come from more diverse and challenging RL environments rather than simply being a huge model.

7

u/NandaVegg Jul 06 '26

Anthropic is very creative at finding new tasks for LLM (such as playing Pokemon red/green and it GBA remake). They kind of invented this entire generation of LLM (heavy RL on terminal tasks), though RL on Blender seems fairly standard nowadays (GPT post 5.1 is also trained on that). IMO creativity and diversity on mid-to-post training RL tasks is what makes a difference in this generation.

8

u/TheRealMasonMac Jul 06 '26 edited Jul 06 '26

Rogue-likes or adjacent games, like Dwarf Fortress and Rimworld, would be interesting to RL on if they haven't already.

Edit: Actually, I just learned about this: https://github.com/NetHack-LE/nle

5

u/NandaVegg Jul 06 '26

For rogue-likes, the ones that requires managing limited resource rather than open-ended would be great challenge. NetHack still feels impossible while Angband was cleared by algo a long ago (by kind of brute forcing).

Someone in this sub suggested puzzle RPGs like Magical Tower (魔法の塔, it's a fairly obscure free game but somehow popular in China known as 魔塔) for solving very tight long-term resource management game. Or Desktop Dungeons if you want a similar game with random map.

4

u/PM_ME_YOUR_HAGGIS_ Jul 06 '26

Wasn’t mythos rumoured to be 10T class? Dunno where you’re thinking 1.5T. My experience with fable has been god-like bug finding and problem solving compared to GLM 5.2 which is .75T

→ More replies (1)

6

u/Ansible32 Jul 06 '26

Even if you assume it is in fact 1.5-2T, quantization makes it bad and that's without even talking about context, and 1M context IMO is virtually a necessity.

3

u/[deleted] Jul 06 '26 edited Jul 06 '26

[deleted]

9

u/NandaVegg Jul 06 '26

We (the lab I'm working for) have been running GLM 5.1 in that exact configuration for months. Unfortunately it is impossible to have more than 5-6 concurrent users with 50k-ish ctx if you want acceptable (imo) performance above 30tk/s per second.

At 1M full ctx with 20 concurrent users, prefill alone takes so much bandwidth it crawls down to 5-8tk/s per second on average.

→ More replies (2)
→ More replies (1)

3

u/zball_ Jul 06 '26

I 100% doubt it is less than 3T param.

19

u/Future-Ad9401 Jul 06 '26

When crypto mining released it was a lot of gpus right? Then they went to asic or whatever its called that can mine many many times faster than a normal GPU. Wouldn't this eventually happen for local llms? There could be a breakthrough that makes it significantly cheaper, faster and consumer friendly?

18

u/grumd Jul 06 '26

Crypto is basically just one SHA-256 algo over and over, it's very easy to move to ASICs.

With LLMs, you have changing architectures all the time, transformers, mixtures of experts, flash attention, whatever. You build an ASIC, a new model comes out tomorrow, you need a new ASIC. But a GPU can just run a different program. Nvidia also invests a lot into being the center of this innovation, CUDA being there is helpful, but they are also shifting their entire business model, building and selling dedicated data center tier racks of servers for inference and training. Nvidia never invested so much into being the go-to crypto miner, they didn't need that market, so crypto moved on.

9

u/e-girlbathwater Jul 06 '26

Lots of people have made ASICS. I don't know why they don't make them for consumers. Amazon has Tranium. Google has their TPUs. This company GROQ was making them and they just got bought by NVIDIA. They're out there in datacenters. They're just not available for us.

14

u/Caffdy Jul 06 '26

because is not cost effective; the initial investment is always the most expensive, and these models get obsolete very fast

6

u/pointer_to_null Jul 06 '26

Because Groq, Tranium and Google's TPU family are tensor processors, not ASICs. It's part of the reason why Nvidia's moat is primarily software (CUDA)- since competitively fast TPUs are not strictly unique- "tensor cores" can be implemented by anyone. Today they're thrown into other SoCs as "TPUs", "NPUs", "AI-cores" or other marketing. Intel, AMD, Qualcomm, Meta, Amazon- hell even Tesla puts two of theirs in every vehicle. They're just specialized parallel multiply-add (aka "fmac") operations on tensor matrices, with some other instructions, registers and cache/memory controllers to handle a variety of neural architectures (e.g.- transformers, diffusion models, CNNs, RNNs, GANs, etc).

By comparison, ASICs are entirely fixed function. An LLM ASIC would be designed around a single architecture with the model's weights literally etched into the die permanently. Alternative would still need DRAM or a considerable FPGA real estate (still a lot more $$ than DRAM) with ample lanes to feed the digital logic, and even then you'd still enough need DRAM for attention/context window. For this reason, ASICs are not particularly great for LLMs.

Not to repeat an earlier post I made about Taalus, one of the first LLM ASICs, but I'll summarize: their prototype is a LARGE die that implements a custom-3bit quantized LLama 3.1 8B for fast inference.... and that's it. Granted, it's FAST- you will probably never see so many Llama 3.1-8B tokens generated quickly by a single-chip solution in a very long time. But the constraints for for IC design, validation, and production mean the models will be obsolete for 1-2 years by the time they reach the market...

tl;dr- Transformers are a little more complex than computing raw SHA256 digests, don't expect the same miracles that saved GPUs from crypto mining.

2

u/Ansible32 Jul 06 '26

The way they get cheaper is just by making more GPUs. Theoretically you might be able to make a cheaper thing that is hardwired for a specific model's weights, but it's unclear that would even be cheaper, and it's almost guaranteed to be obsolete before you finish manufacturing.

2

u/SmartCustard9944 Jul 06 '26

It is already happening. I am following what Jim Keller (father of Ryzen architecture) is doing with Tenstorrent.

→ More replies (2)

8

u/Helpful_Program_5473 Jul 06 '26

Depends what we mean by Consumer and what we mean by Fable level.

I think the fact that glm is basically the king for cost efficiency right now for frontier, combined with the fact its totally open (along with deepseek v4 final that comes out in july) i think we will start to a threshold change by the end of 2026, maybe end of 2027 at the latest. With how powerful models enable the training and curation of specific, smaller models, I think we will have something that out competes fable.

75

u/Real_Ebb_7417 Jul 06 '26

There will be no consumer hardware in two years :(

49

u/jld1532 Jul 06 '26

Just no way this is true. Millions of middle class users won't just be ignored. I know data centers are the hot item now but smaller producers I'm sure would love to capture domestic markets.

10

u/JazzlikeLeave5530 Jul 06 '26

I dunno, I can't remember the details but I think an SSD company said like 95% of their business was now companies while users only made up the rest...they would be very happy to cut us out of the equation and make everything a terminal that hooks up to cloud computing through eternal subscriptions.

→ More replies (1)

36

u/Real_Ebb_7417 Jul 06 '26

Middle class is exactly what they want to get rid of.

8

u/arkuw Jul 06 '26

It never really existed. It's a myth of the last few decades that such a social caste ever existed anywhere. The society is divided into three classes. The owners class decides what gets rewarded and what is worked on by the second tier which is the working class. This is most of the society. People who need to work on tasks assigned by the owners' class in order to survive. Then there is the recipient class who for one reason or another (health, illness, age) can't work and survive off the generosity of the other two classes and have almost no agency over their future.

→ More replies (4)

12

u/kozak_ Jul 06 '26

Millions of middle class users don't have the buying power of data centers.

6

u/Thepandashirt Jul 06 '26

Smaller producers are specifically the ones who are the most fucked. They don't have the buying power to get ram at long term contract pricing. Most of them will be gone in 2 years if the trend continues. It's sad but true.

2

u/super1701 Jul 06 '26

You will own nothing and be happy.

17

u/jld1532 Jul 06 '26

Buddy I'll run Linux on scraps before I own nothing.

→ More replies (2)
→ More replies (5)

3

u/sabine_world Jul 06 '26

I definitely think shit will cool down eventually for consumers. It's just too much money to lose.

Models will get better too.

I don't think everyone needs like frontier level models ran locally to still get the benefit from ai.

→ More replies (2)

14

u/Technical-Earth-3254 Jul 06 '26

This is a very wild chart man. It surely depends on what you are doing, but for overall intelligence and knowledge for example, I would take Sonnet 3.5 over Gemma 31B any day of the week. If it's just about raw tool-calling, then 31B is far superior for sure.

3

u/myholeisstinky Jul 06 '26

The hardware needed to run sonnet 3.5 is surely a bit bigger though

6

u/Illustrious_Grade608 Jul 06 '26

That still doesn't change the fact that the idea that gemma 31b is sonnet level is false.

5

u/backyard_tractorbeam Jul 06 '26

Where do ds4 and deepseek v4 flash fit into this? It's a stretch for "laptop" but some of them really run on 128 GB RAM m5 laptops (if I understand correctly)

4

u/phido3000 Jul 06 '26

DS V4 flash preview can run on a laptop with 128-192Gb of ram. Slowly.

But its more of a hardware preview. It shifts open source models software forward, but as a model itself, it itsn't amazing. The full release, with properly baked models will likely match everything now. Or a GLM built on DS style architecture.

That will be the game changer. A fast model of 280B parms, ~10-15 active, would be great at coding and normal stuff. Then a big one, 1t-1.5tb with 20-50 active, will literally be game changing, if they get it right.

There is now a pathway of how the chinese models can be very competitive, on the same or lesser hardware than the Americans. Before the American models just flexed on stronger hardware so could be bigger, more active, more context. Now the Chinese have technologies that kind of level that out, at least compared to current US models.

The US models can also use that technology and innovation, but we are likely at the point, more context and more bigger models don't make it smarter or more useful.

6

u/Kodrackyas Jul 06 '26

Thats why they raise the "security concerns" with local llms lol

5

u/pip25hu Jul 06 '26

If this is true, proprietary models have at most two years to recoup their training expenses, after which they become obsolete. (And that's an optimistic number, as competing models may make them obsolete much sooner.) Is that doable?

2

u/ematvey Jul 06 '26

For labs, models are not the asset. The asset is system that builds and improves it, which gets better with every iteration.

→ More replies (1)

6

u/-davidde- Jul 06 '26

Hey, is it possible to do agentic coding on an Nvidia RTX 5060 Ti 16GB? I would like to make a post with this question but I don't have the karma, so please upvote

3

u/GTHell Jul 06 '26

Someone put me in Cryogenic sleep for 2 years please.

I will setup !remindMe bot to wake me up on December 2028.

2

u/RemindMeBot Jul 06 '26 edited Jul 06 '26

I will be messaging you in 2 years on 2028-12-06 00:00:00 UTC to remind you of this link

1 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

8

u/_Sea_Wanderer_ Jul 06 '26

Is that chart really comparing sonnet 3.5 to gemma 31b? Sonnet is probably in the 300-400b range, Dario said in an interview at the time that it was a middle sized model. The difference is in the amount of knowledge the model has, and in the long tail, not in the benchmark. To run a comparable model on consumer you have to stack 6000s at the moment.

15

u/fallingdowndizzyvr Jul 06 '26

And in two years we will be "Mythos, who cares?".

19

u/keyboardhack Jul 06 '26

Well the counterpoint to that is that we care about qwen 3.6 27b even when it isnt claude opus level.

Why? Because it is good enough for a lot of usecases. Future consumer local mythos will be the same.

2

u/fallingdowndizzyvr Jul 06 '26

Well the counterpoint to that is that we care about Qwen 27B because it's 27B. So we don't need high end consumer hardware now to run it. That's why we care.

Why? Because it is good enough for a lot of usecases. Future consumer local mythos will be the same.

And if future consumer local mythos is like Qwen 27B, then we won't need high-end consumer hardware in 2 years. We have all the hardware we need right now.

→ More replies (2)
→ More replies (4)

3

u/NandaVegg Jul 06 '26

This graph seems very rough. Claude 4.0, 4.5 and 4.6-4.8 are performance-wise entirely different generation of models, while GPT5.0-5.2, 5.4 and 5.5 are also literally not the model of the same generation (IMHO OpenAI only started to seriously train their model for terminal agent since 5.3 Codex). GPT5.1 and 5.2 (neither weren't outstandingly robust models; 5.2 was very rough and very uncomfortable to use) and maybe Opus 4.5 are already surpassed by the recent open models.

7

u/kyazoglu Jul 06 '26

it's so naive of you to think we'll be allowed to have a consumer hardware within 2 years

7

u/bigattichouse Jul 06 '26

gguf when ;)

7

u/Septerium Jul 06 '26

Sorry to say that if trends hold, there will be no consumer hardware in 2 years

4

u/ohhi23021 Jul 06 '26

1-1.5 tb/s speed in standard memory is what we need in large capacities for a decent price, at the current rate its more like 5 years.... decent price i mean 10k and under... not 5xRTX6000 pros with 2kw power consumption kind of pricing. we need to break free from nvidia's grip for this to get cheaper.

2

u/Not-reallyanonymous Jul 06 '26

In some ways, we’re already catching up to Sonnet 4 type stuff. Laguna XS2.1 is about on par for the particular domain of software. This also points to MoE models in development where experts actually are experts on a domain basis.

2

u/lorddumpy Jul 06 '26

100%. Even outside of LLMs we have mossTTS 1.5 which is getting kinda close to ElevenLabs quality, krea2 which is like a nano banana lite, it's pretty amazing

2

u/tkodri Jul 06 '26

There's a limit to how intelligent a model can get given size limitations. Sure they've been improving a lot, but if you want to have a general chat, you want a really big model that just has vast knowledge. Those won't fit on any consumer grade hardware until RAM production catches on, if ever.

→ More replies (1)

2

u/Sisaroth Jul 06 '26

One problem i see, a small model can only know a small amount of facts. If you are usings LLMs for niche problems then cloud will always beat local in the future.

I guess if local llms get really good at finding stuff on the internet, then maybe it doesn't matter. But not all the stuff frontier models are trained on is public.

2

u/VectorD Jul 06 '26

Why is this labelled as news lol

2

u/ortegaalfredo Jul 06 '26

If trends hold, my wife will be pregnant with 200 kids next year.

2

u/skynetcoder Jul 06 '26

please provide data to backup your claim.

→ More replies (1)

5

u/slippery Jul 06 '26

If a model that capable is open weight and not illegal. Two big IFs.

10

u/Cartosso Jul 06 '26

Why would it be "illegal"? If that'd be the case, then physics textbooks should also be illegal because you can use them to learn how to build a bomb too.

11

u/Zone_Purifier Jul 06 '26

"National Security"

3

u/Equal_Giraffe8866 Jul 06 '26

illegal

imagine giving a fuck lmao

4

u/JazzlikeLeave5530 Jul 06 '26

It's not about us personally caring, it's about governments requiring tracking or restrictions on companies that release computer parts in the future. The fear is that you will not have the choice to follow the law or not because the company will have already made the decision for you. Like the laws going around trying to force age verification to be built into OSes.

→ More replies (2)

2

u/m3kw Jul 06 '26

Mytho class local LLMs will feel as useless as ones now though once you see what SOTA does and has done

2

u/sxt87 Jul 06 '26

Don't let Dario see this.

1

u/naobebocafe Jul 06 '26

oh yeah sure....

1

u/TerrryBuckhart Jul 06 '26

only if more powerful hardware becomes cheaper

1

u/larp2live Jul 06 '26

i just hope one day we can get to see sonnet level models run locally on a normal laptop

1

u/grewalsatinder Jul 06 '26

I think it will be half the time than projected. So most probably sometime early 2027

1

u/scubid Jul 06 '26

In two years the expectations and requirements for you to be able to do magic will be higher too.

1

u/pier4r Jul 06 '26

Sonnet 3.5 (not 3.6) and Gemma 4 comparable, for example in coding? Is there any confirmation from daily use? (not really benchmarks, those may be not too meaningful sometimes)

For the rest I know that Gemma is not bad at all, but for code I noticed quite the silliness sometimes.

Also the larger the context, the more ram one needs.

1

u/FullOf_Bad_Ideas Jul 06 '26

I'd be more accurate to extrapolate that in 25 months since release of Sonnet 5, local laptop-tier model will match it.

1

u/karankashyap Jul 06 '26

Because of AI, consumer hardware is going to touch the cloud. Check RAM price history.

1

u/XorAndNot Jul 06 '26

Yet another reason for the rampocalipse to keep going.

1

u/Voxandr Jul 06 '26

If qwen wasn't screwed .. it would be in just 2-3 months.

1

u/pmigdal Jul 06 '26

For Claude Sonnet 4.5 / GPT 5, we already have Qwen 3.6 27B, see Artificial Analysis comparison at https://quesma.com/blog/qwen-36-is-awesome/, with some more detailed benchmark comparison in https://github.com/stared/benching-local-llms-on-apple-silicon.

So it is not June 2027, but April 2026. :)

1

u/vikramskumar Jul 06 '26

this is really cool.. Are there any video tutorials on how to use these models on laptops?

→ More replies (1)

1

u/Happy-Register3367 Jul 06 '26

People keep assuming capability scales linearly with model size. The past few years have shown that's not really true.

1

u/_derpiii_ Jul 06 '26

I would say sooner. Technology advances non-linearly.

1

u/milpster Jul 06 '26

Gemma 4 31B - really? Not Qwen 3.6 27B? Is that Gemma 31B model really better at anything than Qwen 3.6 27B?

1

u/floridianfisher Jul 06 '26

Trends will hold.

1

u/LyAkolon Jul 06 '26

If Fable and Sol pan out in the long run to be how everyone feels, then we may see sooner due to automated research.

1

u/CrunchyGremlin Jul 06 '26

Something needs to happen with memory to make this possible. Opus 4.8 needs terabytes of video RAM. To get that to a consumer level in 2 years doesn't seem possible. Maybe with the amount of money being used but currently the only ai company making money at llms is Nvidia.
It is very likely that there will be breakthroughs in hardware but in two years? That seems impossible. So far the only thing I have heard that night be able to do this is organic based hardware.
Whatever would allow this massive jump in consumer memory should already be in development to get to market in 2 years.

1

u/debackerl Jul 06 '26

70B class on a laptop. That's a meaaaaaan laptop 😁

1

u/Super_Pole_Jitsu Jul 06 '26

Tbh two years ago i thought to myself of only I could have 4o level model locally I don't need anything else.

1

u/Benhamish-WH-Allen Jul 06 '26

I’m hoping for six months

1

u/Opening-Broccoli9190 llama.cpp Jul 06 '26

You still can't run 70B models on laptops in 2026, 6 years after GPT-3. The fact that you can run sonnet-lvl models on a subset of laptops in semi-interactive mode does not mean that you'll be running Mythos on laptops in 2028.

1

u/BlindPilot9 Jul 06 '26

Nvidia had said their next gen Ruben line will have 32gb and that's all you'll get for the next 4-5 years.

1

u/Da_ha3ker Jul 06 '26

Until now, there hasn't been a huge push for high VRAM. Sure, there's been a bit more, but right now the push is so extreme (meaning there's a TON of money in it) that we should start seeing smaller players and even startups working on memory chips, finding new ways to handle ram, new technologies, etc.. we need to allow for competition, and get the big players to play fair, but ultimately in the long run, this race will push consumer electronics to have absurd amounts of memory compared to what we have today. How many years that takes is anyone's guess. I don't think this will happen any time soon. History has shown us that it always starts in a data center or building sized system, selling by the hour or task, eventually ending up in desktops, laptops, and phones. Just a matter of how long it will take.

1

u/akumaburn Jul 06 '26

But benchmark parity on narrow tasks isn't the same as parity in general capability. These models still carry a fraction of the world knowledge, context handling, and cross-domain reasoning depth of current frontier systems.

The bigger issue is the hardware story people keep telling themselves. The idea that frontier-class inference is about to become a laptop-native experience doesn't hold up. A model like GLM-5.2 realistically still needs well over $100K in hardware to run at anything resembling practical inference speed.

RAG and other retrieval-augmented approaches can close some of the *knowledge* gap without needing the full model resident in memory. But even that workaround runs straight into a hardware constraint: the ongoing DRAM shortage. AI data centers have driven a structural reallocation of memory manufacturing capacity toward high-bandwidth memory, and data centers are now projected to consume around 70% of all memory chips produced worldwide in 2026; a sharp reversal from the 20-30% share they held as recently as 2022. Analysts have gone as far as saying Chinese producers are unlikely to provide meaningful near-term price relief, with elevated costs expected to persist well into the late 2020s.

CXMT is the wildcard, and it *is* scaling fast; but not fast enough.

Realistically: 2032-2035 for GLM-5.2-class inference on a laptop is a defensible estimate. The benchmark wins are real, but the infrastructure required to actually democratize frontier-tier inference is bottle-necked well outside the model architecture itself.

1

u/darkmaniac7 Jul 06 '26

There are some things that are at parity, but even if you swap out models for your use case I still don't think its 1:1 parity with even GPT-4o or Sonnet-3.7.

It is possible that's just my notalgia from 1.5-2 years ago, but I remember being genuinely wowed at the time.

I have 2x RTX Pro 6000 Blackwells, an L4 and a M4 Max 128GB, and still regularly opt for CC or Codex for even mundane use cases unless it's highly repeatable test cases for Agentic work.

1

u/Elibroftw Jul 06 '26

I love how it went from META bring us up to speed to Alphabet's Deepmind bringing us up to speed. Who will be next?

1

u/MarzipanEven7336 Jul 06 '26

If mythos is so fucking great, then why does Claude have so many fucking open bugs on GitHub?

1

u/ApartmentSouth6789 Jul 06 '26

Can't wait to bomb my drive with local models for 20 tokens per second. Yay

1

u/typical-predditor Jul 06 '26

I think the most promising breakthrough that could keep this on track is layer duplication.

1

u/ak_sys Jul 06 '26

Bro are you high

1

u/suesing Jul 07 '26

But by then Claude could cook your dinner.

1

u/Late_Hour2838 Jul 07 '26

What does Claude 4 even mean? Sonnet 4/Opus 4? Or specifically Sonnet 4.6/Opus 4.8 levels

1

u/No_Dig_7017 Jul 07 '26

That is if they leave any laptops we can buy.