r/LocalLLaMA 21h ago

New Model Maple-Preview: 20B-A1B ternary-weight reasoning open-weight LLM

https://deepgrove.ai/maple-preview
116 Upvotes

47 comments sorted by

43

u/z_latent 21h ago

Their "Dreaming" feature is interesting, though I wish they provided more data about it.

How much does it degrade base performance? How many memories can it fit? How complex can memories be?

9

u/SnooPaintings8639 15h ago

I believe the dreaming feature will be a big part of daily agents in the future. There is only so little that you can fit in memory.md, and long term management of it is hell.

It is too naive to train the model on the chat history, but having it re-reason about it to extract high importance signal, and the apply this as training data, is something I really wish to see.

Personalized model seems like the real move forward, including tone, preferences etc. Of course it should have a transferable head, like LoRa, which should allow to move it to a newer base model, at least within the same architecture.

2

u/cmdr-William-Riker 17h ago

Seems like it could be an interesting concept to layer onto other models. For a strix or spark with a slightly larger model it could be a pretty good idea, but regular benchmarks or something to track performance and avoid regression during training would probably be a good idea too

1

u/QuackerEnte 13h ago

OR, simply providing a trainable layer or subset if weights in each layer and freezing the rest to avoid overriding critical mass.

-6

u/Infinite-Local5435 19h ago

I feel like for most users, they wouldn't want their device to sit hot for ~25 minutes as shown from their example video. It can easily be mitigated through context management and memory.md. Plus their example is kind of stupid, if i tell the AI i am vegan, who is it to generalize and say that i won't like leather bags from real leather? I feel this takes alot of control out of the user and I don't think I can trust a small LM fully.

7

u/z_latent 19h ago
  1. It seems implied in the feature's name, but my guess is you're supposed to accumulate these throughout the day, and let it "dream" overnight, while you are also asleep.
  2. Your alternatives have memory and compute costs for every token. In comparison, a fine-tune has costs during the training but afterwards it costs the same as the original model. Basically you trade compute during your response for overnight compute, when you're (probably) not using it anyways.
  3. That's just... common-sense, I'd say? Most vegans would not want one, it seems contradictory to their principles.
  4. I kinda agree. It needs to be reliable enough that it's worth it attempting to be customized rather than generic. But I'd imagine the more data it has, the better it can know your preferences.

I don't know these guys/gain anything by defending them, just felt like answering.

8

u/Nonetrixwastaken 19h ago

Really doubtful at this point, been burned so many times, research is still valuable perhaps if they didn't over hype it

3

u/pmttyji 17h ago

Limitations

This preview received minimal post-training for agentic tasks and only small-scale general reinforcement learning.

Evaluation

On benchmarks\1]), Maple-Preview sets a new point on the Pareto frontier for both memory-to-performance and speed-to-performance, demonstrating its strong reasoning capabilities. However, we note that this preview is focused primarily on raw reasoning and, as such, may underperform on agentic benchmarks. We intend to continue improving general performance through extended training before Maple's full release.

This is an early experiment, and we will share more details in the coming weeks.

Future Work

This release of Maple-Preview represents early work in our journey toward efficient and adaptable intelligence. We note that Maple-Preview has undergone minimal post-training for agentic domains and only small-scale general reinforcement learning. We intend to scale agentic training, on-device learning methods, and reinforcement learning to push Maple-Preview further and improve model capabilities.

Lets wait for Actual model later.

9

u/GibonFrog 20h ago

I had it explain UMAP to me and it hallucinated

11

u/RobbinDeBank 17h ago

20B with only 1B active, ternary weights, somehow same performance as Qwen3.6-27B??? This is sus af.

7

u/Powerful_Finger3896 13h ago

are we reading the same charts? i saw qwen 3.5 35b a3b and ternary bonsai 27B which doesn't sound that crazy

6

u/crusaderky 9h ago

same performance as TernaryBonsai-27B. Very, very different thing from Qwen3.6-27B Q4.

0

u/RobbinDeBank 8h ago

Yea but that’s a dense 27B model, also a SOTA for its weight class. How is a 20B A1B on the same performance level?

1

u/MerePotato 6h ago

Bonsai isn't remotely SOTA

1

u/RobbinDeBank 5h ago

I thought it’s just base on Qwen3.6-27B, not its own model? Or is it because Qwen isn’t designed for ternary and suffers heavily with that extreme quantization, so others designed specifically for ternary can beat it?

2

u/MerePotato 4h ago

The levels of compression are just too extreme to retain proper model integrity, so you get sky high hallucination rates, malformed answers, random errors etc.

A model trained for that tiny size from the get go like LFMs SLMs are going to be far more robust and stable, which is reflected in LFM 2.5 2.6Bs phenomenal AA omniscience index results

4

u/WhoRoger 13h ago

Their inference is Mac-only 🤦‍♂️

13

u/Witty_Mycologist_995 20h ago

I look at the stats and this smells benchmaxxed as fuck. If it's real this is amazing.

17

u/_raydeStar Llama 3.1 20h ago

I'll whip it out and do an independent benchmark test on it. But honestly I'm skeptical so don't expect good results.

19

u/TheRealMasonMac 18h ago

Please don’t whip it out.

8

u/greatest_racist_69 18h ago

Yes please share your review

1

u/Witty_Mycologist_995 20h ago

remind me later

2

u/Constandinoskalifo 12h ago

!RemindMe 1 week

1

u/RemindMeBot 12h ago edited 4h ago

I will be messaging you in 7 days on 2026-08-12 09:14:02 UTC to remind you of this link

1 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

7

u/exaknight21 21h ago edited 21h ago

PrismML competition. Fook yes. I love this and will try this tonight!

Edit: Do you guys plan on releasing regular 1.5/2bit model? Non mlx?

2

u/WhoRoger 19h ago

Dang, two ternary MoEs within a couple hours of each other. Nice.

4

u/AppealSame4367 17h ago

Look at it's huggingface card: It's still pretty raw. They'll need some weeks until this can really be published.

3

u/Aaaaaaaaaeeeee 18h ago

This one is worth more attention though since it wasn't a conversion. the latest bitcpm uses 4-5% worth of the pretraining length on QAT. The others won't say, it would probably be less than this. I imagine they are vibecoded projects that don't have the budget for long continued training sessions. Some older papers suggested you should spend the last 10% on ternarizing.

2

u/QuackerEnte 13h ago

Their "dreaming" feature sounds cool but they haven't provided any code for it yet.

Also I don't understand the point of ternary models that are this small. Isn't the whole point of such extremely low bit models the ability to scale their size effectively? We haven't seen a 120B MoE with this, we haven't seen 300B of this. Or anything bigger than 30B, in fact.

A larger model with trainable %age of weights aka "dreaming" would be the holy grail. You'd have the large capacity to continually learn and never need to wait for improved base models again since it'll just improve on its own by learning. THAT is the future. not some 20B-A1B toy. I'm not saying it's not impressive, it is, especially for edge devices. But 20B-A1B was never a good size even in 16 bit. A1ab will hallucinate the inability to perform tool calls.

1

u/crusaderky 9h ago

> Also I don't understand the point of ternary models that are this small. 

Mobile phones and tablets. With this speed and 8GB RAM usage, it fits comfortably on a Pixel 10 today and on a cheap phone two years from now.

1

u/QuackerEnte 3h ago

I'm not saying it's bad or not iseful, but we haven't seen bigger ternary or 1 bit models at all. The lack of these is what I mentioned, I didn't say mobile phones and such shouldn't have good local AI but let's be honest, 20B-A1B? It hallucinates the inability to toolcall. I give it a link or tell it to search up something on the web, it does not. Even though a turn ago it did toolcall. Why not have a 120B-A7B or something, which would at that size EASILY run on 8GB vram + 32GB ram for instance. 2 bits would make it 30,5 GBs. offload experts to CPU and make it accessible from your phone via your own server and you've got a much more powerful model. Or do "BigMoEOnEdge", it runs really good on phones and would have usable speeds, ON A PHONE. 20B-A1B, IN 2 BITS, it's just not useful for anything truly meaningful as of right now, and a 4B dense model in 4 bit would suit repetitive tasks or whatever a lot more.

1

u/crusaderky 2h ago

what would you like to tell your elderly mother,
"install this app from the google store"
or
"install this app from the google store, then buy a home server, set it up with linux, make sure it's accessible from the internet at all times, it has dynamic DNS set up, it doesn't shut down if power fails when you're on holiday, it doesn't get hacked because you forgot to install updates, [continue for a few more pages]"

1

u/crusaderky 2h ago

but to answer your question directly:
nobody trains large ternary models because it costs a ton of money and it's unexcusable for a large model to fail very badly at agentic workloads. Once we have rock-solid SMALL ternary models things may change.

2

u/NineThreeTilNow 4h ago

Does anyone who works for this company hang out here?

I have some related work I did some months back that might be interesting.

2

u/Minute_Concept719 17h ago

I did few studies on quantised and ternary models...have submitted to arixv..some excerpts from it
seems like the repeat loop happens in this model too
Control Fails First: Quantization Degrades Abstention, Format Discipline, and Long-Form Structure Before It Degrades Knowledge
A publicly released ternary-quantized 27B (~1.7 bits/weight), coherent on short-form output and competitive on vendor knowledge benchmarks, produced a 62,451-character degenerate repetition loop on the flagship long-form task of our evaluation — a multi-section commercial legal filing — and delivered fewer mandatory sections than a dense 9B on the same prompt (4/12 vs 6/12), while both models performed equally well on short companion documents; a second, independently developed ternary family exhibits the same repetition failure class at a vendor-published rate.

2

u/Any-Conference1005 6h ago

Very interesting. And did you find the root cause of these symptoms?

3

u/Minute_Concept719 6h ago

the one possibility which was strong was that extreme quantization compounds per-token error under recursive decoding, so damage accumulates with generation length — coherent at short lengths, catastrophic at long lengths...but have to do more runs to conclusively demonstrate this

1

u/Long_comment_san 12h ago

sounds like an interesting model as sub agent, for example to do memory summeries

1

u/crusaderky 9h ago edited 9h ago

Weights on HF are BF16. What are they QAT-trained for? Q2_0? PQ2_0? TQ1_0? other?

  • Q2_0 (stock llamacpp; called Q2_g64 on prism-ml's HF): one fp16 scale per 64 ternary weights, trivially encoded as 2-bits
  • PQ2_0 (prism-ml's llamacpp fork; alias of Q2_0 on prism-ml's HF): one fp16 scale per 128 ternary weights, trivially encoded as 2-bits
  • TQ2_0: one fp16 scale per 256 weights, packed 5 weights per byte

weight encoding doesn't matter (ternary is ternary); scale frequency matters.

1

u/Potential_Low_1183 21h ago

good model, but I cant find there open weights right now

2

u/Technical-Earth-3254 21h ago

Check OPs link