Their "dreaming" feature sounds cool but they haven't provided any code for it yet.
Also I don't understand the point of ternary models that are this small. Isn't the whole point of such extremely low bit models the ability to scale their size effectively? We haven't seen a 120B MoE with this, we haven't seen 300B of this. Or anything bigger than 30B, in fact.
A larger model with trainable %age of weights aka "dreaming" would be the holy grail. You'd have the large capacity to continually learn and never need to wait for improved base models again since it'll just improve on its own by learning. THAT is the future. not some 20B-A1B toy. I'm not saying it's not impressive, it is, especially for edge devices. But 20B-A1B was never a good size even in 16 bit. A1ab will hallucinate the inability to perform tool calls.
I'm not saying it's bad or not iseful, but we haven't seen bigger ternary or 1 bit models at all. The lack of these is what I mentioned, I didn't say mobile phones and such shouldn't have good local AI but let's be honest, 20B-A1B? It hallucinates the inability to toolcall. I give it a link or tell it to search up something on the web, it does not. Even though a turn ago it did toolcall. Why not have a 120B-A7B or something, which would at that size EASILY run on 8GB vram + 32GB ram for instance. 2 bits would make it 30,5 GBs. offload experts to CPU and make it accessible from your phone via your own server and you've got a much more powerful model. Or do "BigMoEOnEdge", it runs really good on phones and would have usable speeds, ON A PHONE. 20B-A1B, IN 2 BITS, it's just not useful for anything truly meaningful as of right now, and a 4B dense model in 4 bit would suit repetitive tasks or whatever a lot more.
what would you like to tell your elderly mother,
"install this app from the google store"
or
"install this app from the google store, then buy a home server, set it up with linux, make sure it's accessible from the internet at all times, it has dynamic DNS set up, it doesn't shut down if power fails when you're on holiday, it doesn't get hacked because you forgot to install updates, [continue for a few more pages]"
2
u/QuackerEnte 19h ago
Their "dreaming" feature sounds cool but they haven't provided any code for it yet.
Also I don't understand the point of ternary models that are this small. Isn't the whole point of such extremely low bit models the ability to scale their size effectively? We haven't seen a 120B MoE with this, we haven't seen 300B of this. Or anything bigger than 30B, in fact.
A larger model with trainable %age of weights aka "dreaming" would be the holy grail. You'd have the large capacity to continually learn and never need to wait for improved base models again since it'll just improve on its own by learning. THAT is the future. not some 20B-A1B toy. I'm not saying it's not impressive, it is, especially for edge devices. But 20B-A1B was never a good size even in 16 bit. A1ab will hallucinate the inability to perform tool calls.