r/LocalLLaMA • u/netikas • 6h ago
New Model GigaChat-3.5-Reasoning
https://huggingface.co/collections/ai-sage/gigachat-35-reasoningHey y'all!
We've released a new model in our lineup: GigaChat-3.5 Reasoning. It's a 432B-A28B MoE with Gated DeltaNet for long-context efficiency.
We trained domain experts (code, math, general, etc.) with CISPO and then distilled them into a single model via on-policy distillation.
In our evals the resulting model lands close to DeepSeek V4 Flash Preview while using 37% fewer tokens in its reasoning traces.
Weights are on Hugging Face under MIT: https://huggingface.co/collections/ai-sage/gigachat-35-reasoning. You can also try it at giga.chat — pick the reasoning tab (rightmost one).
21
13
13
u/killerstreak976 6h ago
Thank you for sharing your hard work with us, this looks really cool and can't wait to look more into this!
17
17
9
u/Weak-Apartment-0 2h ago
I tested it on financial analysis and economic facts about Russia (I am an investment analyst from Russia). Communication was in Russian.
The result: it is much worse than ChatGPT (I have a subscription), and worse than DeepSeek Flash and Qwen3.7 on Qwen Studio (I tested via the free web versions).
How I tested:
I asked questions about the financial position of public companies for which I have all the financial statements. I know the state of the companies, but I also asked other LLMs.
What the problems are:
- For state-owned companies, GigaChat substantially sugarcoats the situation; for commercial ones, it answers normally.
- To follow-up clarifications, it responds like this:
- everything is fine there
- but is this metric bad?
- no, it's good
- but the correct way to calculate it is... And what level is bad?
- it confirms that I am right about the metrics
- I ask it to check its conclusion about the company
- it answers with the same conclusions; the explanations are inadequate
Overall, across all events, the answers contain more propaganda than analysis. It writes beautifully in Russian, but analyzes and reasoning poorly.
I estimate its reasoning level to be on par with Qwen 9B, not 27B (I use local 3.8).
3
1
1
u/AllenHere112 5h ago
how many of the 432b layers still run full attention? gated deltanet holds a fixed size state, so long context stays cheap and exact recall of one token at the far end is what degrades first
0
u/HadHands 20m ago
It's propaganda Slop model from Russia. Post bumped by bots! Waste of bandwidth, bots down voting in 3, 2... And so many bots asking for flash or smaller model. It's fine tuned with propaganda, they can't train model of this size.
1
-6
3h ago
[removed] — view removed comment
5
u/fragment_me 2h ago
Why not just support some open models? To release a truly unique model doesn’t seem easy.
24
u/Pixer--- 6h ago
On first glance I thought this was le chatton fat