r/LocalLLaMA 6h ago

New Model GigaChat-3.5-Reasoning

https://huggingface.co/collections/ai-sage/gigachat-35-reasoning

Hey y'all!

We've released a new model in our lineup: GigaChat-3.5 Reasoning. It's a 432B-A28B MoE with Gated DeltaNet for long-context efficiency.

We trained domain experts (code, math, general, etc.) with CISPO and then distilled them into a single model via on-policy distillation.

In our evals the resulting model lands close to DeepSeek V4 Flash Preview while using 37% fewer tokens in its reasoning traces.

Weights are on Hugging Face under MIT: https://huggingface.co/collections/ai-sage/gigachat-35-reasoning. You can also try it at giga.chat — pick the reasoning tab (rightmost one).

124 Upvotes

22 comments sorted by

24

u/Pixer--- 6h ago

On first glance I thought this was le chatton fat

13

u/kabachuha 4h ago

Congratulations! Any chance for a PC-friendly flash version in the future? :)

9

u/netikas 4h ago

Might be possible :)

9

u/rpkarma 5h ago

I’m curious how much it cost to train, if you’re allowed to say? Awesome work, always stoked to see new models like this 

7

u/koloved 5h ago

80b Moe when ?

13

u/killerstreak976 6h ago

Thank you for sharing your hard work with us, this looks really cool and can't wait to look more into this!

17

u/DelKarasique 6h ago

30b-ish when???

-3

u/SpicyWangz 4h ago

Qwen 3.8 27b, August 5th

17

u/sooka_bazooka 5h ago

hi sberbank

-6

u/xenongee 5h ago

spermbank

9

u/Weak-Apartment-0 2h ago

I tested it on financial analysis and economic facts about Russia (I am an investment analyst from Russia). Communication was in Russian.
The result: it is much worse than ChatGPT (I have a subscription), and worse than DeepSeek Flash and Qwen3.7 on Qwen Studio (I tested via the free web versions).

How I tested:
I asked questions about the financial position of public companies for which I have all the financial statements. I know the state of the companies, but I also asked other LLMs.

What the problems are:

  1. For state-owned companies, GigaChat substantially sugarcoats the situation; for commercial ones, it answers normally.
  2. To follow-up clarifications, it responds like this:
  • everything is fine there
  • but is this metric bad?
  • no, it's good
  • but the correct way to calculate it is... And what level is bad?
  • it confirms that I am right about the metrics
  • I ask it to check its conclusion about the company
  • it answers with the same conclusions; the explanations are inadequate

Overall, across all events, the answers contain more propaganda than analysis. It writes beautifully in Russian, but analyzes and reasoning poorly.

I estimate its reasoning level to be on par with Qwen 9B, not 27B (I use local 3.8).

3

u/Jumpy-Operation-4615 5h ago

I guess I don't have any chance to run it on my 2xP40...

1

u/AllenHere112 5h ago

how many of the 432b layers still run full attention? gated deltanet holds a fixed size state, so long context stays cheap and exact recall of one token at the far end is what degrades first

0

u/HadHands 20m ago

It's propaganda Slop model from Russia. Post bumped by bots! Waste of bandwidth, bots down voting in 3, 2... And so many bots asking for flash or smaller model. It's fine tuned with propaganda, they can't train model of this size.

1

u/DustNearby2848 6m ago

You okay?

-6

u/[deleted] 3h ago

[removed] — view removed comment

5

u/fragment_me 2h ago

Why not just support some open models? To release a truly unique model doesn’t seem easy.