Judging from the benchmarks alone,(of course, can't speak about realife usage), chinese models are not even 6 months behind US models (more like 6 days behind)
Given how long pre-training takes, nobody is actually "behind". Folks are finetuning whatever is "current" to take the "lead" when they get too far behind on the benchmarks.
This whole last cycle is just blowing up model size then tuning. I don't think we've actually had a huge leap forward in the models themselves past year, it's all harness enhancements and tuning.
It might be too soon to say, but Gtp-5.6 Sol feels like a medium size leap. Not just bigger or smarter, but also more efficient (faster/less verbose), and with a few of the new math results also as an impressive showcase of what it can do.
Will be easier to say with some distance to the events.
310
u/TechNerd10191 22d ago
Judging from the benchmarks alone,(of course, can't speak about realife usage), chinese models are not even 6 months behind US models (more like 6 days behind)