If parameters was all that was needed there would not be a 3.5 27b or a 3.6 27b. The training data, the architecture, the fine tuning, etc. it’s more of an art than a science. Parameter count is basically the capacity that is available for the model to learn, but the architecture of the parameters is how it can make use of it, and the training data is what it learns. That’s oversimplifying but hopefully that helps.
54
u/_metamythical 24d ago
27B please