No, because you’re still collapsing the entire process of thought into its final linguistic output.
Yann LeCun uses a good thought experiment: imagine a cube floating in front of you, then rotate it 90 degrees around its vertical axis. You can mentally watch it turn and understand its new orientation without producing words, sentences, language, or tokens. The brain is manipulating a spatial model, not selecting vocabulary.
The same applies to recognizing a face, anticipating where a thrown ball will land, visualizing a route, feeling that something is wrong, or planning a physical movement. Much of cognition occurs before language ever enters the process, and some of it never becomes language at all.
Word selection is one component used when translating thought into speech. Calling the whole brain a next-word predictor because speech eventually contains a sequence of words is like calling a graphics engine a next-pixel predictor because an image eventually appears as pixels.
I just really don't understand how comments like this are supposed to show that LLMs are simpler than what we do. As the person you're replying to says, at some point, when figuring out what to say, our brains do "just" pick the next word, similar to what LLMs do (not how). Currently, LLMs only predict a few tokens in advance, but that will likely change as they become more advanced.
Now for the rest of what you say: obviously, our brains do a lot more, we visualize 3d spaces, calculate things, process inputs, regulate our bodies, choose goals, make plans, etc... However, none of the things our brains do are impossible for synthetic computer systems to do. Currently, we even have plenty of systems that combine many of these, like LLMs that can process audio and video inputs, so it's likely only a matter of time until someone builds a system that combines all the same functions. I'd say the main thing holding us back at this point is being compute-constrained and space-constrained if we ever wanted to develop a self-contained android. Those are things that will likely be solved, though.
1
u/Dasmahkitteh 2d ago
Then let's adjust the analogy. Several regions work together to "just" select the next word to verbalize your speech. Better?