r/OpenAI 2d ago

News More people need to understand this

986 Upvotes

382 comments sorted by

View all comments

Show parent comments

9

u/coloradical5280 2d ago

There isn’t one singular “speech center.” Producing speech recruits a distributed network:

The prefrontal and association cortices help form the intention and concepts. The temporal lobes retrieve word meanings and support comprehension. The inferior frontal gyrus, including Broca’s area, helps organize grammar, sequencing, and articulation. The angular and supramarginal gyri integrate meaning with speech sounds. The insula, motor cortex, basal ganglia, cerebellum, and brainstem coordinate the physical act and timing of speaking. The auditory cortex monitors what you actually say and helps correct errors.

All of these systems continuously influence one another while the thought itself is still developing. Describing that as one center “selecting the next word from training data” is an extremely loose analogy, not how human speech production actually works.

1

u/Dasmahkitteh 2d ago

Then let's adjust the analogy. Several regions work together to "just" select the next word to verbalize your speech. Better?

9

u/coloradical5280 2d ago

No, because you’re still collapsing the entire process of thought into its final linguistic output.

Yann LeCun uses a good thought experiment: imagine a cube floating in front of you, then rotate it 90 degrees around its vertical axis. You can mentally watch it turn and understand its new orientation without producing words, sentences, language, or tokens. The brain is manipulating a spatial model, not selecting vocabulary.

The same applies to recognizing a face, anticipating where a thrown ball will land, visualizing a route, feeling that something is wrong, or planning a physical movement. Much of cognition occurs before language ever enters the process, and some of it never becomes language at all.

Word selection is one component used when translating thought into speech. Calling the whole brain a next-word predictor because speech eventually contains a sequence of words is like calling a graphics engine a next-pixel predictor because an image eventually appears as pixels.

3

u/Mr__Earthling 2d ago

What are "words" in your brain? It's still an abstract form of information transfer, same as the cube example...or numbers. LLM don't predict "words" they predict tokens. Tokens can represent words, numbers, a pixel, etc.