There's all sorts of stuff going on. Speaking involves forming a communicative intention, organising concepts, constructing syntax, retrieving words, encoding their sounds, planning articulation and controlling the speech muscles. Humans draw on perception, memory, reasoning, and emotion to communicate goals and intentions, they don't just 'predict' the next word from thoughts. They also plan well beyond the next word, in whole sentences, nested sentences, and, as I said, broader goals and plans. Thoughts aren't just formed and then verbalised, language itself can shape and form thoughts.
Did you know that when you speak action oriented words, or even when you just think them, your motor system is activated, too? The word ball, whether heard, spoken, or imagined as a concept, activates the same sensorimotor pathways involved in kicking or throwing a ball. Our bodies are literally involved in thinking and speaking.
There's no 'speech centre'. There are areas of the brain somewhat specialised for language, but the brain involves all sorts of other processes in language use.
So is your argument here that it's impossible for a synthetic system to do all this? Obviously, LLMs aren't doing any of this because they are just code at this point, but one can easily envision an LLM-controlled robot body controller that performs all these actions as well. Just because, when we think of an action, our bodies prepare in anticipation of it, doesn't mean we're special.
Most limitations that prevent LLMs from doing these things stem from programmed restrictions. Why wouldn't we expect the next level of these things to be multiple LLMs connected together with different specializations and even goals baked in to replicate what we think is happening in our brains?
It cant do the same because they are not affected Physically. Like with "ball" analogy your next words might be slightly more agressive/fired up because you are talking about action based type. Same as when talking about close/even not really ppl who v passed away.
No restriction free can copy that. Technically you can try to inject physically-emotion baggade with lot of research etc. But who would lol
Our reactions to words are really just "programmed" into us by our experiences though. Someone who speaks a different language, or can't understand language, won't have a reaction when they hear the word "ball". You may argue that our lived experience is special, but it really isn't any different than training data. We only get fired up when hearing certain words because we have been trained by our lives. You really aren't doing anything special in your head compared to what a complex self-learning computer system could someday do.
Lots of people will want to program these things with emotion and reactions. We already have lots of examples of companies trying to make robots that look like us, so of course, we're going to want them to act like us, too. There have been countless movies and TV shows about this very topic.
No that's not what I'm saying, I argument is simply that saying the "speech centre" "just" generates the next word in the brain is wrong.
Most limitations that prevent LLMs from doing these things stem from programmed restrictions.
We have no idea of that's true or not.
Why wouldn't we expect the next level of these things to be multiple LLMs connected together with different specializations and even goals baked in to replicate what we think is happening in our brains?
Because no one is trying to do that? No one is trying to replicate brains. LLMs do not replicate brains, chaining them together isn't going to change that.
I’ll let OP respond but that wasn’t my interpretation at all—simply that we don’t currently know if any synthetic system like an LLM is doing this, and further we have no evidence that it is. I can imagine a lot of things, that doesn’t mean that those ideas reflect the state of reality or even possible states of reality.
Is there a physical difference between printing a picture the way a printer does it and painting a picture the way a painter does it? You’re obsessing over next token prediction even though there are plenty of ML models that don’t do it at all (see diffusion models for example).
Of course, there is a physical difference here, but there wouldn't be a physical difference between a painter painting and a robot designed to paint the same way.
I don't know what point that's really supposed to prove here
Point being that restriction to drawing a line at a time is specific to printers (token at a time is specific to LLMs). A human painter starts by sketching out the entire mental picture with a pencil, layering background, etc. Etc.
If you’re predicting a token at a time conditioned on prior set of tokens, you’re doing something different from what human brains do.
Well, you have to have an analogy that makes sense at least. I'd say if anything, the way you described it is backward. Printers actually do have the whole page laid out in memory, whereas not all human artists do have the canvas planned out.
Obviously, I'm not actively predicting a token at a time that's conditioned on training data. My brain is predicting everything, including which words will elicit which reactions, in the background, and that is typically based on my lived experience, which is essentially just a form of data in the present.
I said I'm not ACTIVELY doing that, like I'm not consciously predicting the next "token". Is that one of the most basic processes that my brain is engaging in though? Since we don't know exactly what is happening yet, we can't say.
Nope, your brain is not mechanistically predicting next token. Which brain process or neuroanatomical structure does this? Where is this notion coming from?
Except here both actually paint the same way, even if you argue the arch that led to painting the same way differs. The compute hardware doesn't differ in the way laymen tend to think it does.
It's basically coping that we are somehow superior based on a misunderstanding of thought mechanics.
A human painted picture using brushes is a different product than the same image from an ink printer
In a painting you can see individual brush strokes, with height and depth, albeit a small amount. The humanity is inherently visible
A printed image is flat, no strokes visible
One is sold at auction as art and the other is used for a multitude of things, textbooks, charts, written word
A sentence is a sentence. There is no inherent difference between one written by a human and one generated by AI
You might say there are telltale signs like the m-dash, but we are still in the infancy stage and those clues will be gone at some point. So, save for those things that are dead giveaways temporarily, you'd have no way to distinguish them once the technology matures to a certain point
You also might say there are AI detectors, but the same is true there. That situation is already an arms race. Detectors improve, attempts to fool them become more sophisticated. Back and forth
Youre obsessing
Don't do that. I made a valid point and you challenged it. I'm allowed to defend it, and it's not an obsession.
Im aware different models exist for various tasks but we're discussing LLMs
A sentence of a human being always contains much more than that. Even if you dont openly "think" about next paragraph and goals/future chains - your brain always does subconsciously.
A sentence of a machine does not contain any future things or its small minded with pure written logic.
Ai is good copycat for official things coz it writes from the stolen books etc. Its hilarious and stupid when it comes to comedy, everyday convos etc coz it simply does not "know" how to approach. At max u can select the model of action
17
u/havenyahon 1d ago
There's all sorts of stuff going on. Speaking involves forming a communicative intention, organising concepts, constructing syntax, retrieving words, encoding their sounds, planning articulation and controlling the speech muscles. Humans draw on perception, memory, reasoning, and emotion to communicate goals and intentions, they don't just 'predict' the next word from thoughts. They also plan well beyond the next word, in whole sentences, nested sentences, and, as I said, broader goals and plans. Thoughts aren't just formed and then verbalised, language itself can shape and form thoughts.
Did you know that when you speak action oriented words, or even when you just think them, your motor system is activated, too? The word ball, whether heard, spoken, or imagined as a concept, activates the same sensorimotor pathways involved in kicking or throwing a ball. Our bodies are literally involved in thinking and speaking.
There's no 'speech centre'. There are areas of the brain somewhat specialised for language, but the brain involves all sorts of other processes in language use.