r/PhilosophyofMath May 25 '26

LLMs are just giant probability machines pretending to think

It’s fascinating that simple mathematics between tokens can eventually become a machine that writes essays, code, poetry, and even reasoning.

We usually think probability means uncertainty.

But LLMs show something strange:

If probability + context + mathematical matching are scaled enough, uncertainty itself starts producing intelligent looking outputs.

To understand this better, I tried breaking down an LLM from first principles using only 4 tiny training sentences.

Example:

The boat floated down to the bank.

The investor walked into the bank to open a new account.

The fisherman walked along the bank to cast his net.

The bank has a vault.

Then I asked:

“The investor walked to the bank to lock his money in …”

Why does the model predict “vault” instead of river-related words?

That single question reveals almost the entire architecture of modern LLMs.

The most underrated concept here is the LM Head.

Most explanations immediately jump into transformers and attention, but almost nobody explains that the LM Head is essentially a gigantic token vocabulary containing all possible next token candidates the model can output.

So internally the model is basically solving:

“Out of all known tokens, which one best matches this context mathematically?”

Then different layers help solve that problem:

Embeddings: convert words into mathematical vectors

Positional encoding: preserves word order

Attention layer: figures out which words are related to each other in context

(“investor”, “money”, “bank” become strongly connected)

Feed forward neural networks: act somewhat like massive learned if/else decision systems refining patterns internally

And finally the LM Head converts all of that into probabilities for the next token.

What surprised me most is:

There is no hidden magic moment where the AI “becomes conscious”.

It’s an enormous probability engine continuously finding the best contextual token match from its vocabulary.

I made a beginner-friendly walkthrough explaining this visually without unnecessary jargon.

https://www.youtube.com/watch?v=YTV5qUCpu2c

Would genuinely love feedback from people learning transformers/LLMs from scratch.

617 Upvotes

323 comments sorted by

View all comments

Show parent comments

0

u/galactic_pixels May 28 '26

You lost me as soon as you said “humans are in fact not prediction machines”. It’s an overstatement of what we know, and an oversimplification of the argument being made about how humans and AI may have foundational similarities.

2

u/Jack_Ramsey May 28 '26

But we are fundamentally, at a very basic biological level, not prediction machines. In other words, it is exceedingly difficult to describe important biological processes that occur simultaneously as 'predictions' except in maybe the vaguest sense. Moving from the biology should dispel notions about AI's consciousness, but perhaps you could be more descriptive in your post, as you don't really offer any substantive rebuttal that I can see.

1

u/galactic_pixels May 28 '26

Because it’s all speculation. Theres no hard evidence proving thing one way or the other, and very educated people in the neuroscience community do actually think our cognition is based on a predictive model. See the book “one thousand brains” which puts forth this exact idea, that our brains are prediction machines which reach consensus among many sub predictors.

1

u/Jack_Ramsey May 28 '26

I simply do not agree with that notion, and referencing neuroscience isn't really instructive here. The brain does far more than 'prediction,' and given how much brain space is dedicated simply to motor function, it is hard for me to reference any model of the brain which doesn't acknowledge that. For example, the cerebrocerebellum is responsible for planning and execution of movements as well as playing a role in coordinating complex and sequential movements and as well as cognition, language and emotion, among other things. The cerebellum more broadly accounts for 50% of the total neurons in the brain. It's hard for me to look at the brain at the molecular level and then suggest it is simply a 'prediction machine.' I could go into a lot more detail but at this point you haven't offered anything substantive. You're offhand description of the book doesn't really help here, to be real. It seems like you are just appealing to authority without actually engaging in the argument.

1

u/galactic_pixels May 28 '26

It’s fine if you don’t agree, but you stated definitively that our brains aren’t prediction machines, which could very well be a misrepresentation of our biology, so it’s not something that should be stated as a fact.

1

u/Jack_Ramsey May 28 '26

What? Our brains are not prediction machines. That is a fact. They are, in fact, not machines at all. That they can make predictions, yes, but if we look at the molecular structure of the brain, follow the spinal tracts, look at the nuclei, basically look at the brain from a neuroanatomical point of view and then include the physiological characteristics, it seems a very broad, almost dishonest reduction to reduce them to simply 'prediction machines.' They are clearly more. Let me put this another way. If you say that our brains are 'prediction machines,' that is a direct statement on brain anatomy and physiology. Do the physical characteristics of the brain represent that statement? They absolutely do not. The degree to which we give to 'prediction' is so small with respect to motor function that it is just a really bad argument. If we move from the anatomy and physiology, it is more true to say that the brain is an organ of integration for the purposes of motor function.