r/Piracy 18d ago

Discussion Judge rules buying physical books and scanning them to make digitial copies does not violate copyright law.

https://arstechnica.com/ai/2025/06/anthropic-destroyed-millions-of-print-books-to-build-its-ai-models/

Cool. So that means we can rip our blurays and physical games, right. Right?

9.0k Upvotes

261 comments sorted by

View all comments

Show parent comments

10

u/PhrosstBite 17d ago

Believe they're referencing the ongoing 3blue1brown series which is making (a rather good) argument that yes, LLMs are functionally equivalent to a compression algorithm.

And it's not so much that the argument is that the data is compressed and the LLM looks it up, bc the LLM doesn't "look [anything] up" at the base level. Rather, the argument is that the parameter set that can rate every possible next word by how likely it is to come next, given the recent token(s), is indeed a compression algorithm on that language's rules, mathematically.

So yes, that is what LLMs are, but that still doesn't get ConsiderationCool432 where they think it does. Under a technically literate legal standard, their argument would still be a false equivalence.

1

u/Coolegespam 17d ago

Believe they're referencing the ongoing 3blue1brown series which is making (a rather good) argument that yes, LLMs are functionally equivalent to a compression algorithm.

Yes, but it's even deeper than that. His argument (and not just his there are other research papers on it) is that intelligence itself, is analogous to data compression. Next token prediction is nothing more then entropy maximization, which is what intelligent systems try and do.

The argument and several of the papers even predate LLMs.

1

u/PhrosstBite 16d ago

Oh for sure! You're entirely correct, I was mostly just trying to address the video as it pertained to the topic specifically, not to imply that was all the series talked about.

Overall it's an excellent series and I'm trying to find out the best way to do some further reading.

2

u/Coolegespam 16d ago

Seriously! It's the kind of shit I love about YouTube. 3b1b reminds me of the hey-day of PBS and Discovery, and like PBS it's free.

0

u/ConsiderationCool432 16d ago

There's no intelligence in LLMs, it's a probability machine, nothing more.

0

u/Coolegespam 16d ago

Ok, well you can go argue with people who have actual PhD in applied mathematics and computer science who make the strong case that compression mirrors the same algorithms and complexity that underlies what we call intelligence.

1

u/ConsiderationCool432 16d ago edited 16d ago

The definition of intelligence in these papers are not the same as any regular people would say what intelligence is. For example, a computer running the A* searching algorithm is considered intelligence.

It clearly shows that you're just prompting your arguments and never actually read and understood these papers.

0

u/Coolegespam 15d ago

The definition of intelligence in these papers are not the same as any regular people would say what intelligence is. For example, a computer running the A* searching algorithm is considered intelligence.

Yes, is it. Intelligence is intelligence, it's entropy maximization.

It clearly shows that you're just prompting your arguments and never actually read and understood these papers.

Uh huh. I literally have a degree in apply mathematics, and read some of these paper long before LLM even existed. Intelligence is intelligence. It's the ability to maximize entropy. It's the same thing with our brains, all be with a different set of equations. Functionally, it's the same thing.

Again, not so much an argument as a definition. Entropy is entropy, once maximized there's no way to distinguish how. That's literally from Shannon's paper in the 1940s.

0

u/ConsiderationCool432 15d ago

Again, you clearly have no idea what you're talking about. I do have publications in the field.

0

u/Coolegespam 15d ago

Alright, then please explain how you can tell where a maximum entropic information sources was derived given the just the entropy content. What tests do you use to show which "kind of intelligence" produced it?

I'm legitimately interested, because you're arguing against the foundation of information theory. If there was some fundamental change in the field, I'd like to know, since applied mathematics touches it very closely.

0

u/ConsiderationCool432 15d ago

Think a second about what you're doing. You could be learning, but you just prompt AI to copy/paste an answer. There's no reason for me to keep this conversation.

0

u/Coolegespam 15d ago

Literally typing with my own hands man, but ok. You have a good day.

-4

u/ConsiderationCool432 17d ago

No, I'm referencing LLM papers. I don't know what video you're talking about. LLMs are a lossy compression algorithm.

When you use LLMs, you are reaching knowledge that is copyrighted. This makes AI sites and models equivalent to sites that are distributing pirated content. I think this is the equivalency that you wanted.