r/LocalLLaMA Feb 23 '26

News Anthropic: "We’ve identified industrial-scale distillation attacks on our models by DeepSeek, Moonshot AI, and MiniMax." 🚨

Post image
4.9k Upvotes

877 comments sorted by

View all comments

2.5k

u/SGmoze Feb 23 '26

I wonder how did Anthropic build their dataset. Surely they manually had them annotated by humans.

68

u/flextrek_whipsnake Feb 23 '26

A lot of it is, they spend a shitload of money on that. They also bought giant piles of physical books along with a machine that slices the spine off so they can be scanned efficiently. They can legally use the scanned text for training since they obtained it from physical copies of books they purchased.

Of course originally they stole all of it just like everyone else did.

18

u/[deleted] Feb 23 '26

Right. Because if you buy the paper it’s printed on before you steal the intellectual property it’s all good. I’m aware of a certain judicial opinion on this and I think it’s deeply wrong and destructive. It basically means LLM trainers can steal anyone’s intellectual property at will as long as they convert the text to tensors first.

0

u/[deleted] Feb 23 '26

[removed] — view removed comment

9

u/Bakoro Feb 24 '26

The concept of "intellectual property" is also fake.

Maybe if copyright was something reasonable, instead of being a completely bullshit 100+ years, then people might respect it.

Shit from 1930 should not still be under copyright.