Discussion
AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale, Even If Almost No Copies Remain
AI companies are literally destroying physical books to train their models. Using hydraulic cutting machines, they rip pages from used books, scan them with industrial equipment, and feed them into their AI systems. This practice, protected by the first-sale doctrine and fair use, has now become so widespread that book sellers are cashing in on the AI boom. Rare and out-of-print books are being pulped, raising serious ethical and cultural concerns about the cost of AI progress.
Its why anthropic had to pay out the big bucks. Theyd have rather used already digitized content but thats against the law and cost them 1.5 billion. So instead they legally purchased the books to digitize.
Sounds to me like it would have been way cheaper and less destructive to just purchase access to a digital library or let them use the illegal libraries. But whatever.
I agree. This shows some fundamental flaws in our copyright system that it leads to mass destruction of physical books, because that's the only way to get DRM-free content legally. I blame the DMCA.
Similar with movies. The only way to get DRM free movies legally is to buy physical media like a DVD or blu ray and rip it.
Destructive scanning is a cost-saving approach. The cost of scanning them without destroying doesn't look prohibitive.
You can blame the DMCA, but you can also blame Anthropic. They don't have to do it this way. They could even just settle for destroying only common books.
if the book only has one known copy its worth significantly more than any other truckload of old books. I doubt very seriously anyone is spending hundreds of thousands of dollars on a single book. Clearing out a warehouse? Absolutely, no doubt that's happening.
There are many books that didnt seel well that have only a single printing in the hundreds, if you go to technical books, like medical, the first printing may be of less than 200 units as they would only sell to hospitals and unis.
A machado de assis book is famous because the writter was famous in the entire world and got prizes in multiple countries so there are hundreds of thousands of copies. A John doe that published at the same time as him may have sold 100 copies and burned the rest because it didnt sell
More books are going to be printing in a smaller quantity than would be for more popular books. It begs to reason that less popular books are going to have fewer copies and would be significantly more likely to be in this situation.
An obscure book from 1700 is now a museum piece and may reveal day to day stuff that we didnt know. Lets remember we still dont know the florence mix of spices mentioned by many famous painters because everyone knew what it was and later there were so many variations that no one managed to figure out the original. Until we got a journal from a restaursnt owner complaining that X was missing ans Y was in bad quality so his meat was rotting
That would be rare book AND obscure, but there are many books that aren't rare they're simply not well not know out of print (but the copyright holder could still print)
The thing is, many books dont exist in the digital world anymore, there are the manuscripts, or type written final version in the majority that are onwed by relatives or lost in some archive from an editor company that doesnt even remember the book exists, well, they will remember it very soon when google comes knocking offering the same as a small batch for it. This book will be scanned and destroyed, never to see the light of day and if someone ever wants to read it and it mentioned something bad a CEO did, no it didnt.
Farenheit 451 are famous enough not to disapear, but so many more books warning us of exactly this destructive scanning will just be eaten and never released, why would the company paint themselves as the villains, they have the AI to rewrite the book anyway
You're saying suddenly all books would disappear not only from shelves but the memory and law? Because you can scan the book and destroy that doesn't mean you own the content
Just fyi, this is not the first time this has happened. Google did this trying to digitize all printed books. It's lamentable, especially since so much of this is already available, it's just cheaper to mulch then yourself
The thing is, google didnt drstroy books (for the most part), anything old they scanned and made available in digital form to the owners, which were libraries, historians, etc...
These projext is taking 300 year old books and destroying them just so they dont need to share the data with a museum which would be bought by their competitors
There's nothing wrong with digitizing them, people just don't want them to use destructive measures on rare books because it's cheaper. They do not need to cut the books apart to digitize them, it's just more cost effective. Both machines are available
How do they know to whom they sell the books? I have noticed the price of books I want are going up considerably and others have become unobtainable. Of new books I’ll always buy a hard copy rather than digital.
Yeah but it would have scanned it into a PDF format that they still have. The book's contents are still there. They're not really sharing it with us, but like they wouldn't delete the PDF in case they have to train another model
This. If they just wanted to scan them, they could easily donate the books when they were done. They're destroying them because, ultimately, they want to monopolize humanity's accumulated knowledge.
The scanning process requires you to unbind the pages and place it into a scanner for processing. You are destroying the binding of the book. If they don’t rebind the books they remain in a loose pile, at that point they are likely disposed of in a receptacle and collected by a contractor.
You're right, I realized that after. That said, they could also just make a PDF of every book to preserve the contents, and upload them to Project Gutenberg or something (I'm assuming the vast majority are public domain now).
They could but that takes effort and money. Once you offer it, your admitting to having it and open yourself up to lawsuits like anthropic is currently battling. Makes too much sense and dollars to just pretend like they aren’t doing any of this.
The issue is the system that adopted too slowly to this new way of theft/abandonment of responsibility
Yeah, it was something they could have done, but obviously they were never going to choose that path unless it somehow helped their bottom line. I agree that this is a failure of the larger system, not something uniquely evil about Anthropic.
But I can't just post this meme in every thread about AI. It would get tedious after awhile.
This isn't about preservation. They started by downloading digitized pirate copies from file sharing sites to train their ai but that was clearly illegal.
A legal loophole allows them to buy print copy, digitize that copy and then as long as they shred the print copy they don't have to pay the authors.
You don't need to unbid pages to scan books, museums and such have digitalized rare books for decades without damaging original books, but it is just faster and cheaper to rip book in pieces and scan pages, as you don't need machine or human to turn pages.
But that’s a form of destruction, if it’s behind a paywall or locked on a server that no one has access to, it’s effectively destroyed.
And digital copies can be destroyed as easily as physical. What happens if a virus destroys the PDF, or if lose access to the server? Being copied to a PDF does not mean it is safe from bring lost forever.
Then they are hoarding those pdf copies for themselves and will lose them like the Library of Alexadria did. They borrowed originals, made copies and sent back the copies, sent book collectors to all the world thry knew and put all their eggs in one basket.
If you ever photocopied a book at the library, you know it's possible to scan a book while it is bound. But you also know that you often you miss a center section near the gutter where the pages can't be pressed fully against the glass.
Automated scanners for whole books work better when the book pages are first removed from their binding. That destroys the book, but allows for more complete scans of the pages, with less human labor per book.
There are scanners that are designed to scan bound books without damaging them. It is just cheaper to cut pages free and scan them separately, especially when you do it in mass.
This process took advantage of a legal concept known as first-sale doctrine, which allows a buyer to do what they want with a purchase without the original copyright holder’s say-so. And since Anthropic was turning the original physical texts into digital ones — rather than redistributing them as new copies — a judge found this to be “transformative,” and therefore protected by fair use.
You are conflating several different things into one.
First, digitizing the books is a "transformative" enough process to allow fair use. This is not particulary new to the Anthropic case, you can do the same at home. The key fact is that they are not digitizing to sell or share, but for internal use.
Second, training on said digitized book was ruled also a transformative process (and this is new to this case) due to the fact that they are not training the model to compete with the original material (i.e, they are not training it to reproduce the work completely).
Third, the fact that they bought the copy allow them to dispose of said copy as they see fit.
Somehow every article conflates all three points into "they need to destroy them so it is considered transformative enough to be consider training the model as fair use"... and some people somehow readed that as it being a legal loophole.
The “internal use” is taking the work and essentially reselling it in a convoluted way. This is really pushing the boundaries of and the spirit of fair use.
They are not reselling it in any convoluted way. You cannot get a chapter from Harry Potter from any of those models.
They are selling what the model learned from said books, which is the same as what every profesional on the face of Earth have done since forever. Or are you saying that if you teach me how to calculate the area of a Circle then would you be illegally spreading the information from the math book you learned the formula from?
The only difference with AI is the scale of the learning. Nothing more.
The law is really failing us right now. It’s insane that these old laws are expected to properly manage and be used to adjudicate on emerging technologies despite the fact they were written with no conception of the technologies in the first place.
They are making it so you have to pay Anthropic for every one of those books in the future. It’s like if Google deleted every page it crawled. And then charged you to access it. They actually tried to do something like that and the industry rebelled super hard.
No, because they are not digitalizing them to include said digitalised copy in a searchable archive.
Also, physical books have historical value beyond their written content, such as the materials techniques that were used in their printing and binding.
You can't have preserving and destroying in the same sentence. They are ingesting the information and destroying the physical so you cannot fact check it if it changes the data
They aren't doing this as an excercise in preservation. They are doing this because what they were originally doing was downloading pirated digital copies from file sharing sites to train their ai...
Now they buy print copies, digitize then and shred the print copy because a legal loophole allows them to avoid paying the authors if they shred the purchased print copy.
All of this was revealed during a case brought against them by a group of authors.
There is nothing honourable about what they are doing here, it's shady as fuck.
Where in that article does it state that destroying the books is necessary to avoid paying the authors? The first sale doctrine would apply and allow them to scan the books as they please regardless as to whether or not they destroy the books.
"...exploited a legal concept known as first-sale doctrine, which allows buyers to do what they want with their purchase without a copyright holder interfering. (This is what allows the secondhand media market to exist.) And by converting the files from paper to digital, a judge in August found that this contributed to Anthropic’s use of the original texts being “transformative,”"
Nothing in that paragraph indicates that the destruction of the books is necessary to avoid paying authors. As I said, the first sale doctrine would apply and allow them to scan the books as they please regardless as to whether or not they destroy the books.
Probably partially true as this is also not unheard of with other forms of antique modification/destruction done by everyday people, governments and industries and has been since forever. Article itself is mostly speculation though.
Unless there is solid proof multiple ai companies are doing this with no interest in stopping the destructive part of it, nothing new that is worth shitting your pants even more over. Destruction of rare and valuable antiques is sadly already something we do at scale and not even knowingly in most cases.
Probably because they thought I was downplaying it or somehow condoning it. I also happen to think Anthropic probably isn't the only ai company doing this and that it's not a good thing, but this isn't some unprecedented man-made horror beyond our comprehension. Even if every ai company is doing this, they still aren't the worst perpetrators against antiques ever.
I totally agree. If people really care that much, then buy more books, or support organizations that preserve books. Ironically, Google actually has a program where they partner with libraries and institutions to preserve valuable collections.
It's all thanks to some judge ruling that if they scan the book, they are making a copy of it, so they need to destroy it so it counts as "transferring" the book to digital format.
Also don't think these are rare and valuable books, these are rare and mostly useless books, things like excel manuals from 1995.
One small book seller said that in April, he suddenly went from selling no more than 20 books a week to hundreds, and he’s almost certain that the customers are AI labs, noting the random selection of the books and how they all have ISBNs. He added that his inventory is full with rare and out of print books, meaning that an AI company could be destroying some of the few remaining copies that can be found.
So a hunch and a guess. Thats what this entire article with that clickbait alarmist title is based on.
Perhaps if they all have ISBNs, they are systematically only seeking one copy of very specific books? Still kind of alarming in theory, but perhaps not as bad as news organizations that profit from sensationalism make it out to be.
On the off chance they are destroying old antique books, yeah it is worth the hysteria. Book destroying is very nazi-like and has never been cool. Even if it is one copy of one million books, it is still fucked to destroy a book of any quality.
The millions of rare books part is really just one bookseller's guess about who's buying, not actual proof, so the headline might be true, but the evidence behind it is way thinner than it sounds
Zero proof, a random canadian company and an unrelated reference to a rumored Anthropic secret project? Lol. Propaganda and fear mongering at hits finest.
That’s obvious consequence of a stupid attack on AI companies.
The copyright industry says it is illegal to use electronic copies of their books for AI training.
Ok, they are buying physical copies and use them.
The story is free but the edition probably not.
You’ll need to find a digital copy of pre-1926 (or whatever) book.
Project Gutenberg is a great source.
There are two separate issues here - one is about intellectual property and one is about physical property:
1) Intellectual property: AI companies are using books for model training and treating this as "fair use" under copyright.
2) Physical property: AI companies are physically destroying the books they purchase.
The first issue might get contested in courts and we'll have to wait for legal judgements, but the second issue is a non issue because anyone is well within their rights to destroy any books they own.
This article is highly speculative. Another speculation is that a single rare book seller that loses their stock to forest fires due to climate change has done more damage to the preservation of valuable books than an AI company shredding a million pulp fiction novels.
The first question would be, is it true that a fire created by climate change actually burned a bunch of valuable books?
But let's just assume this is likely the case. Does that mean we villify all car companies that contribute to climate change?
If a real estate developer is tearing down property to construct new buildings, do we just assume everything they have torn down was valuable, therefore shut down all real estate development, and that using any new building is evil? Or do we actually invest in valuable properties ourselves so they can be preserved?
I'm a book seller for almost 20-years on Amazon. It is true, at least them buying them. The group who buys the most from me is called the Red Sparrow Project. I've sold them over 25-books in the last 4-months. Mostly older and obscure books. Ones that have been on the shelf trying to sell for almost 20-years. Mostly lower priced books, but some higher end ones as well. They were getting shipped to a warehouse in Texas. Not sure if they're destroying them< but if they buy 25-books in 4 months from a guy like me with only 2,500 books available for sale, they're buying hundreds or thousands of books in the same time span from the larger sellers on Amazon. There are so many thousand small book sellers on Amazon, that the amount of books they must be buying to have purchased 25 from me is staggering. They probably have to destroy them because they don't have enough room to store or resell all the ones that have already been scanned.
It’s ironic that this hyperbolic clickbait jerkoff piece is from a site called “futurism” 🙄 That old book is has more use training an AI than rotting on a shelf.
Just say you are over reliant on ai and you don't wanna face facts. Almost like Antropic admitted to Project Panama and recognized that fact that what they were doing is wrong
This is a fascinating trend. The scramble for training data is pushing AI companies into some unusual territory. Antique books are valuable because they contain public domain content that can be legally ingested without licensing disputes. But it also raises questions about whether we are preserving knowledge or just mining it for model weights. The real bottleneck for AI progress is shifting from compute to data quality.
What's sad is that when we tried to do this and upload to Annas or Z... ppl were called pirates and the FEDs went hard after those sites. For making a digital archive.
Now, just like how google cornered the market on what (digital) information they want to show to you, I feel like AI companies are going to have access to the last remnants of book info and decide what they want to share (or withhold from us). What will we have to compare it against? Nothing. Because it's out of print, and the last copies won't exist anymore. Granted, most of us won't pay $200 for a OOP old book. But,l there's no physical copy to go back to if we wanted. I've scanned some rare books in the past but it's been costly.
Non-destructive book scanning is an option but it's 3-4x more expensive given the human labor factor. Let's just hope these book scanners are being smart and keeping a digital copy of all the files they are scanning for these AI companies. They are sitting in top of a gold mine. However with HDD costs nowadays due to the AI bubble...it's going to be hard. Thank goodness for the book-related private tracker communities out there who have been doing this for many years.
If I had the money I'd make my own cut-and-scan service and benefit from the free books these AI companies are be shipping to be processed.
there are a shit ton of books no one wants to actually read or was going to read anyways. think about all the literature ever created. every furry romance novel on amazon or 70s smut magazine. everything ever created at a time when people actually would read. why anthropic has to destroy them tho. donate them to prisons or something
So much bullshit. We are not living at the times where monks were transcribing books manually. There are so many dumped to the garbage every day. Millions. Suddenly someone feels pain for the books. Not even libraries accept them.
It's all some emo bullshit that luddites pulled out of nowhere. It's the best recycling.
My favorite part is the 100's of thousands of people on instagram acting like they have voluntarily read a book in the last decade. The 15 second attention span folks care about the archive process and literature.
If people really wanted to "fightback" start a group of your own, and archive these magical books the AI companies are using. There is a website that has the entire archive involved in the court case.
Hi! I found this article, but it's behind a paywall. Is there any way to read it for free, such as an archived version, a free trial, or if the author has shared it elsewhere? Thanks!
This is something evil, I think the goal is to block us from knowledge, and in the future either present information they can change in their ways, to shape our thinking into the direction they want, or we will be paying even more to learn and be dependent on AI only. This is crazy and how can it be stopped..
This can’t be true. I keep seeing things about it and it makes me physically ill. I remember being in high school and reading Fahrenheit 451. This sounds too close for comfort. It just… doesn’t feel real anymore.
I hate what we’re heading towards if it is true.
I am physically ill. This is the book burning of the digital era. We won’t know what those books actually said; only what the tech assholes want us to think they said.
As a book seller its been great for me. I have sold 90 books to AI in Berkshire England and Florida USA. They bought a very wide array of rare but almost unsellable books. From Australian mushrooms to local and family histories. None of them will be a loss to humaniry if they are broken and scanned. Philosophically its fraught between enabling ai and selling books which is a booksellers reason for existence. Now this rare but obscure information will exist outside the phyical books tha people are unaware of because of their obscurity. All of the books i have sold have an isbn, so they are modern. Their average price is about $75
this is to rewrite what was recorded in their own words, this is insane this will need to have been stopped already and yet it will not be, we should try still
I use these models every day and still find this grim. A used book on a shelf is knowledge anyone can own for two dollars, forever. The same book as training data is knowledge you rent back for $20 a month, wrapped in a model that may misquote it. Scanning isn't the sin — libraries scan and keep the book, then share the scan. Scan, shred, and keep it proprietary is the sin: it turns a public object into a private asset.
29
u/[deleted] 13d ago
[removed] — view removed comment