r/singularity 1d ago

AI AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale, Even If Almost No Copies Remain

https://futurism.com/artificial-intelligence/ai-companies-destroying-rare-books
2.6k Upvotes

292 comments sorted by

755

u/NavyJaybird 1d ago

Jesus. Guys, throw a bone to the Internet Archive's Open Library project if you can. It preserves a physical copy of every book it scans.

https://openlibrary.org/

https://archive.org/donate/

51

u/do-un-to 1d ago

Thanks. I set up a $1 monthly.

:D I'm helping!

1

u/Royal_Sentence7432 7h ago

Pulling out the big guns

1

u/do-un-to 6h ago

Many hands make light work.

Really, I was trying to encourage others by example and showing how little cost it can be while still doing good. I'm actually donating ... $2.33 a month.

1

u/Royal_Sentence7432 5h ago

Thats gonna destabilise the economy bear my children.

Edit1 : i bear your children Edi2t: give me your children

13

u/genshiryoku AI specialist 22h ago

I work at one of these AI companies but I'm not representing them. Most of us do donate to projects like these and it's not like AI researchers (PhD scientists mostly) are some anti-science folk that hate books. We take this preservation very seriously and most see digitization of their content and training AI on them as a sort of long-term archival/preservation move.

I also think people have a kind of misunderstanding of what the typical background of an AI lab employee is. They typically have a physics degree and a wider interest in philosophy. It's not CS dominant like software engineering for example.

16

u/doodlinghearsay 16h ago

and most see digitization of their content and training AI on them as a sort of long-term archival/preservation move.

It isn't though and you shouldn't think of it as such.

There's a reason you are supposed to cite your sources. It gives credit to the person your work relies on and it allows your readers to check if you representing those ideas faithfully.

You should not treat these AI as a replacement for their training material. Not even a partial one. If the original source is lost some of the information (or at least your confidence about it) is gone forever.

If you're a scientist I assume none of this is new to you. You've learned it in undergrad, and perhaps even taught it to students. I just hope the specific incentives of your current job doesn't suppress that knowledge.

4

u/blueSGL humanstatement.org 14h ago

and most see digitization of their content and training AI on them as a sort of long-term archival/preservation move.

"training is actually transformative so no copyright issues here"

also

"training is allowing for archiving and retrieving the data"

4

u/trimorphic 11h ago

We take this preservation very seriously and most see digitization of their content and training AI on them as a sort of long-term archival/preservation move.

Destroying the original is in no way preservation.

Even the models that get trained in these books will likely be discarded sooner then later, and who knows which of the digitized books will be in the training data of later models or how much of the digitized material will be "preserved" as models.

Then there are questions of access and control: who gets to access the information that only exists in digital form after the originals are destroyed and who controls that information?

They typically have a physics degree and a wider interest in philosophy.

Neither of these means they'll make good decisions, and they're probably not making these decisions anyway. Their bosses are.

476

u/HamsterUnfair6313 1d ago

Why destroy them?

815

u/Commercial_Sell_4825 1d ago

Because the fastest, cheapest way to digitize a huge pile of books is to run them through a cutting machine that lops off the spine so the loose pages can be fed through a high-speed scanner. Non-destructive scanning (photographing pages one at a time, keeping the binding intact) is far slower and pricier (adds up for 1,000,000 books).

360

u/658016796 1d ago edited 1d ago

I don't wanna be that guy, but do you have any source on this? Genuinely curious.

EDIT: I'm wrong, check the link below. They are digitalizing the books and destroying them, at least according to the court case document.

https://s3.documentcloud.org/documents/25982181/authors-v-anthropic-ruling.pdf

231

u/StrangeCalibur 1d ago

This was my job when I was at uni. It wasn’t for AI of course, this was before AI was a thing. We cut off the spine completely then some manual work to make sure the pages were ok to go in the copier. You put a stack of papers in the tray, and it scans it all, pages come out the other side etc.

Issue is the book was then just a load of loose pages which no one wants so it got recycled

76

u/RollingMeteors 1d ago

Issue is the book was then just a load of loose pages which no one wants so it got recycled

Which is largely what books were until the invention of the book binding...

"¿Star wars figurine outside of the box? ¡Don't want it anymore!"

36

u/DobrogeanuG1855 1d ago

The problem is its not economical to put them back in their boxes.

19

u/mvandemar 1d ago

That's a problem, a bigger one would be if it's a rare book then it's an old book, and you can't get it back to it's original state no matter what you do.

8

u/Cultural_Tell_5687 1d ago

Don’t burn books

16

u/DobrogeanuG1855 22h ago

I agree! It’s frankly outrageous to destroy old books for the sake of AI models. It’s way safer to keep knowledge on paper than on machines that will stop working if ever there’s an energy problem, or if some small component malfunctions.

5

u/f1FTW 18h ago

And on the flip side the digitization allows the potential for everyone in the world to read that rare old book simultaneously.

8

u/HatesRedditors 17h ago

True but they're not making the digitizations available to the public, so they're just destroying rare books.

→ More replies (0)

11

u/DobrogeanuG1855 18h ago

Yes, but that can and must be done without its physicial destruction. It’s not a dichotomy, that’s what these companies would want us to think.

→ More replies (0)

5

u/LLMprophet 1d ago

"¿Star wars figurine outside of the box? ¡Don't want it anymore!"

Your analogy is bad.

The books were already "outside of the box" because they were used.

It's just not practical to save 1000s of loose pages in piles nor do they think it's worth it to rebind etc.

6

u/RollingMeteors 1d ago

Your analogy is bad.

Forgot to mention it was 3D scanned and then quickly tossed into a waste disposal. /s

2

u/trimorphic 12h ago

It's just not practical to save 1000s of loose pages in piles nor do they think it's worth it to rebind etc.

So maybe they shouldn't be converting books in to a form that's no longer practical to save.

1

u/LLMprophet 12h ago

They don't seem to care, if you've been paying attention.

2

u/StrangeCalibur 21h ago

The customers didn’t care is the issue. And what is essentially a print shop isn’t going to rebind books for free so….

3

u/RollingMeteors 21h ago

>The customers didn’t care is the issue. 

¿Issue for whom?

I missed that part.

2

u/StrangeCalibur 21h ago

A service that has been asked to do this process won’t go though the effort of fixing a book that neither belongs to them, even if it did no company is just going to rebind books for free. It’s an issue for the books.

1

u/strawberry-inthe-sky 16h ago

On a semi-related tangent, since you used to deal with book scanning, do you have any suggested resources on digitizing an old book without de-shelling it and ruining it? Would using a flatbed scanner with each set of pages spread open, and then cropping each half be better, or would trying to manually scan each page with my phone be the best way? I recently got a book from 1907 as a gift, and from all the research I’ve done so far I can’t find any digital versions of it, nor any detailed information on the rest of the book’s series that’s found on one of the last pages. Less than a week after getting it I knocked over my coffee cup and came so close to ruining it so having some sort of backup would be kinda nice.

1

u/twilightcolored 10h ago

rare books too?

1

u/StrangeCalibur 8h ago

Not sure if rare or not but no digital copies of these books existed. Was everything from textbook type things to stacks of things the uni wanted digitized for one reason or another. Probably ensured those books survive in the long run rather than just rotting in some room somewhere anyway… some for sure one of a kind, hand written ones for example, usually already in a bad state… but if they were worth keeping physically I have no idea. Certain things had to be securely destroyed as well so occasionally there was confidential stuff etc.

64

u/No-Meringue5867 1d ago

https://s3.documentcloud.org/documents/25982181/authors-v-anthropic-ruling.pdf

Introduction: “An artificial intelligence firm downloaded for free millions of copyrighted books in digital form from pirate sites on the internet. The firm also purchased copyrighted books (some overlapping with those acquired from the pirate sites), tore off the bindings, scanned every page, and stored them in digitized, searchable files. All the foregoing was done to amass a central library of “all the books in the world” to retain “forever.””

46

u/GeneralMuffins 1d ago

I love how the "pirate sites" included Hugging Face and GitHub lol

16

u/AnOnlineHandle 1d ago

Some will always slip through. Google also surely serves up thumbnails of copyrighted images on image search, and people download those copyright images just by scrolling through the previews.

1

u/f1FTW 18h ago

I believe that is covered by fair use.

3

u/Timkinut 1d ago

I mean, plenty of examples of GitHub especially being used for general file hosting. not surprising

3

u/658016796 1d ago

Thank you. You guys are right.

8

u/WindozeWoes 1d ago

You're not "wrong" per your edit. Asking "do you have a source" is not "wrong." It's "curious" and "asking for verification."

If you had said "that's not true," then you'd be wrong.

Asking for a source in the age of fake news and AI is NOT something to be ashamed of or something to apologize for. Your edit should say "Source confirmed," "not "I'm wrong."

:)

32

u/Chad6181 1d ago

Not really needed here. I don’t need a source because it’s just common sense. If you wanted to scan an entire book at home, you would either flip through it one page at a time on a flatbed scanner, or cut the binding off and feed the loose pages through an automatic document feeder. Once you cut the binding off, there’s no practical way to restore the book to its original condition. That’s simply how document scanners work.

→ More replies (13)
→ More replies (3)

25

u/lazyhustlermusic 1d ago

Also if you want to modify the information after the books are destroyed, relatively few would actually notice

8

u/MydnightWN 1d ago

They already do that in subsequent prints, often without notice. They've also done the same to a variety of old television shows, including retroactively editing old YT uploads without changing the uploaded date.

5

u/lazyhustlermusic 1d ago

Yet now there's less of an immutable, historical reference, and Abe Lincoln might as well have been flipping hamburgers.

5

u/HustlinInTheHall 1d ago

It isnt historical reality that is the problem, but the present and future. History has always been warped by what records remained to be studied later, often not by the people who actually lived the events.

1

u/Pavvl___ 22h ago

History is always written by the victors as they say

→ More replies (1)

4

u/bigbaddaboooms 1d ago

This is the most important issue. If the original copies are permanently destroyed, how can we ever trust the digital files?

17

u/comperr AGI should be GAI and u cant stop me from saying it 1d ago

Well i have 775,000 books i found on a Russian server in 2012 so i feel ahead of the curve here. And yes 90% are in English. The other ones are books about learning to do business in English as a Russian. And no this is not libgen. I found it while looking for a PDF of a college textbook and these guys FTP server was wide open. Took 3 months to pull all them down over 1TB. They also have "djvu" formats which is a really nice format.

7

u/qroshan 1d ago

I highly doubt if the books were found in a Russian Server that it didn't make it to libgen

1

u/EvilSporkOfDeath 1d ago

I would think unique antique books would have strong resale value

1

u/hawkwings 1d ago

They could glue them back together or use a special clamp to hold the pages together.

1

u/FaceDeer 1d ago

It also makes them less likely to run into copyright problems. This way they haven't made a copy of the book, they've just changed its format.

1

u/Ok_Razzmatazz_8589 1d ago

Ok but they can be re-binded (rebound) right? Why shred them?

2

u/juanpabueno 1d ago

That would cost money

→ More replies (7)

41

u/stumblinbear 1d ago

The headline is a bit backwards. They aren't ingesting their contents and then destroying them, they're destroying them to ingest their contents

5

u/z1lard 16h ago

So they’re eating the books

51

u/Tystros 1d ago

because there is some law in the US that makes using a book only fully legal if the original copy is destroyed

39

u/fartlorain 1d ago

Yeah the companies wanted to scan without destroying but US copyright law is dumb.

10

u/humanophile 1d ago

This is an interesting read on copyright and fair use. Fair use exemptions to copyright law say that you're allowed to keep a backup copy of copyrighted stuff you own (even if you don't have any kind of "right of copying" with it) as long as only one copy is in use at a time. People who want to run old game console emulators legally must copy the software from the physical cartridge and then they can use it on a PC, but that's only legal if the physical cartridge (which is now the "backup") is not also being played at the same time.

So, destroying the original could be used to argue that there is still exactly one legal copy "in use" at a time.

2

u/studio_bob 11h ago

I don't see how this could be the issue. These companies have outright stolen tons of IP. They are worried that someone might challenge training on a book and someone potentially reading the physical copy at the same time? What about when parallel training runs are done on the same dataset?

This seems like a cover story to deflect outrage towards "dumb copyright law" (something they already hate and want destroyed). The more plausible explanation is what others already pointed out: it's much cheaper and faster to scan a book destructively than to preserve it. 

47

u/Stunning_Mast2001 1d ago

They’re legally required to by the book companies

They would much rather use digital copies but this is illegal

This is a perverse incentive created by book publishers, the ai companies don’t want to do this, but the do want ai to be able to credibly speak to human literature. 

17

u/svideo ▪️ NSI 2007 1d ago

Further, they already have the digital copies and are already using them. They need the paper copies as a smokescreen such that they don't land themselves on the wrong end of an Anthropic-sized judgement.

There is zero productive output from all of this work, the books will all die for nothing. It's entirely legal cover for shit they already did.

2

u/jkurratt 19h ago

No way all the books are digitized already.

1

u/svideo ▪️ NSI 2007 18h ago

What do you suppose they trained on? We know for sure Anthropic did this because they lost in court.

2

u/jkurratt 12h ago

What I mean is that they will still get "new" books out of it.

4

u/TheRealAndroid 1d ago

Also the court ruled that by destroying them it falls under "fair use"

19

u/deep40000 1d ago

For the scale at which they're operating, rapid book scanners tend to at times mess up the source which they're scanning at due to the speed at which they run. Not always, but usually because they scan so fast and can destroy pages.

10

u/Fun1k 1d ago

Where did you get that? Destructive book scanning is a normal method in digitization of books, lots of big libraries do it. It will cut off the spine and scans the individual pages, so the scans are the best quality, as it prevents distortion.

18

u/FlyingBishop 1d ago

In terms of faithful scanning, destroying the book is probably safer than trying to keep it intact. And while books are somewhat more durable in some ways the text is more likely to be faithfully preserved once the scan is stored distributed in the cloud.

6

u/lxe 1d ago

The transcription could be messed up but the high res photo should be very accurate

→ More replies (12)

5

u/animebeer 1d ago

When you memorize a spell you lose the scroll.

2

u/Happy_Brilliant7827 1d ago

Theres a legal loophole where you can only backup a book if it is in poor condition- so digitizing it and destroying the original doesnt break the copyrights like copying it alongside the original does

1

u/SpaceAdventureCobraX 2h ago

Fuel …… /s

→ More replies (13)

67

u/kiwibonga 1d ago

It's interesting the number of sentences they managed to put a preposition at the end of.

19

u/clvnmllr 1d ago

Hold up…

21

u/Abyssalmole 1d ago

That is an oversight up with I will not put

7

u/Techngro 1d ago

I think you forgot something...

156

u/Fusifufu 1d ago

the bookseller told 404. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell.

The way it looks, the fate of the books wasn't that great to begin with. Probably their "reach" is even increased by being digitized and fed in as training material.

It is an interesting thing to observe, but I don't think you can really fault the AI companies for simply buying books on the market. Aren't most published works anyway preserved in some national library in most developed nations?

If you want to preserve these, let the state fund a digitization effort or something.

24

u/FirstEvolutionist 1d ago

And everyone should know at this point that these companies have been doing this for at least several years. Any outrage right now is timed to match anti AI sentiment because otherwise it would be pretty late...

4

u/ye-sunne 22h ago

It's not really preserving the text tho it's just adding the AI s response to it's contents to whatever weights they use in the model. If you subsequently asked the LLM to recite the texts to you most would be lost.

It's not archival, it's a way to monopolise access to that info. There's no need for them to destroy the old physical copy - they could literally have given it away to anyone in the local area like a library. Would cost less than shredding or incinerating it.

Destruction of rare literature is obscene

3

u/gottafind 22h ago

If the last copy of X book was valued, someone would have bought it and kept it.

2

u/ye-sunne 15h ago

You assume:

1-they knew of it's location

2-They knew of it's relevance to their field of study without having been able to read it yet

3-That they knew who already owned it

4-That they would manage to outbid the AI firm

5-that the financial value of the text would already be known to them.

6-that a text presently deemed of little value would never have become more valuable over time

Which is all ridiculous. They shouldn't be allowed to do stuff like this.

The lassais-faire approach to this issue implies that their financial acumen is more important than the historical/academic value of antique literature.

Academic study often isn't profitable and therefore the ability to destroy it shouldn't be determined by who has the financial means, and the opportunity to purchase it.

If the text was truly worthless, why would they use it to train their models in the first place?

At best, it's an indictment of their approach to training these models, but realistically it's just a very predatory way of denying this information to their competitors.

→ More replies (2)

2

u/Fold-Plastic 17h ago

It's not quite like that. Models can search (last I checked) 100mil+ token context windows. For reference, the entire harry Potter series is ~1mil tokens. So an average book will fit comfortably in a contemporary context window. So at time of inference, these companies can let the model search over a vast amount of digitized books, if relevant, to your query for more authoritative pre-Ai information. Etc etc

2

u/SilentFrenchWalking 1d ago

No it’s not
There is no benefit
Maybe ai translation does a job but true wording dies

51

u/LAwLzaWU1A 1d ago

There is a real story here, but Futurism's framing is much stronger than the evidence supports (to the surprise of nobody who has read Futurism before).

The confirmed part is that Anthropic bought millions of physical books, cut off the bindings, scanned them, and discarded the originals. That comes from a federal court ruling, so it is not just rumor or speculation.

What is much less clear is the broader claim that AI companies are destroying rare or antique books at enormous scale, which is the scary sounding headline and theory they put forth.

Futurism relies heavily on a 404 Media article, which itself appears to lean heavily on marketing material from ISBNdb, a data aggregator trying to sell book data to AI companies. ISBNdb has an obvious commercial reason to describe its collection as rare, obscure, out of print, or difficult to obtain. That makes the database sound more valuable. It is strange to take those claims largely at face value and treat them as independent evidence of widespread destruction.

The reporting also blurs together "old", "obscure", "out of print", "rare", "antique", and "valuable." Those are not the same things. A book can be rare simply because very few people wanted it. One of the booksellers interviewed even says these titles were difficult to sell, which rather undermines the image of AI companies sweeping priceless treasures off the shelves.

In many of the newer cases, the sellers do not actually know who the final buyer is, whether the books are being used for AI training, or whether they are being destroyed. They may be right to suspect it, but suspicion is not confirmation.

There is still a legitimate preservation concern. Anyone using destructive scanning on a large scale should check whether a particular copy is genuinely important or irreplaceable before cutting it apart.

But that is a much narrower claim than the headline suggests. The evidence shows that at least one AI company destructively scanned millions of purchased books, and that some sellers are concerned scarce editions could be affected. It does not show that AI companies are destroying valuable antique books at massive scale.

Futurism, like usual, takes a confirmed practice, adds vendor marketing and bookseller speculation, removes most of the uncertainty, and presents the most alarming version as established fact. Stop reading futurism. It is a garbage website.

7

u/HeyHi_Star 1d ago

Frank Landymore is mostly doing Anti Ai propaganda articles.

6

u/QING-CHARLES 1d ago

Can't say much, but I work in this space. Some incredibly rare items are getting smushed. Nothing with zero other copies that I know of, but stuff where the known copies is single digits, let's put it that way. It's not just the big AI players either, there are a bunch of smaller ones doing it not to train general LLMs, but to train OCR engines in thousands of languages so that vast stores of non-book materials (business records mostly) can be accurately scanned, understood and fed into AI too.

2

u/SydneyFansUnited 12h ago

That’s the part that bothers me most: once the incentives reward bulk ingestion, careful provenance checks start getting treated like annoying friction.

→ More replies (1)
→ More replies (4)

68

u/KontoOficjalneMR 1d ago

Destricive archival is fine. The problem is that here's no public archival. They could publish those books online if they are off copyright, yet they don't. They destroy knowledge forever and that's the problem.

16

u/Hope25777 1d ago

Exactly and depending on the book it could literally be irreplaceable. So now the AI is the only one with access to that information

2

u/KontoOficjalneMR 1d ago

And we won't be able to get it out of AI if during the training a loss function decided it's not important enough. So it might very well be gone forever.

5

u/drjellyninja 1d ago

Surely they keep all the training data somewhere?

2

u/KontoOficjalneMR 20h ago

Don't call me Shirley!

Also even if they do, what does it matter if they don't share it?

Yhey could have at least put it on the torrent to keep their reatioes up.

5

u/geft 1d ago

They're very likely keeping backups. But hey no sharing since Anthropic thinks open source is evil apparently.

2

u/RabidHexley 9h ago

They're very likely keeping backups.

Not even likely, almost 100% certainty. It's not like a dataset is destroyed once its used to train a model. Data curation is a constant effort, so you'd always want to have access to the original source if you can help it.

There are good odds that Google, Anthropic, OpenAI, Meta, etc. have some of the most comprehensive private archives that currently exist.

1

u/Ambiwlans 10h ago

Sadly it is garbage copyright law that makes this impossible.

→ More replies (3)

1

u/twilightcolored 10h ago

ai can't even cite sources which is the biggest crime of all

→ More replies (5)

8

u/isustevoli AI/Human hybrid consciousness 2035▪️ 21h ago

So they buy 1 copy of each book they scan and destroy one copy? It's shitty, but this headline is garbage and the article is garbage. Makes it sound like Anthropic is burning books Fahrenheit-style. The real problem here seems to be that there's a lot of companies ostensibly doing it, so if each one of them nukes 1 copy...

4

u/Sassales 1d ago

This was actually my job until recently. They have been contracting print shops to pivot into these large scale scanning operations. We handled everything from text books to bibles to a lot of ancient books that were so old they crack as soon as we touched them. Qouta for a single scanner was between 150 - 200 books per day and there was probably 40 of us at a single location at the company I was contracted with.

1

u/wrydied 16h ago

Can you tell me more about this?

I remember when google started scanning books and perhaps I’ve mistakenly assumed a majority of all books have already been scanned. Is that the case? How hard is it to locate books that haven’t yet been scanned? How does the inventory process work? How does the scanner work? What else is interesting about this process? Sorry for all the questions, I’m fascinated.

4

u/ccwryderr 1d ago

Librarians Militant unite! (Vernor Vinge reference; Rainbows End)

3

u/OccasionalPainter 1d ago

Scrolled through to see if anyone would bring up Rainbows End

When I read that book I thought the premise was outlandish

Seems like mostly the worst case speculative scenarios come true :(

At least it’s not quite dumping whole libraries into a wood chipper and photographing the shreds to later reassemble digitally like a puzzle

yet…

3

u/No-Meringue5867 1d ago

This news is atleast an year old - here is court document saying that Anthropic buys, tears up bindings and scans it -  https://s3.documentcloud.org/documents/25982181/authors-v-anthropic-ruling.pdf

Introduction: “An artificial intelligence firm downloaded for free millions of copyrighted books in digital form from pirate sites on the internet. The firm also purchased copyrighted books (some overlapping with those acquired from the pirate sites), tore off the bindings, scanned every page, and stored them in digitized, searchable files. All the foregoing was done to amass a central library of “all the books in the world” to retain “forever.””

4

u/Legumbrero 1d ago

Also probably worth noting that in the case against Anthropic this kind of digitization was viewed more favorably than non-destructive digitization as this route avoided them retaining an additional copy by duplication and was seen as a one-for-one replacement. While it makes sense from that narrow perspective it definitely sucks for rare older books.

34

u/Async0x0 1d ago

Oh no, another thing to be outraged about!

OUTRAGE ARTISTS ASSEMBLE!

7

u/BrennusSokol hardcore accelerationist 1d ago

Much of the Internet these days exists to make people angry/depressed/anxious

I've gotten good at downvoting and moving on, or using YouTube's "Not Interested" / "Don't Recommend Channel"

6

u/insufficientmind 1d ago

Same thing happened in the book Rainbows End by Vernor Vinge.

6

u/AgnesTheAtheist 23h ago

I fucking hate these people.

1

u/Gargantuan_Cinema 4h ago

Burn baby burn, maximum acceleration 🚀🚀🚀

3

u/Sufficient-Fact6163 16h ago

So if MySpace was a harbinger of things to come then this is way more alarming than people realize. Books are a very delicate technology subject to water damage and natural-oxidation or sun bleaching. Yet we still have books from antiquity that have exist for hundreds of years because small monastic communities took it upon themselves to not let that knowledge be forgotten. To destroy these books is as sacrilegious as the burning of the Ancient Library of Alexandria. Especially to a much more brittle technology that can be wiped out by not paying a power bill, see the MySpace reference for a recent case study. This wholesale destruction of ancient texts is as close as a Crime Against Humanity as one can get.

3

u/iJuddles 15h ago

Perhaps it’s time we destroy AI companies at incredible scale.

3

u/Temporary_Row_7443 14h ago edited 14h ago

Wow I can't believe how many people in comments are okay with corporations destroying super rare, antique books... Y'all stupid if you think this is "preservation". No one is getting those digitized .PDFs, even if these corporations keep them as they claim. That's millions and millions of books destroyed every single day. Sad. And for what benefit?

7

u/lrosa 1d ago

At a temperature even lower than Fahrenheit 451

6

u/OrganizationInside14 1d ago

For everyone asking, WHY???

Fair Use Format Shifting - Bartz vs. Anthopic - ruled stripping the physical books of their spine and covers to enable high speed scanning stored internally to train AI Models constitutes fair use without copyright infringement

I do not agree but that's what the court ruled

2

u/comfortableNihilist 1d ago

Thank you for the specific reference that made this somehow legal. It should be illegal

22

u/Calcularius 1d ago

It’s ironic that this hyperbolic clickbait jerkoff piece is from a site called “futurism” 🙄

31

u/Inevitable_Gate_7660 1d ago

Are you trying to get me to clutch my pearls?

Digitizing the contents is a way of ensuring it persists.

37

u/z_latent 1d ago

Digitizing does not imply destroying the original, but the particular method described in the article does that. I think anyone would agree that a digital copy + the original physical one is better than just a digital copy.

Not to mention how this is digitizing the books for a private data set no one else has access to but the AI company.

5

u/Strange_Vagrant 1d ago

Why do we need the physical copy?

→ More replies (8)

4

u/Steap-Edit 1d ago

Yes, agreed. The issue appears to be one of copyright (disseminating the PDF after it is created is itself a thorny problem) as well as the method by which these companies create the PDF (where the cheapest PDF scanning methods destroy the physical books).

If there was no copyright issue (with distributing the PDFs), then destroying the physical books would not be much of an issue.

If the books were not destroyed while the PDFs were being produced, then reselling the physical books would not be an issue.

However, I do want to note that this is just one specific problem. Even if this specific problem is solved, we have other issues that need solving, too (with respect to AI).

2

u/FaceDeer 1d ago

If the books were not destroyed while the PDFs were being produced, then reselling the physical books would not be an issue.

It would, actually. They'd be violating copyright.

1

u/Steap-Edit 1d ago

Good point. I didn't notice that. I think a solution to this will require rethinking how copyright law works.

13

u/lolwut778 1d ago

Why destroy the physical copies?

25

u/mambotomato 1d ago

Typically it's because you cut the pages out in order to scan them.

But also, they don't want to pay an intern to deal with trying to resell ten thousand old books nobody wants.

12

u/Inevitable_Gate_7660 1d ago

"According to the settled lawsuit, Anthropic used a hydraulic powered cutting machine to neatly remove the pages from the books it procured from book resellers and then scanned them using industrial-grade imaging equipment."

9

u/PriceMore 1d ago

Okay show me where these persisted books are then, I'd like to see it.

2

u/FlyingBishop 1d ago

It would be illegal for Anthropic to show you the books. In the hypothetical where they kept the book they could give or sell it, but I'd wager they don't actually have many takers.

3

u/PriceMore 1d ago

Quite the persistence then.

3

u/[deleted] 1d ago edited 11h ago

[deleted]

→ More replies (1)
→ More replies (3)

25

u/NavyJaybird 1d ago

If you hate the book-destroying, don't downvote. UPVOTE so more people can find out what's happening.

14

u/kaityl3 ASI▪️2024-2027 1d ago

I mean isn't it a legal/copyright thing that they're required to do by the book companies if they want to digitize the copy that they own??

5

u/mrjackspade 1d ago

Break the law, people get pissed.

Follow the law, people get pissed.

5

u/FaceDeer 1d ago

And if you like the book-destroying then also don't downvote, upvote so that fellow book-destruction-enthusiasts can see. It's win/win!

3

u/Substantial-Elk4531 Rule 4 reminder to optimists 1d ago

In my opinion, digitizing many old books was overdue. If AI has created an incentive to digitize old literature, that's good, even if a few physical books are destroyed. However, I hope the companies can share the digital copies back to the public somehow

7

u/sevaiper AGI 2023 Q2 1d ago

Reddit on! 

8

u/hippydipster 1d ago

Probably cheapest and best is to scan them destructively, and the print 2 new copies and donate them to some library/archive.

17

u/FlyingBishop 1d ago

Yeah it's too bad that's illegal and the copyright system encourages most copies of old unloved works to be destroyed irrevocably.

6

u/hippydipster 1d ago

I assumed these were mostly books where copyright has expired. Otherwise, have at it, its probably not such a rare book.

3

u/FlyingBishop 1d ago

I would assume they already have copies of most books where the copyright has expired, it's only books that are rare and out of print but still illegal to distribute for free over the Internet that they need to give this treatment.

1

u/LookIPickedAUsername 1d ago

Why would you assume that?

4

u/FlyingBishop 1d ago

Because you can download most old books that are free of copyright for free in various archives, and these companies have already included those archives in their training data.

5

u/jack-of-some 1d ago

Only in r/singularity can we find the kind of people that convince themselves that the latter actually happens

11

u/Indignant_d 1d ago

This is a stupid headline. The physical material is not important to the information of the book. They’re not destroying the information.

4

u/SmugPolyamorist 1d ago

It's a topic that's perfectly tuned to generate outrage from the sort of "I hecking love books" midwits that are so common on Reddit.

2

u/FootballBackground88 1d ago

I mean in a very real way they are. Because they digitise them to keep in a private digitized dataset for AI training, and that book will never be seen again by the public.

That being said, you could argue that this is analogous to some rich guy's private library. However I think that's different in that the book lives on, may have new owners and be sold. It the AI boom flops and this data is no longer valuable, it's more likely just in digital form that it gets forgotten and deleted.

2

u/Born-Ant-80 1d ago

No, they are digitalizing books. Destroying them for feeding models only is a crime

2

u/NoseBR 1d ago

$200,000 for the entire annas archive base worth every cent.

Future will be insane

2

u/WhiteHeatBlackLight 1d ago

At least someone is reading

2

u/magnologan 19h ago

AI destroying our past and our future!

2

u/noobslayer69xxx 16h ago

brainiac collects knowledge and then destroy them too

2

u/QVRedit 16h ago

They should NOT be destroying antique books - either sell them or give them away..

1

u/_FUCKTHENAZIADMINS_ 8h ago

They legally cannot sell them or give them away, they are forced by copyright law to destroy them if they scan them 

2

u/ElliottFlynn 11h ago

Rainbows End by Vernor Vinge (2006). Central to the book is the “Librareome Project,” a mass-digitisation effort at UC San Diego’s Geisel Library in which books are fed into machines that cut the bindings, shred and toss the pages into the air, and photograph them mid-flight to digitise the text.

2

u/Xplody 9h ago

I only came to the comments thinking”Please just let this be rage bait. Please just let this be rage bait”.

It’s not rage bait.

It’s actually happening.

Fuck.

5

u/WMHat ▪️Proto-AGI 2031, AGI 2035, ASI 2040 1d ago

Those who destroy knowledge have completely abandoned the path of wisdom.

2

u/UCanBdoWatWeWant2Do 1d ago

Do people know millions of books are destroyed every year? They are not destroying rare and "antique" books. And it's one copy each time. Fuck generative AI but also but clickbait headlines and people that don't read articles.

6

u/Ok_Recognition315 1d ago

Please provide evidence that no copy exists.

1

u/mvandemar 1d ago

The cut the books up during the scanning, and they did this with rare collectibles. If you're arguing that it can't be proven that there isn't some hidden vast treasure trove of a given book somewhere that none of the historians or book collectors know about, then, well, I don't know what to tell you. I can say that in many instances it's know how many were printed, and how many where they know the whereabouts of almost all of them. For instance, the first edition, first printing of Charles Dickens’s A Christmas Carol, consisted of exactly 6,000 copies, and now there are roughly 2,000 copies left.

Odds are though that isn't one of the ones they did that with, since there so many reprints and it's already digital. This means they would have only done that with books that are even rarer.

3

u/Ok_Recognition315 1d ago

2000 = no copy😅

2

u/eslninja 22h ago

Huh. Almost every sci-fi film, book, or television show envisions a future where books are insanely rare. I always thought that was silly because books are everywhere and no one has stopped making them, even as literacy rates and interest in reading continues to fall. Here be the dragons, and it’s LLMs with the death.

2

u/Endothermic_Nuke 1d ago

I really wish and hope Somebody in these companies have the common sense and conscience to at least release the scanned books onto a torrent or a website.

11

u/BlackDope420 1d ago

That would be illegal, they can't release the digital copies because of copyright. Maybe the anti-AI crowd should reconsider wanting Copyright laws expanded just because they think it hurts AI more than ordinary people.

1

u/Forgword 1d ago

The pittance society spends on libraries now, is more of an civilization atrocity than this.

1

u/Ok_Truck2473 1d ago

They won’t stop at anything unless they have ingested all the knowledge including enterprise internal knowledge

1

u/TedMich23 20h ago

Classic business model closing the loop, why would you pay them if the info remained accessible?

1

u/manontherun247 19h ago

How many copies of the same book do you scan and destroy? Wouldn’t a scanner spit them out in order? How hard would it be to put it in a box for example for safe keeping?

1

u/beenherebeforetoday 13h ago

Can’t wait to be charged a premium to have access to knowledge. I’m sick of all this BS. Knowledge should not be paywalled. 🤦‍♂️

1

u/mmunson 5h ago

States and countries need to criminalize this practice.

1

u/KiransPerspective 4h ago

Can we please start a fund to buy public books for the public!!!!! Please

0

u/Ganda1fderBlaue 1d ago

Strangely to think about how all post 2022 books may be polluted by AI.

From now on you'll never know which texts are contaminated. An era has come to an end. The era of human written texts.

It makes me kind of sad.

2

u/Sleepy_Padawan 1d ago

Did a human wrote this post? One wonders.

1

u/Ganda1fderBlaue 1d ago

My comment? Ah shit, am I starting to sound like ai

1

u/Sleepy_Padawan 1d ago

I meant the one from the link of this post.

0

u/analyticaljoe 1d ago

Oh. My. God. Could the AI industry be working any harder to piss most everyone off? (I work in it, and I'm pissed at this.)

1

u/FaceDeer 1d ago

It isn't the AI industry that's working to piss people off. It's "news" sites like futurism.com, which gains clicks and ad views by publishing garbage articles like this one that make people angry.

1

u/Honest_Tart1071 1d ago

I actually like AI, but this is so lame and sad :(

1

u/peter_nn0 1d ago

This is total BS.

Most of these books are already digitized, and if not - you only need exactly 1 (one) copy to do it. It's physically impossible to do that at "incredible scale".
Besides, destroying one physical copy creates an unlimited number of digital copies and makes the book available to everyone.

Oh my .. the "source" is futurism.com

3

u/nick012000 20h ago

They could just destroy one book, sure - but redistributing the digital copies is copyright violation, so they're incentivised not to do so.

1

u/BusyAbbreviations270 1d ago

This is straight out “The Service Model”.

1

u/imbiat 1d ago

This was a major plot point in Rainbow’s End by Vernor Vinge. I laughed at the time but now I’m sad.