I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
RajT88 9 hours ago [-]
The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
card_zero 9 hours ago [-]
Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.
scarmig 8 hours ago [-]
The hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.
asaddhamani 8 hours ago [-]
But that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.
scarmig 8 hours ago [-]
Dumpsters also don't typically come equipped with a robot scanner and network uplink built in.
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
ralferoo 3 hours ago [-]
> I really don't know what people objecting to this imagine typically happens to old, unwanted books.
In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.
Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.
theshrike79 3 hours ago [-]
So trashed, but with a bunch of extra steps then?
soperj 7 hours ago [-]
They're buying the books from resellers, not rescuing these books from dumpsters. Stop being an apologist.
skeledrew 7 hours ago [-]
Dumpster is where they go when the resellers fail to complete sales.
scarmig 7 hours ago [-]
The magical library in the countryside, to be painfully explicit, does not exist.
Ekaros 7 hours ago [-]
And they should not even be needed. In many places the issue is solved at start. Copy or copies of each commercially produced book is send to national library. Which with tax payer money keeps an archive. Meaning that at least one copy exist for research purposes if needed.
scarmig 7 hours ago [-]
Unfortunately, that's not the case in the United States. The LOC only selects around half of published books to be permanently held. The rest are disposed of (usually returning them to the publisher, donating them to a library, or destroying them).
fmajid 6 hours ago [-]
They should send them to The Internet Archive instead.
exe34 5 hours ago [-]
The internet archive
soperj 7 hours ago [-]
How many times can you post the same thing in a thread?
fmajid 6 hours ago [-]
The Internet Archive tries to be that magical library, but they can only scan and physically archive what is sent to them.
dukeyukey 6 hours ago [-]
If it were legal they may well do that as a public branding exercise. Google already tried and got punished for it!
skeledrew 7 hours ago [-]
It was never available to you/us in the original form either.
red75prime 7 hours ago [-]
...because it is illegal to copy copyrighted material. 70 years later they might do it.
mrweasel 6 hours ago [-]
The thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, when it's free to share.
At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.
novok 6 hours ago [-]
This is solved by law, which is solved by 'we the people' and I bet many AI companies would be fine with something like the equivalent to patent law with bankruptcy escrow to the library of congress, where they must release the scans in 10 years for books that the vast majority will not give a flying shit about. By then the advantage is long gone in data moat.
fmajid 6 hours ago [-]
Since they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this.
rhdunn 6 hours ago [-]
95 years after publication. Many other countries also have an X years after the author's death clause where X varies between countries but is at least 70.
There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.
In short, it's a mess.
mejutoco 4 hours ago [-]
In my opinion this is one of the reasons why libraries should accept any book, even if all they do is examine it and throw it in the trash. This way they would have a chance at finding any treasures that could be regularly dumped in that way.
bulbar 8 hours ago [-]
> They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies?
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
margalabargala 8 hours ago [-]
Right, but if an AI company buys some vanishingly uncommon book, digitizes it, shreds it, and adds the information it contains to their permanent digital library and digests its contents into an AI that is then publicly accessible...are they making that book less accessible, or more?
scarmig 8 hours ago [-]
You've got to compare it to the alternative. Books have a half-life, and the vast majority of these books being purchased are grody, moldering ex-lib copies of books that no one has read in decades. Their other likely outcome is mulching.
margalabargala 8 hours ago [-]
Right, that's my point.
These generally are not books people care about. The information contained therein was doomed.
Now the information has been digitally preserved and a digestion of the information will be made publicly available.
michaelmrose 6 hours ago [-]
[dead]
halsafar 8 hours ago [-]
Can you get the exact text back out with a prompt or not? Having or not having a book isn't fuzzy.
skeledrew 6 hours ago [-]
Funnily the argument made just a few months ago by many rights holders who wanted their pound of flesh was that, if prompted a certain way, exact text could be retrieved.
margalabargala 8 hours ago [-]
Having or not having a book is absolutely fuzzy. If you have a translation, do you have the book? Even if, like the Odyssey, there are hundreds of wildly varying translations? What about an abridged copy? What about the Sparknotes version? If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice? Certainly more so than if you have neither.
fluoridation 6 hours ago [-]
>If you have a translation, do you have the book?
No, you have a translation.
>Even if, like the Odyssey, there are hundreds of wildly varying translations?
Precisely why translations are not considered equivalent to the original text.
>What about an abridged copy? What about the Sparknotes version?
An abridged copy is not a copy of the unabridged version.
>If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice?
No.
I'm honestly surprised these were the questions you chose to ask, when you could have asked what if you have 90% of the pages, or what if most of the pages are missing pieces because the book was shot with a shotgun, or what if the book was scanned and OCRed and all the "rn"s were replaced with "m"s and all the lower case Ls with ones. Hell, is a scan of the book close enough to having the book, or is it far enough that one can no longer be said to have the book anymore?
close04 3 hours ago [-]
The scan and destroy method is what a judge allowed to do in order to have a copy of the book in the training dataset. With the physical copy destroyed there's still only 1 copy in circulation. Once "inside" an LLM I don't know if anyone decided unequivocally that it's copyright infringement or not, and if that counts as a second copy.
There's no technical reason why an LLM couldn't reproduce verbatim some of the training material. It's sort of a lossy statistical compression engine. Enough of the info will survive to the output in the original form. With the amount of data and the commercial nature it's hard to argue fair-use. But nobody tested this in court. I'm not even sure the US wants to ever test this. Why even attempt something that has a non-0 chance to sabotage your most promising industry/bubble in ages?
GPerson 8 hours ago [-]
They’re not supposed to be storing a copy. What they’re doing is destroying their copy after training a model on it.
monocasa 8 hours ago [-]
They're absolutely keeping the digitized copies. They're not going to just train a single model.
derektank 7 hours ago [-]
No, US copyright law allows them to keep a single digital copy. The hypothetical issue is with them maintaining two copies, one digital and one physical, when they only purchased one.
Natsu 8 hours ago [-]
AIs are weirdly bad at quoting stuff in my experience.
But you'd think that the Library of Congress and such would actually prevent stuff from vanishing just by collecting it themselves.
margalabargala 7 hours ago [-]
Bad at quoting, good at digesting and regurgitating. The concepts are preserved even if quotes aren't.
I'd rather a digital copy exist in someone's hands than a rotting physical copy.
subscribed 6 hours ago [-]
But you don't have access to this copy. The digital copy is removed and all you get is paraphrased content. Some frontier models have been explicitly forbidden from recalling exact quotes in system prompt.
It almost seems like you're suggesting that having Claude generate a paraphrased book is as good as having the original book but i don't think that could be your intention?
sharpshadow 7 hours ago [-]
On a similar topic are all those artifacts kept in museum storages for literally eternity.
Maybe AI money can crack open access to it.
qingcharles 7 hours ago [-]
Many are just copyright "orphans", nobody knows who owns the copyright any longer. Maybe the author died and the copyright passed to their estate, but they're not even aware of it.
One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.
willy_k 9 hours ago [-]
Is there a specific book from 40 years ago you have in mind? Asking out of curiosity.
ipaddr 8 hours ago [-]
Books by Zolar are interesting hard to find all editions. The Fearful Void by Geoffrey Moorhouse probably still has 100s of copies available but hate to see it lost.
ErigmolCt 6 hours ago [-]
I suspect rare here often means out of print or commercially obscure, not unique
alightsoul 9 hours ago [-]
There's just a few copies in a single library worldwide which is probably a national or a university library
tptacek 9 hours ago [-]
The 404 story suggested that these are largely vanity press books and instruction manuals for things no longer sold. Implying that these books were almost certainly headed for the recycling center had the AI companies not snatched them up.
alightsoul 8 hours ago [-]
They are useful as a source of non synthetic data, replacing synthetic data consisting of rephrasings of common topics I assume?
enraged_camel 9 hours ago [-]
Also, a lot of these "rare books" are stuff like TV programming magazines from October 1994.
card_zero 9 hours ago [-]
So, have you tried finding out what the programming was in October 1994? Or what cultural ephemera appeared in the TV guides of that era alongside the schedules? Either there's a copy for the week you want in an archive, or somebody's got one for sale, or most often neither. This can piss you off, if as it happened you had a reason to care.
tptacek 9 hours ago [-]
That would make sense as an argument if the natural endpoint of these copies was preservation, and AI was disrupting that. But the natural next step for virtually all these books is to be recycled, not preserved. Books are generally not preserved. It is extremely normal for them to be pulped. Millions and millions of books are pulped every year.
mslt 8 hours ago [-]
Might as well grind up the tablet of complaint to Ea-nāṣir and make cement out it, right? What use could there be in preserving the mundane facets of everyday existence?
dukeyukey 6 hours ago [-]
It's more that, we don't need to preserve a million copies of the same mundane book. Losing a few copies to AI training is fine.
3 hours ago [-]
pfdietz 9 hours ago [-]
Around a million books are destroyed each day in the US.
8 hours ago [-]
mslt 8 hours ago [-]
To play devils advocate, completely on the terms of your argument, would it be better for that particular human artifact to be shredded and its contents melted into an anonymized data pool, or for it to exist in a museum archive, in its original form, such that future generations can better understand what it was like to be alive in 1994?
I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.
Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?
skeledrew 6 hours ago [-]
A museum - or any other building - can only hold so much physical stuff. How much of it do you really want preserved? How do you choose what is preserved (it's an eventually inevitable choice)? Do you save the 1980s stuff but not the 90s? Or save every even/odd year? Some other method? How much direct access do you think people need to pre-digital history?
5 hours ago [-]
throwaway219450 7 hours ago [-]
The BBC has both in some cases, but we know for sure what was broadcast when:
Historic TV guides are also the sort of strange ephemera that people collect. They ought to be digitized like newspapers and other magazines, but this was always the purview of libraries anyway.
unleaded 8 hours ago [-]
Whose job is it to dictate what is and isn't worth saving?
azan_ 8 hours ago [-]
Well for example yours - if you don’t pay for these rare books and don’t store them in good condition, then you have decided that they are not worth saving. Many of these rare books would run you like I don’t know, 1 buck?
runarberg 9 hours ago [-]
At this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.
eru 9 hours ago [-]
My personal wastebook at home is so rare, it's unique. That doesn't mean it needs preservation.
card_zero 9 hours ago [-]
Your what now? Made from your personal waste? That does sound unique.
Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.
eru 9 hours ago [-]
The alternative for many of these un(der)appreciated books is that they will get unceremoniously dumped in the future anyway. The publishing industry and libraries etc dispose off lots and lots of books.
So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs.
P.S. I'm not sure why you need to make fun of your own ignorance? Just look up the word you don't know and don't mention it?
fmajid 6 hours ago [-]
The AI companies’ working assumption is that if someone found it worth printing, it has enough information content to help train a model. That assumption might be invalid with some of the more rambling self-published books, however.
You are giving me too much credit: it's made up evidence. It's an illustration that rarity doesn't equal value, and doesn't depend on whether I actually own a wastebook or ten.
asdfsa32 4 hours ago [-]
Your point is understood but the crux of the issue is a bit similar to capital punishment, the argument is that risk of losing even one innocent person or useful book isn't worth taking; specially considering the value created for society as part of such risky undertaking, whatever it is exercising capital punishment or scanning and destroying books.
eru 3 hours ago [-]
Whenever you build a highway or a bridge or a power plant, your engineers have to put a money value on human life, or at least the worth of a statistical human life. Just to make ordinary engineering decisions.
And refusing to do this exercise just means that you behave as-if you put a really silly number on the value of human life, and probably not consistent between different parts of the project.
So I don't quite agree with these taboos in the absolute.
(I'm still against capital punishment on practical grounds.)
For books it's similar: if you taboo book destruction for the AI training folks, that doesn't rescue books from their ordinary pre-AI life cycle of getting destroyed all the time in the course of running a publisher or a library or a second-hand book store.
In fact, the AI craze is what's giving rare books _value_ and incentivises people to dig them up and preserve them. Or at least preserve them long enough to be scanned.
The scanning might destroy the physical copy of that book, but they save the contents. That's the whole point of scanning after all.
asdfsa32 2 hours ago [-]
I am well aware of the Statistical Value of Life and how it impacts projects. But that is generally used for allocating preventative measures. No road is going to get approved if it is going to result in random deaths by design.
runarberg 9 hours ago [-]
Like I said, at this scale, there are no guarantees for anything. Very likely will there be a unique copy of an invaluable book or letter an the person feeding it to the scanner will not know and the book get destroyed.
Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered.
My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.
eru 9 hours ago [-]
Well, at least afterwards the book is scanned and preserved digitally in their archives of training data.
If the book was just rotting away in some forgotten bookstore, it would more likely be unceremoniously disposed off in the future without anyone scanning it first.
zmmmmm 9 hours ago [-]
It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
hn_throwaway_99 9 hours ago [-]
There is a whole legal saga here that is often misunderstood. Googling "Project Panama" should give more information.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
ivell 7 hours ago [-]
Can they keep backup of the digital copy?
alightsoul 9 hours ago [-]
Ai companies don't use ebooks, because they are more expensive than second hand books
breezybottom 9 hours ago [-]
They absolutely do. Meta torrented 81 terabytes of ebooks. They just have no incentive to pay when the law looks the other way.
hn_throwaway_99 9 hours ago [-]
The entire ironic thing here is that a huge part of those 81 terabytes of ebooks that Meta torrented were directly pirated books from Anna's Archive.
alightsoul 9 hours ago [-]
I meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks
warkdarrior 6 hours ago [-]
On Amazon right now, retail prices for e-book copies are higher than for the corresponding paperbacks.
steelframe 4 hours ago [-]
This is exactly why Amazon has also been doing this acquisition and destructive scanning of millions of books for some time now.
bawolff 9 hours ago [-]
i imagine its because the doctrine of first sale does not apply to ebooks.
freejazz 8 hours ago [-]
> I doubt AI companies would use a single physical book if they could avoid it
They just don't want to pay what the copyright holders want to charge
cm2012 9 hours ago [-]
Yes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.
wesleywt 6 hours ago [-]
They are scanning "rare" books. I presume there are not a lot of copies left to throw out.
joshstrange 2 hours ago [-]
Rare by whose definition?
I’m not aiming this at you directly by: ISBNs or STFU
Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for.
People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy.
It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.
squidbeak 24 minutes ago [-]
> Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this
"A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
"But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
"It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
joshstrange 15 minutes ago [-]
It would be if it made the point you think it is. All of this is more and more hand waving. 75 copies in museums? Then I think we’re good. As for the 18th century books, that’s pure speculation. It’s like a museum saying, “yes we have they prints for sale that they keep buying and destroying but wouldn’t it be a shame if someone destroyed the actual Mona Lisa?”.
And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them. This isn’t complicated. Amazon/etc aren’t breaking into museums and libraries, they are buying books on the open market.
If these books are so rare and important, then why has no one cared until now to actually preserve them?
ranit 2 hours ago [-]
> Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
The OP article sounds quite opposite though - that AI companies are doing exactly this - destroying books so only they have the scanned content.
ctm92 4 hours ago [-]
They also only scan books that are easily and cheaply available, which means they are either not rare or have no significance.
Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.
asdefghyk 3 hours ago [-]
RE "...Instead, they enforce the copyright and force AI companies to shred books they want to ingest...."
Why are AI companies forced to shred books?
jeroenhd 3 hours ago [-]
They want to have a digital copy and the judge ruled they can only keep one copy.
drtgh 1 hours ago [-]
You do not need to shred books to scan them. This only happens if you don't care about preserving the integrity of the books and you want to scan more cheaply.
jeroenhd 26 minutes ago [-]
But their legal framework for being permitted to scan the books en masse (they are "transforming" the book from physical to digital) requires destruction of the original. Otherwise it wouldn't be transforming, it would be duplicating.
ErigmolCt 6 hours ago [-]
Copyright holders are certainly responsible for keeping unavailable works inaccessible, but shredding is mostly an industrial scanning decision, not a copyright requirement
breezybottom 9 hours ago [-]
They don't "force" anything. Trillion dollar AI companies and their owners have as much agency as book publishers.
parineum 9 hours ago [-]
They are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.
freejazz 8 hours ago [-]
If that's true, then why did they pirate so many books?
scarmig 7 hours ago [-]
Is your complaint that they follow copyright law, or that they don't follow copyright law?
freejazz 2 hours ago [-]
I'm not complaining
Planktonne 4 hours ago [-]
A company with no respect for copyright law can't pretend they respect it when it is convenient for them; it's disingenuous, and shows that there is clearly another explanation.
breezybottom 9 hours ago [-]
Since when do AI companies care about copyright law? They're destroying them so their competitors can't use them.
gpt5 9 hours ago [-]
A little meta - I want to point out demagogic/populist comments like these that try to clear all nuance and brush a topic in black and white tend to come from a really small portion of the users here, but the same user (whose account is only 4 months old), dominates posts like this by posting many many bait-like comments in the same post that deviate the discussion away from meaning and insight.
This is just one example, but it has become unfortunately common across all social media platforms.
alightsoul 9 hours ago [-]
Since they had to pay 1.5 billion for it
eru 9 hours ago [-]
> They're destroying them so their competitors can't use them.
That's pretty silly. My competitor can't ride my bike either, and I didn't have to destroy the bike for that to be true.
tptacek 9 hours ago [-]
To do what?
runarberg 9 hours ago [-]
To not destroy rare books.
fenomas 9 hours ago [-]
Not under US copyright law. The Bartz case ruled that if you scan a book and destroy the original it's considered format-shifting and you're likely fine. But not so for keeping the original and using the scan in its place - Internet Archive tried that (in an incredibly limited way), and publishers sued and won.
So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.
tptacek 9 hours ago [-]
When you read "rare books", what are you thinking these are? The 404 article that spun this story up goes into more detail. These are like vanity press books. They're rare because nobody cares about them. The book industry already destroys these books.
runarberg 9 hours ago [-]
I am thinking about a long essay Icelandic author Þórbergur Þórðarson wrote to his pals abroad, and were left abroad. I am thinking about a photo book by an Indonesian naturalist who is famous on Bali, and took amazing photos of wildlife on Sulawesi in 1926, and colored in, and somehow ended up in New Jersey in the 1980s. I am thinking about a collection of essays written by a teenage J.D. Salinger who he left unsigned at a café thinking nobody would want to read them but just leaving it up to chance. Or maybe a Jackson Pollock sketchbook he lost at a party which ended in the host’s bookshelf, and finally at an estate sale.
Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.
tptacek 9 hours ago [-]
You're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines.
This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them?
You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.
dukeyukey 6 hours ago [-]
Can you explain why AI companies destroying these is worse than a library or bookseller destroying these?
breezybottom 17 minutes ago [-]
"Can you explain why Soviets persecuting Jews is any worse than Nazis persecuting Jews?"
9 hours ago [-]
HedonicEscal8r 9 hours ago [-]
If only this complaint was being posted by an organization ideologically opposed to copyright itself!
raincole 9 hours ago [-]
> Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
remus 1 hours ago [-]
> ...AI companies will still do scan'n'destroy because it's just cheap.
There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.
jscd 8 hours ago [-]
Sorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense?
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
scarmig 8 hours ago [-]
Relinquishing copyright does not imply any of the labor you're suggesting. It's the opposite: you're just committing not to perform the labor of pursuing legal action against someone who does digitize and freely distribute the work.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
jeroenhd 3 hours ago [-]
AI companies are buying the physical books, they can turn them into confetti if that's what they want to do. If the physical books are running out, the authors can print and sell more. Or they can sell digital copies so the information is not lost.
The law is currently forcing these companies to destroy the books after scanning them.
bondarchuk 4 hours ago [-]
We the people in the society who have the power to make laws through democratic means are the ones locking these books up.
wotamess 7 hours ago [-]
"want to ingest"
Not "need to ingest"
Copyright holders are capitalizing on laws on the books just like Jeff Bezos companies buying their own copies to shred
So in the end it's really a Congress problem as usual
signa11 5 hours ago [-]
mr. vernor-vinge's "Rainbows End" is oddly prescient ! Highly recommended nevertheless.
sophacles 9 hours ago [-]
If you're buying second hadn books by the lot, you'll get a lot of duplicates and its eaiser to scan wholesale and dedupe in the computers than it is to try to run a sorting operataion on "things".
customguy 7 hours ago [-]
> force AI companies to shred books they want to ingest.
Nothing forces them to shred books, they do it because it's slightly cheaper that way.
hparadiz 7 hours ago [-]
There was a court case where they said that if they copied the books it's not fair use because they didn't own it but if they bought physical copies and then destroyed them then somehow it was fair use because it fell into the niche of personal backups. I forget the details but basically they buy one time prints and destroy them immediately.
customguy 2 hours ago [-]
I had no idea about that, or how fucking bad this actually is:
So I stand corrected: at least some don't do it because it's cheaper (than to buy a license, or simply forego some things), but because they're fucking evil, or so stupid that it effectively is the same as being extremely evil.
skeledrew 7 hours ago [-]
They do it because, for each work, they bought one copy, which they scan and no longer need the physical version of, and would be in copyright violation if they keep more copies than they bought.
annapanna 5 hours ago [-]
>and would be in copyright violation if they keep more copies than they bought.
They can contact the copyright holder and ask/buy a license to make multiple copies.
hparadiz 5 hours ago [-]
yes and they can ask for a pony and then a unicorn too
skeledrew 5 hours ago [-]
For what?
postepowanieadm 7 hours ago [-]
By destroying them they don't copy only convert them into another format.
ajsnigrutin 4 hours ago [-]
It's also the regulation, where most systems still look at "one pirate copy" = "one sale of lost profits", especially when pirates end up in court. If the book (or game or whatever) is not sold anymore in any way where you could give the copyright holder money in an easy accessible way (eg. buy it on amazon, or a local bookstore), they shouldn't be able to claim losses from piracy, since they clearly don't want your money.
On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.
watwut 6 hours ago [-]
Copyright allows you to sell book you have and does not force you to shread it.
They are not forced to shread them by copyright.
9 hours ago [-]
jacobo37 9 hours ago [-]
this is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...
rpdillon 9 hours ago [-]
Wait: the entire premise of copyright is to prevent someone from publishing a book, and a competitor buys a copy, clones it, and sells copies way cheaper because they don't have to pay the author.
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
alightsoul 9 hours ago [-]
This is not a technical problem at all. This is a copyright problem. Anthropic thought it was just a technical problem until they had to pay 1.5 billion after they lost a copyright court case
Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule
rpdillon 2 hours ago [-]
> publishers and probably authors no longer exist
Who is going to claim a copyright violation?
Sha1rholder 8 hours ago [-]
> Anthropic thought it was just a technical problem
Did they? Then why did they "don't want anyone to know about this"? Or do you think their lawyers are dumb?
dukeyukey 6 hours ago [-]
I imagine they know the public would react like the public is reacting. Like obviously this is not worse than what libraries and thrift shops do daily, but it is bad optics.
Sha1rholder 3 hours ago [-]
I don't think so. Libraries and thrift shops usually don't consider it a good idea to damage those precious, hard-to-reprint books. None would say anything if Anthropic only destroyed 1 million Harry Potter or Foundations
footydude 4 hours ago [-]
> "ai" companies are simply burning our cultural context ...
Most developed countries have a 'legal deposit' system with a national archive/national library that requires publishers to send a copy of their works to them. They've existed in some form for centuries in some countries.
Many of them don't have current publisher because no one wants to buy them. I would bet that vast majority of these books have no commercial market.
wesleywt 6 hours ago [-]
I was wondering what the pro book shredding take was going to be. Why destroy the book after scanning? You can create a beautiful library of rare books with all the AI debt bubble.
ls-a 9 hours ago [-]
[dead]
HedonicEscal8r 9 hours ago [-]
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
Cider9986 9 hours ago [-]
You can donate with over 20 different payment methods after making an anonymous account.
Anna's archive is selling data to AI companies. They're essentially saying "hey, don't sell your books to be scanned by AI companies, scan them yourself, and give us the data, so we can sell it to AI companies."
urbnspacecowboy 43 minutes ago [-]
With the important difference that scans sent to Anna's Archive are, you know, archived.
Levitz 9 hours ago [-]
It's an excellent play by them, using moral outrage to the benefit of the project. When life gives you lemons...
tene80i 7 hours ago [-]
Plenty of written works aren’t “information” but rather art. Most piracy is just about people preferring not to pay for novels, TV and film.
TFNA 3 hours ago [-]
Indeed, art. And it is a pretty common position that all people should have access to art and culture. Add up the cost of buying the DVD/Blu-Ray releases for the 1500 or so films that make up the canon of cinema. That's a sum of money daunting even for people in developed countries, let alone most of the world. Piracy is going to be the realistic solution. (And before you say "Use the library", you know well-stocked libraries don't exist in most of the world, right?)
guax 6 hours ago [-]
Or not being able to. Regional licensing, missing and shuffling content.
To watch the world cup I had to spin up a VM in Brazil to watch it with Portuguese narration because the free transmissions are region locked.
I would gladly pay 5 bucks for it if it was possible otherwise and avoid the hassle.
tene80i 5 hours ago [-]
There was no way in your country to pay and watch it? FIFA will have sold the tv rights there to someone, surely. In which case your complaint is what, that it was expensive?
guax 3 hours ago [-]
Free on both, but not with the narration in my native language.
On Brazil the world cup was being transmitted on youtube. in NL only on traditional TV channels or Online for the same channels (all for free but in Dutch).
And literally as I write this I receive an email saying that my youtube premium was raised from 33 to 38 EURO. So there we have piracy getting juicier and juicier.
Xunjin 4 hours ago [-]
We humans, often forget that logistics is always the bottleneck in any Industry.
mrweasel 6 hours ago [-]
There is also a ton of tv shows, movies, music and books that you cannot buy, for now real good reason. I wouldn't be surprised that if in a few years there will be shows and movies that are only exists as pirated versions. With things increasingly only being available on streaming platforms or behind DRM in other ways, we risk looking back on the current era as a black hole 50 years from now.
My concern is that less popular content is just erases, lost in mergers or lost in massive datacenters, never to be seen again.
tene80i 5 hours ago [-]
I agree there is an archival justification. I just don’t think that’s why most people who pirate things are doing it.
dombiscoff 6 hours ago [-]
Art is information, always.
tene80i 5 hours ago [-]
It’s not tautological. Explain why, and particularly why it’s only information, which is the thing that purportedly wishes to be free.
skeledrew 7 hours ago [-]
Maybe begging the question here. If a physical book is rare, doesn't that mean it wasn't available to many in the first place? It seems to me providing its knowledge via LLM, even if it's a private company, benefits more people than if it were sitting in a library somewhere maybe read by a few, or worse in some private collector's set.
I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.
guax 7 hours ago [-]
I think is the scanning for profit and destroying them in the process that angries people.
If they we’re just kept where they were you can always say they’ll eventually be scanned or have that potential.
I do believe the issue is a bit overblown but the core of it sounds reasonable to me.
skeledrew 6 hours ago [-]
Shouldn't they be allowed to do whatever they want with the copy they bought, and so legally own? Isn't that the entire purpose behind "copy right"?
impossiblefork 6 hours ago [-]
Of course they shouldn't.
If the books are rare enough there's a shared cultural value that is being destroyed.
skeledrew 5 hours ago [-]
Why should they buy the books then if they can't do what they want with them? Just leave them wherever they are to rot and eventually be unceremoniously dumped anyway, without being preserved in any form. Some things are just unavoidable.
impossiblefork 56 minutes ago [-]
In order to read them, presumably.
Just because they are for sale doesn't mean that no one else would have bought them. Destroying them obviously destroys them, which obviously destroys culture. After all, if the books is destroyed, it will not be read.
Think of it like this. If I buy an island, and there are bunch of people who live on it, maybe they're day laborers or whatever, is it right for me to expel them, if I don't want them? Can't that be straight up genocide, if they have some culture indigenous to the island?
The same is true for books. If you destroy culture, you destroy culture. There is no magic which triggered "but I owned the physical book" which means you didn't.
It doesn't matter what property relationship you have to a thing: whatever property relationship you may have to it, you still do what you do, i.e. you take whatever action with respect to it as you take with respect to it.
Another good example is food. If there's a shortage of it, and you still have a great deal, you may think "eh, what does it matter if I accidentally burn some, I save time through my carelessness" but if there are others who aren't getting any, you may be killing people, and may be despised for how you use "your property". The same is true here. Why should one not despise one who deprives others of rare texts, that may even be lost, and thus cause cultural destruction of some culture that may actually be rare, by treating the carelessly. That rare book on model building or whittling from the 70s may actually be important.
skeledrew 31 seconds ago [-]
You can't compare an island or food to books though, because the former are only really valuable in their atomic form. Like if you take detailed pictures, people can't live on or eat those pictures.
You scan books though and all the value they provide is still fully available in bit form, because books merely contain information. And of course being in bit form means they can be trivially copied, technically. But there's also the legality, which determines how many copies it's legally permissible to retain. The latter becomes the bottleneck for your preservation of "culture" because said culture could be made trivially available to anyone with an internet-connected device, if only rights holders weren't screaming foul (remember what happened with Internet Archive during/after Covid?[0]).
Ultimately there's no destruction involved when books are scanned, just a change of format which helps to ease management and for legal compliance.
They are, they are in no risk of being arrested. But that does not make it free of moral judgement.
skeledrew 6 hours ago [-]
Why not?
guax 6 hours ago [-]
That's just a fact, not something to argue around. People will judge you for many reasons, some cultural, some political, some ethical, some personal, some valid, some not. Its human nature.
Which is also not illegal and within some bounds and exceptions, a protected right across the globe.
7 hours ago [-]
hoppp 14 minutes ago [-]
Should be illegal to burn books for this reason. Burning 1 book as a protest, fine. But burning a lot to destroy information? Hell no.
sieve 7 hours ago [-]
Physical books and digital content is special in that you can mostly archive their content almost permanently for cheap. Buildings, paintings, idols, living things, natural features of the environment ... not so much.
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
skeledrew 7 hours ago [-]
> copyright
This is what caused the problem in the first place. If people had unrestricted access to content then the world would be a better place. And works wouldn't be so rare that it's worth AI companies buying and destroying them to gain some edge, as well as remain in legal compliance.
sieve 7 hours ago [-]
I completely agree. IPR as a concept is suspect. But it is the world we live in, with timelines extending like crazy. There are works produced before I was born that will remain under protection till long after I am dead.
A targeted modification to the laws could produce most of the benefits for a minor cost.
theshrike79 3 hours ago [-]
Copyright should be globally determined so that if I cant pay fair market value for a piece of content, it's deemed out of copyright protection and I can use whatever means to get it digitally.
Like if a game isn't available for sale anywhere in my region, I can get it without breaking any laws. A book is out of print and I can't pay money for it digitally -> free game.
(Why "fair market value?": So that skeezy publishers don't have an online shop with one physical copy of every book they own for $1Trillion just to fulfill the law)
xvxvx 9 hours ago [-]
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
glimshe 9 hours ago [-]
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
TiredOfLife 6 minutes ago [-]
Also there was a highly discussed paper talking about how “touched by machines” content will kill llms. About a month after the papers first llms trained with “touched by machines” content appeared an the capabilities of the models got huge upgrade by using that dirty content
eru 9 hours ago [-]
> Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time?
No. They also use lots of other methods to get training data.
npn 7 hours ago [-]
No but with 100% clean data you can easily train a model to filter ai generated content.
ACCount37 4 hours ago [-]
As a rule: all high quality text is useful.
There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.
At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.
Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.
jupp0r 9 hours ago [-]
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
shakna 9 hours ago [-]
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
altcognito 9 hours ago [-]
"You own a particular physical copy; you don't possess an abstract transferable 'one-copy license."
There is nothing that says you have to destroy something because you scanned it. This argument has been confusing me since I've seen this pop up.
Edit;
despite the above, looking at the court documents from the Anthropic case, this is pretty close to what they were arguing: “we are just transferring the physical form we purchased, therefore it is legal.”
I still dont think there is a requirement to destroy the book, but since there isn’t a reason to store the book and they can’t sell it, they probably just took the cheapest route. It might be worth an argument that they only purchased the right to use the digital copies while the physical copies exist, but I’m in over my head from a copyright standpint
breezybottom 9 hours ago [-]
Since when do AI companies care about the law? Most of their training data is pirated.
shakna 9 hours ago [-]
The headlines about destruction came not soon after they got rapped on the knuckles and told "no more pirating".
tkel 4 hours ago [-]
No, the destructive book scanning was happening before that.
blooalien 8 hours ago [-]
> Since when do AI companies care about the law? Most of their training data is pirated.
I imagine since the law recently cost one of them truckloads of money for their violations of it?
silcoon 9 hours ago [-]
As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge.
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
Cider9986 8 hours ago [-]
>In a world like that, what motivation would we still have to read?
Why would a reduction in human writers cause a complete reduction in motivation to read? There's millions of books already written and it makes zero sense that people would stop writing. People write for hundreds of reasons other than to make money and they created literature before copyright was a thing.
msftgreed 8 hours ago [-]
The human tradition is storytelling. The idea that storytelling was something a company could own and other's weren't allowed to tell is very, very new in our history.
People write without any profit motive today. It's weird of the OP to think of writing in such a narrow space as commercialization.
silcoon 8 hours ago [-]
> Why would a reduction in human writers cause a complete reduction in motivation to read?
Because there would not be human written books about the present. All books would be about the past. But literature is not stuck in time. Today writers talk about topics and feelings that writers of the last century might never know or experienced. Many people read books to better understand the today world (non-fiction) and to better understand their today feelings (fiction).
> People write for hundreds of reasons other than to make money.
Agree, but most of the writing that we have from the past still came with some form of financial incentives. Shakespeare didn't write all of the compositions just because he wanted to express himself. He was making money with theater performances. Many religious writing got patronage by the church. Dante Alighieri had a career as politician, Plato came from an aristocratic family. Writing was reserved to elites because education was expensive and people had to work for food.
Today we are lucky because education is accessible and printing is cheap.
> they created literature before copyright was a thing.
Copyright wasn't a thing because replicating content was hard. Try to manually copy a book...
womble2 3 hours ago [-]
Notably, Shakespeare lived in a time before copyright and famously ripped off plots from his contempories. I dont think he's a great reference for the argument that there would be no writing about the modern world if you removed copyright.
watwut 6 hours ago [-]
First, people want tonread books that speak to them and their lives, older books are not that.
Second, people write to be read. It takes huge amount of effort too. With no potential reward for it at all, they stop.
Third, we are social animals. If you dont see people reading, if you dont read yourself, you wont even think of writing.
internet_points 6 hours ago [-]
> physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write
there have been many times I've wanted to pay for an ebook but found that the only places that sell it apply DRM to it, so shadow library it is
7 hours ago [-]
protocolture 7 hours ago [-]
>On one side, there’s copyright law, which exists to support the work of creative people.
Which exists to enrich Disney and other large corps, while they hide behind artists as a human shield.
>Artists need some form of reward.
Right, as do artists who use the work of other artists as their starting point. Copyright holders aren't bill and bob artist, they are massive corporate trolls throwing around the weight of almost 100 years of our cultural heritage, sucking the marrow from its bones.
>On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
The books are getting digitised into a permanent record of all our cultural heritage. It just sucks we don't have control over it. If only there was a way we could get them digitised AND control our cultural heritage. HMMMMMMM.
>Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write.
The incentive to write is being killed by slop groups like 20Booksto50K and Kindle which predate AI by at least a decade. AI just lets them work faster.
>We might end up with all of humanity’s books digitized and accessible for free
Excellent
>But there would be no human writers left.
Unlikely, but there would definitely be no Disneys or Conde Nasts left, which is a massively pro social outcome.
>In a world like that, what motivation would we still have to read?
In a world with all books digitised and accessible to read? A huge huge huge incentive. I already partake if books are too expensive where I am. It would take me the rest of my life to read all the books I already want to read. What kind of inane dribble is the idea that copyright makes it interesting to read? I havent even read all of Howard and he's in the public domain (in cool countries at least)
Cider9986 9 hours ago [-]
The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires.
Unrelated: So with this one copy BS are you not allowed to have backups of the data?
0x0000F8_ 9 hours ago [-]
I entirely believe the litigation brought against Internet Archive was secretly sponsored by these exact organizations, because they want to monopolize information to train models.
No data => No models => No competition.
QuantumNomad_ 8 hours ago [-]
IA was in hot water already even before ChatGPT came out.
> ChatGPT […] originally released on November 30, 2022
> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.
Yeah, great logic, that way they are sure there are no extra copies around.
tptacek 9 hours ago [-]
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
Levitz 9 hours ago [-]
The outraged people don't care. They hate AI, and so anything surrounding AI that can be evil is evil. Books are good, AI destroys books, AI is bad.
Furthermore, they like that AI is bad. Because they think it's bad, and being right feels good.
tptacek 9 hours ago [-]
I think people genuinely don't get that book destruction is like a pretty natural part of the book lifecycle.
eru 9 hours ago [-]
And the solution to the water issues can be found in any introductory textbook on the subject: a water price.
> People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book.
Please pardon the tangent: that's what always bothered me about the Borg in Star Trek. Why do they need to assimilate whole species? I'm sure there are enough volunteers in the federation that would join the Borg collective. Even a handful should be enough.
lanyard-textile 5 hours ago [-]
> Why do they need to assimilate whole species?
So then they're gone :)
When the borg is threatened ("threatened") by a race, they choose to effectively end their way of life one way or the other. That is done by assimilation or by death.
If only a single outstanding member of a species becomes a threat, that's a strong signal for their overall potential; therefore everything needs to go, the sooner the better.
eru 3 hours ago [-]
I thought the Borg wanted to get better and better? Leaving the species essentially intact, and absorbing a few members every so often gives you much more material to work with than a one time infusion: they are still evolving and changing.
ivell 7 hours ago [-]
Knowledge of every member of a species is better than few samples? LLM became better due to its huge dataset instead of just few samples.
eru 3 hours ago [-]
You hit diminishing returns pretty quickly. And LLM training only bothers with a single copy of each book; they don't even need to destroy the documents. (Eg they didn't destroy the American declaration of independence to train on it.)
guax 6 hours ago [-]
Its nuanced and complicated but its not fully without reason. I think there would be a compromise of using those books while playing the "we're helping preserve them part" but I don't think that even crosses the mind of most Ai CEOs in an honest way.
The crux of the issue is that it is mostly a PR problem. AI is amazing but being promoted, in the eyes of many, by the worst people imaginable. Very akin, and overlapping in many ways to the crypto crowd.
imperfect_light 7 hours ago [-]
I don't know how it works today, but 20 years ago bookstores wouldn't return unsold books (too expensive to ship) but would simply tear the covers off and throw them in the garbage.
globular-toast 7 hours ago [-]
They pulp books that have many copies surplus to requirement, not the last few copies in second hand book shops.
What if people like food more than AI? Have you considered that?
fenomas 8 hours ago [-]
More and more I feel like anti-AI is a bigger bubble than AI. It seems like every week it expands into a new dimension - anti-Flock protesters tearing down years-old traffic cameras that were used for research into auto accidents, etc.
Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
customguy 7 hours ago [-]
> my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
There are were polls about people being "more worried than hopeful" about satanic cults or alien abductions? With the youth being the most worried about satanic cults? Just like that knee-jerk "it's the bigger bubble", that makes zero sense.
As per the GDC 2026 State of the Industry poll, "52% said gen AI is bad for the industry, nearly double the 30% who held that view last year". But sure, everybody but HN, LinkedIn, and X bros are just luddites clutching pearls in their tiny bubble. They're the weak and stupid ones, and that is why the stupid shit said about them, day in and day out, isn't actually stupid. It all checks out.
fenomas 6 hours ago [-]
[dead]
emtel 7 hours ago [-]
“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
thuruv 9 hours ago [-]
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!
Cider9986 6 hours ago [-]
They likely would not destroy the original books after scanning but apparently it's the legal way to do things because of copyright.
everyday7732 2 hours ago [-]
How long until someone starts making fake rare books to sell to AI companies?
jsphweid 9 hours ago [-]
To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?
Ekaros 8 hours ago [-]
They are ordering books with ISBN. So I take that they are tracking what books they have scanned or pirated already and only picking up what they are missing. As just ordering mass bulk and getting 20 of the same encyclopaedia would be waste.
And I guess something like encyclopaedia would be good example of book they scan. At one point popular, but with most copies destroyed as no one actually wants them anymore.
philipswood 6 hours ago [-]
I'm interested in learning/teaching technologies. Naturally science fiction examples are interesting.
I have found that often LLMs are familiar with the contents of SF books.
But I feel poorer, almost deprived, by the fact that all the LLMs I've checked with have NOT been trained on the contents of Eon by Greg Bear.
Alifatisk 6 hours ago [-]
Does it exist in the web at all? Could be the contents have not been scraped for the dataset yet.
finn888 4 hours ago [-]
The scan existing but staying locked inside a training pipeline is barely better than the book going to a landfill. At least make the raw scans available.
jeroenhd 3 hours ago [-]
Lossily stuffing books into a model through a training process has so far been deemed legal. Enabling piracy by giving away digital copies is not. The internet archive tried to give away digital copies and they got sued to hell and back (which they should've seen coming from miles away).
I'm sure they have some repository available somewhere. They can even sell the digital copies down the line if they're done with them (only once, of course).
Joel_Mckay 2 hours ago [-]
>Lossily stuffing books into a model through a training process has so far been deemed legal
Not in the EU, UK, or US. "AI" companies were already forced to settle their piracy cases, but often they get a free pass by law enforcement via regulatory capture.
The problem is a book author contracted publisher does not assign legal rights of duplication to a company/individual that buys a legitimate print. It can take over 70 years in some places to become public domain.
The core issue is "AI" firms have so much borrowed cash around, that getting a $1.5B fine for being a pirate is taken as a cost of doing business. The law is simply not equipped to handle this type of hyper-scaling criminal act. =3
The piracy lawsuits have so far only deemed that obtaining books through pirating is a violation of copyright. Training the models on them and reselling model access has yet to be ruled illegal, from my understanding, even though there has been plenty of opportunity to.
Anthropic's crime wasn't stealing the contents of books and making a derivative work of it, but torrenting a shitload of books. Had they bought all the ebooks, I don't think the lawsuit would've gone anywhere.
Joel_Mckay 2 minutes ago [-]
>I don't think the lawsuit would've gone anywhere
That is not how copyright/trademark/contract laws treat similar works. Most LLM know about Disney Mickey Mouse, and LLM vector search space proximity results will gravitate more accurate reproductions of protected works regardless of granularity of data.
OpenAI simply canceled a popular service to avoid Disney wrath.
I would say the "AI" firms will keep buying time with all that borrowed cash. Everything that could be scraped has already been stolen, and thus the problem should begin to self-correct. The Shrek movie release market correction history correlation is funny, and a new film is due 2027 in July. =3
xvilka 6 hours ago [-]
It would be nice if we have some "tracking" e.g. 30% of all known books are scanned. So far all information I searched in the Internet about the progress has been patchy. It's also impossible to understand if exact book was ever digitized or not.
Cider9986 6 hours ago [-]
Anna's archive estimates that they have preserved 16% of the world's books.
A genuinely useful project would identify publications that are actually rare, poorly catalogued, or held by only a few libraries, then prioritize those for careful preservation
What evidence do we have that they are "destroying" books?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
9 hours ago [-]
knowaveragejoe 8 hours ago [-]
Yes, we know they're destroying the books. Whether this is actually a problem or another convenient "AI bad" trope remains to be seen.
Aren’t publishers required to deposit a copy with the Library of Congress (in the US), or the British Library (in the UK) etc. to claim copyright?
qwertytyyuu 9 hours ago [-]
I'm sure the AI companies will retain scans of the books for training on newer models
WillAdams 9 hours ago [-]
"Whoever destroys a book destroys a link in the chain of human knowledge"
-- Thos. Jefferson
pkaye 8 hours ago [-]
Public libraries destroy unsold book donations all the time. I often tried to give away some old books I have online and nobody wants them. Some of these books have some nostalgic value to me so I hate to see them just get destroyed so they just lie in my shed.
WillAdams 2 hours ago [-]
Which is something I argue against constantly (and at least my local library tries hard not to) --- discarded books are placed on tables near the children's area usually and folks are free to pick them up (it might be that a few of them are sold, I certainly get a lot of ex-library books when buying on Thriftbooks and Better World Books).
A partial solution there is of course a larger budget and a "last copy" policy where the last copy of a text at least is stored away in deep storage against a future loan.
The local libraries also accept book donations for an annual fund-raising sale.
winrid 8 hours ago [-]
I have a copy of Michael Abrash's Graphics Programming Black Book (it's like 1k+ pages) with DESTROY written in red on the sides. I appreciate that someone saved it and sold it to me for cheap :)
userbinator 7 hours ago [-]
That's an example of a "very much NOT rare" book, as you can easily find dozens of sources of scans online.
pkaye 7 hours ago [-]
I have one of those I got second hand also. :)
8 hours ago [-]
landgenoot 9 hours ago [-]
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
eru 9 hours ago [-]
Most old books aren't worth protecting, and the publishing industry destroys oodles of them as waste that no one wants.
ColdStream 9 hours ago [-]
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
luciana1u 9 hours ago [-]
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
globnomulous 9 hours ago [-]
Is this a reference to Fahrenheit 451?
imperio59 9 hours ago [-]
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
qingcharles 6 hours ago [-]
Not everyone should scan stuff. If you spend any serious time looking through stuff that randos on the Internet have scanned the quality fits the Bell Curve perfectly.
Biggest problems:
- scanning items that are bigger than the scanner platten so the start/end of every line is cut off.
- becoming an "editor": scanning only the pages you think are interesting and skipping intros, forewords, title pages, copyright pages etc
I work in this space. I now require that before scanning a video is made carefully flicking through every page of the item so it can be checked after scanning to ensure all the pages are present and in the original order.
Even the big libraries fuck up. I wanted an intact copy of Harper's Weekly from 1900 that has a big fold-out map in it. It's not clear to the libraries scanning this issue that the map is missing from their copies. None of the copies for sale from dealers have the map. Even when it is still glued into the middle it gets missed by industrial scanners. Google's scan only includes the (blank) back of the folded map.
Luckily GPT was able to track down a copy in a university special collections and fired off an email asking them to scan it. I just got the scan today:
I spend a lot of tokens getting LLMs vision tools to find the missing pages in vintage items and then try to reassemble them from other scans where available.
I'm also splitting up volumes to reupload. A lot of periodicals are only available online as giant multi-gig volume PDFs with all the issues in one file. I have a separate app I wrote to scan all the pages looking for covers so they can be split into PDFs and then identifying the volume/issue/month/year data from the cover or title page.
userbinator 7 hours ago [-]
In the context of books, "scanning" is now more commonly something that should be called "camming" --- you simply point a camera at the book, and take a picture of every page.
Being purchased and juiced for model weights is about as noble of an end as any book could hope for.
jiaosdjf 4 hours ago [-]
BOYCOTT.
It's that simple, these corps have once again broken the social contract, you must not reward them. OpenAI and Anthropic especially, both owned by schizo sociopathic elites. Just use Chinese open models on 3rd party providers or more ethical companies.
This is literally the only power you have outside of Luigi, you're not going to fix anything with a letter writing campaign. We are entering a fight for survival so you really need to step up your game and stop letting elites run over you.
Today it's just books and manipulating society, tomorrow they will track and punish your behaviour and the control will only get worse. These people are pure fucking evil and we need to start acting like it while we literally still have the freedom and privacy to organise.
bawolff 9 hours ago [-]
This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.
derektank 7 hours ago [-]
I mean, everything about intellectual property is kind of inherently artificial tbf.
blooalien 8 hours ago [-]
> It is 100% the consequence of stupid laws.
More like the consequence of being unwilling to change stupid laws once the stupidity of them is discovered. Nope. Gotta double down on the stupidity instead...
mplewis 9 hours ago [-]
Can someone name a rare book that was destroyed as part of AI scanning? I want to know what kind of thing we're losing.
TomK32 7 hours ago [-]
I've read a few of those articles in recent months, both in English and German, and I did read any book title that was rare. The Rare Book & Special Collection Div at the Library Congress considers books published before 1801 as rare. Searching for old books on abebooks is surprisingly hard but I didn't find any for less than 10 Euro and nothing in those articles suggested they were buying up anything but cheap books.
mycall 9 hours ago [-]
Aren't AI companies all about the rare book auctions now?
christkv 3 hours ago [-]
Do you mean rare as in old? I doubt they are destroying old books because most of them are out of copyright and probably available already as text. I imagine this applies to in copyright works and I do NOT condone it but the way this is told it sounds like they are raiding old libraries to destroy first editions of Cervantes.
c0lpan1c 9 hours ago [-]
that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.
alightsoul 9 hours ago [-]
Use a vpn
protocolture 7 hours ago [-]
>It’s outrageous is that it’s legally permissible
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
SanjayMehta 9 hours ago [-]
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
userbinator 7 hours ago [-]
Does Anna's Archive have it now? They scrape tons of sources, Google Books and HathiTrust included.
qingcharles 6 hours ago [-]
A lot of Hathi is locked behind university and library access restrictions. I sometimes have to track down students or someone who has a local library card to get items I need.
SanjayMehta 1 hours ago [-]
Not that I know of, I did upload the PDF to the original libgen but that particular domain has disappeared now.
qingcharles 6 hours ago [-]
I agree they're the best currently available, but a lot of their scans are straight garbage and need to be redone.
SanjayMehta 1 hours ago [-]
A garbage scan is better than no book at all. Fortunately the book I needed was a clean scan.
lazzlazzlazz 5 hours ago [-]
Aren't the AI companies just buying one copy of each book?
And this is supposed to be concerning?
tkel 3 hours ago [-]
The blog post is about rare books. Meaning, AI companies scanning and destroying rare books.
whycombinetor 9 hours ago [-]
It's giving Vishnu, but the world cannot exist without Shiva.
BrenBarn 9 hours ago [-]
It's not a bad idea but we need a multi-pronged approach, with at least one other prong being "destroy the companies that are doing this".
8 hours ago [-]
imperfect_light 7 hours ago [-]
People keep repeating the "rare books" without providing any evidence that they are rare. Anyone who has collected books knows there are massive volumes of old books that can be bought by the pound.
Ekaros 7 hours ago [-]
I feel they are thinking first printing of some mega classic or hand drawn book from antiquity.
I am thinking of some generic romance or thriller from no name author bought at airport... Or some generic book on say birds or animals... Stuff that they would not take if given a full box for free.
tamimio 9 hours ago [-]
I can imagine 100y from now, most if not all books and knowledge are in electronic format or even just as part of an AI, then a wild solar flare wipes out all electronics in a minute..
TomK32 6 hours ago [-]
Books have been declared dead several times of the last quarter century, yet in the EU alone it's still half a million new titles every year and a 40 billion Euro market.
ivell 7 hours ago [-]
Good premise for a scifi novel.
partiallypro 9 hours ago [-]
The hysteria around AI and data centers has hit a precipice. It's actually a bit embarrassing now. I am pretty sure there are foreign adversaries that are trying to stop the US, but I also really blame the AI companies for doing the most horrendous job imaginable in pitching AI to the public. Not a shock that people are against something that tech bros have claimed will destroy everyone's lives in the next 5 years. These books were probably going into a landfill without AI companies getting them, regardless. Tons and tons of books go into the garbage every day.
wotamess 7 hours ago [-]
[dead]
aaron695 9 hours ago [-]
[dead]
Rendered at 12:44:57 GMT+0000 (Coordinated Universal Time) with Vercel.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.
Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.
At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.
There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.
In short, it's a mess.
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
These generally are not books people care about. The information contained therein was doomed.
Now the information has been digitally preserved and a digestion of the information will be made publicly available.
No, you have a translation.
>Even if, like the Odyssey, there are hundreds of wildly varying translations?
Precisely why translations are not considered equivalent to the original text.
>What about an abridged copy? What about the Sparknotes version?
An abridged copy is not a copy of the unabridged version.
>If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice?
No.
I'm honestly surprised these were the questions you chose to ask, when you could have asked what if you have 90% of the pages, or what if most of the pages are missing pieces because the book was shot with a shotgun, or what if the book was scanned and OCRed and all the "rn"s were replaced with "m"s and all the lower case Ls with ones. Hell, is a scan of the book close enough to having the book, or is it far enough that one can no longer be said to have the book anymore?
There's no technical reason why an LLM couldn't reproduce verbatim some of the training material. It's sort of a lossy statistical compression engine. Enough of the info will survive to the output in the original form. With the amount of data and the commercial nature it's hard to argue fair-use. But nobody tested this in court. I'm not even sure the US wants to ever test this. Why even attempt something that has a non-0 chance to sabotage your most promising industry/bubble in ages?
But you'd think that the Library of Congress and such would actually prevent stuff from vanishing just by collecting it themselves.
I'd rather a digital copy exist in someone's hands than a rotting physical copy.
It almost seems like you're suggesting that having Claude generate a paraphrased book is as good as having the original book but i don't think that could be your intention?
One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.
I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.
Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?
https://genome.ch.bbc.co.uk/about
Historic TV guides are also the sort of strange ephemera that people collect. They ought to be digitized like newspapers and other magazines, but this was always the purview of libraries anyway.
https://en.wiktionary.org/wiki/wastebook
Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.
So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs.
P.S. I'm not sure why you need to make fun of your own ignorance? Just look up the word you don't know and don't mention it?
And refusing to do this exercise just means that you behave as-if you put a really silly number on the value of human life, and probably not consistent between different parts of the project.
So I don't quite agree with these taboos in the absolute.
(I'm still against capital punishment on practical grounds.)
For books it's similar: if you taboo book destruction for the AI training folks, that doesn't rescue books from their ordinary pre-AI life cycle of getting destroyed all the time in the course of running a publisher or a library or a second-hand book store.
In fact, the AI craze is what's giving rare books _value_ and incentivises people to dig them up and preserve them. Or at least preserve them long enough to be scanned.
The scanning might destroy the physical copy of that book, but they save the contents. That's the whole point of scanning after all.
Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered.
My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.
If the book was just rotting away in some forgotten bookstore, it would more likely be unceremoniously disposed off in the future without anyone scanning it first.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
They just don't want to pay what the copyright holders want to charge
I’m not aiming this at you directly by: ISBNs or STFU
Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for.
People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy.
It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.
BBC good enough for you?
https://www.bbc.com/news/articles/cp3rprx2wl4o
"A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
"But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
"It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them. This isn’t complicated. Amazon/etc aren’t breaking into museums and libraries, they are buying books on the open market.
If these books are so rare and important, then why has no one cared until now to actually preserve them?
The OP article sounds quite opposite though - that AI companies are doing exactly this - destroying books so only they have the scanned content.
Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.
Why are AI companies forced to shred books?
This is just one example, but it has become unfortunately common across all social media platforms.
That's pretty silly. My competitor can't ride my bike either, and I didn't have to destroy the bike for that to be true.
So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.
Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.
This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them?
You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
The law is currently forcing these companies to destroy the books after scanning them.
Not "need to ingest"
Copyright holders are capitalizing on laws on the books just like Jeff Bezos companies buying their own copies to shred
So in the end it's really a Congress problem as usual
Nothing forces them to shred books, they do it because it's slightly cheaper that way.
https://en.wikipedia.org/wiki/Project_Panama
So I stand corrected: at least some don't do it because it's cheaper (than to buy a license, or simply forego some things), but because they're fucking evil, or so stupid that it effectively is the same as being extremely evil.
They can contact the copyright holder and ask/buy a license to make multiple copies.
On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.
They are not forced to shread them by copyright.
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule
Who is going to claim a copyright violation?
Did they? Then why did they "don't want anyone to know about this"? Or do you think their lawyers are dumb?
Most developed countries have a 'legal deposit' system with a national archive/national library that requires publishers to send a copy of their works to them. They've existed in some form for centuries in some countries.
Example for the UK - British Library guidance: https://www.bl.uk/services/legal-deposit
I support Anna's Archive, by the way. Information wants to be free.
https://annas-archive.gl/donate
To watch the world cup I had to spin up a VM in Brazil to watch it with Portuguese narration because the free transmissions are region locked.
I would gladly pay 5 bucks for it if it was possible otherwise and avoid the hassle.
On Brazil the world cup was being transmitted on youtube. in NL only on traditional TV channels or Online for the same channels (all for free but in Dutch).
And literally as I write this I receive an email saying that my youtube premium was raised from 33 to 38 EURO. So there we have piracy getting juicier and juicier.
My concern is that less popular content is just erases, lost in mergers or lost in massive datacenters, never to be seen again.
I can't help feeling there's some hypocrisy or something here with this call to be outraged at AI companies and scan books now. What about before when they were still mostly locked away from the world? It's only when they're actually being made available to - at least a part of - the broader world that they're a "cultural heritage" worth preserving. Shame.
If they we’re just kept where they were you can always say they’ll eventually be scanned or have that potential.
I do believe the issue is a bit overblown but the core of it sounds reasonable to me.
If the books are rare enough there's a shared cultural value that is being destroyed.
Just because they are for sale doesn't mean that no one else would have bought them. Destroying them obviously destroys them, which obviously destroys culture. After all, if the books is destroyed, it will not be read.
Think of it like this. If I buy an island, and there are bunch of people who live on it, maybe they're day laborers or whatever, is it right for me to expel them, if I don't want them? Can't that be straight up genocide, if they have some culture indigenous to the island?
The same is true for books. If you destroy culture, you destroy culture. There is no magic which triggered "but I owned the physical book" which means you didn't.
It doesn't matter what property relationship you have to a thing: whatever property relationship you may have to it, you still do what you do, i.e. you take whatever action with respect to it as you take with respect to it.
Another good example is food. If there's a shortage of it, and you still have a great deal, you may think "eh, what does it matter if I accidentally burn some, I save time through my carelessness" but if there are others who aren't getting any, you may be killing people, and may be despised for how you use "your property". The same is true here. Why should one not despise one who deprives others of rare texts, that may even be lost, and thus cause cultural destruction of some culture that may actually be rare, by treating the carelessly. That rare book on model building or whittling from the 70s may actually be important.
You scan books though and all the value they provide is still fully available in bit form, because books merely contain information. And of course being in bit form means they can be trivially copied, technically. But there's also the legality, which determines how many copies it's legally permissible to retain. The latter becomes the bottleneck for your preservation of "culture" because said culture could be made trivially available to anyone with an internet-connected device, if only rights holders weren't screaming foul (remember what happened with Internet Archive during/after Covid?[0]).
Ultimately there's no destruction involved when books are scanned, just a change of format which helps to ease management and for legal compliance.
[0] https://www.libraryjournal.com/story/internet-archive-loses-...
Which is also not illegal and within some bounds and exceptions, a protected right across the globe.
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
This is what caused the problem in the first place. If people had unrestricted access to content then the world would be a better place. And works wouldn't be so rare that it's worth AI companies buying and destroying them to gain some edge, as well as remain in legal compliance.
A targeted modification to the laws could produce most of the benefits for a minor cost.
Like if a game isn't available for sale anywhere in my region, I can get it without breaking any laws. A book is out of print and I can't pay money for it digitally -> free game.
(Why "fair market value?": So that skeezy publishers don't have an online shop with one physical copy of every book they own for $1Trillion just to fulfill the law)
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
No. They also use lots of other methods to get training data.
There's no "2022 split", and the "untouched by machines" bit came from the marketing blurb of a company offering book scanning services - not the AI labs themselves.
At the AI lab level: the book scanning seems to be driven by copyright concerns, not data contamination concerns. There was a concern about AI contamination, but there's no measurable performance loss from ingesting post-2022 data with minimal filtration, and some tests attribute small but persistent performance gains to post-2022 AI contamination. It's unclear where exactly do those gains come from.
Why is all high quality text useful? The "inverse problem" framing is that all text reflects the thinking behind it, somewhat, and by learning to reproduce it, LLMs implicitly learn to reproduce some of the thought process too. They don't just memorize the dry factual knowledge, but also learn how that knowledge fits together, and how to reason about that knowledge - both in the specific case and in general. And that "in general" then surfaces in an LLM's ability to generalize. Which is very desirable.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
There is nothing that says you have to destroy something because you scanned it. This argument has been confusing me since I've seen this pop up.
Edit; despite the above, looking at the court documents from the Anthropic case, this is pretty close to what they were arguing: “we are just transferring the physical form we purchased, therefore it is legal.”
I still dont think there is a requirement to destroy the book, but since there isn’t a reason to store the book and they can’t sell it, they probably just took the cheapest route. It might be worth an argument that they only purchased the right to use the digital copies while the physical copies exist, but I’m in over my head from a copyright standpint
I imagine since the law recently cost one of them truckloads of money for their violations of it?
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
Why would a reduction in human writers cause a complete reduction in motivation to read? There's millions of books already written and it makes zero sense that people would stop writing. People write for hundreds of reasons other than to make money and they created literature before copyright was a thing.
People write without any profit motive today. It's weird of the OP to think of writing in such a narrow space as commercialization.
Because there would not be human written books about the present. All books would be about the past. But literature is not stuck in time. Today writers talk about topics and feelings that writers of the last century might never know or experienced. Many people read books to better understand the today world (non-fiction) and to better understand their today feelings (fiction).
> People write for hundreds of reasons other than to make money.
Agree, but most of the writing that we have from the past still came with some form of financial incentives. Shakespeare didn't write all of the compositions just because he wanted to express himself. He was making money with theater performances. Many religious writing got patronage by the church. Dante Alighieri had a career as politician, Plato came from an aristocratic family. Writing was reserved to elites because education was expensive and people had to work for food.
Today we are lucky because education is accessible and printing is cheap.
> they created literature before copyright was a thing.
Copyright wasn't a thing because replicating content was hard. Try to manually copy a book...
Second, people write to be read. It takes huge amount of effort too. With no potential reward for it at all, they stop.
Third, we are social animals. If you dont see people reading, if you dont read yourself, you wont even think of writing.
there have been many times I've wanted to pay for an ebook but found that the only places that sell it apply DRM to it, so shadow library it is
Which exists to enrich Disney and other large corps, while they hide behind artists as a human shield.
>Artists need some form of reward.
Right, as do artists who use the work of other artists as their starting point. Copyright holders aren't bill and bob artist, they are massive corporate trolls throwing around the weight of almost 100 years of our cultural heritage, sucking the marrow from its bones.
>On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
The books are getting digitised into a permanent record of all our cultural heritage. It just sucks we don't have control over it. If only there was a way we could get them digitised AND control our cultural heritage. HMMMMMMM.
>Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write.
The incentive to write is being killed by slop groups like 20Booksto50K and Kindle which predate AI by at least a decade. AI just lets them work faster.
>We might end up with all of humanity’s books digitized and accessible for free
Excellent
>But there would be no human writers left.
Unlikely, but there would definitely be no Disneys or Conde Nasts left, which is a massively pro social outcome.
>In a world like that, what motivation would we still have to read?
In a world with all books digitised and accessible to read? A huge huge huge incentive. I already partake if books are too expensive where I am. It would take me the rest of my life to read all the books I already want to read. What kind of inane dribble is the idea that copyright makes it interesting to read? I havent even read all of Howard and he's in the public domain (in cool countries at least)
Unrelated: So with this one copy BS are you not allowed to have backups of the data?
No data => No models => No competition.
> ChatGPT […] originally released on November 30, 2022
https://en.wikipedia.org/wiki/ChatGPT
> On March 24, 2020, following shutdowns caused by the COVID-19 pandemic, the Internet Archive opened the National Emergency Library, removing the waitlists used in Open Library and expanding access to these books for all readers. More than one user could borrow a book at the same time. Two months later, on June 1, the National Emergency Library (NEL) was met with a lawsuit from four book publishers. Two weeks after that, on June 16, the Internet Archive closed the NEL, and the prior Open Library CDL system resumed after the 12 weeks of NEL usage.
https://en.wikipedia.org/wiki/Hachette_v._Internet_Archive
It is not natural in the industry to preserve books! It's tricky to even give most books away. Our library has big donation boxes, and my understanding is: most of those books are destroyed.
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
Furthermore, they like that AI is bad. Because they think it's bad, and being right feels good.
> People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book.
Please pardon the tangent: that's what always bothered me about the Borg in Star Trek. Why do they need to assimilate whole species? I'm sure there are enough volunteers in the federation that would join the Borg collective. Even a handful should be enough.
So then they're gone :)
When the borg is threatened ("threatened") by a race, they choose to effectively end their way of life one way or the other. That is done by assimilation or by death.
If only a single outstanding member of a species becomes a threat, that's a strong signal for their overall potential; therefore everything needs to go, the sooner the better.
The crux of the issue is that it is mostly a PR problem. AI is amazing but being promoted, in the eyes of many, by the worst people imaginable. Very akin, and overlapping in many ways to the crypto crowd.
What if people like food more than AI? Have you considered that?
Like, the current thing in the news cycle is a poll that young people are now more worried than hopeful about AI. Which sounds scary, but my first thought is that one could find similar polls from the 80s and 90s about satanic cults or alien abduction..
There are were polls about people being "more worried than hopeful" about satanic cults or alien abductions? With the youth being the most worried about satanic cults? Just like that knee-jerk "it's the bigger bubble", that makes zero sense.
As per the GDC 2026 State of the Industry poll, "52% said gen AI is bad for the industry, nearly double the 30% who held that view last year". But sure, everybody but HN, LinkedIn, and X bros are just luddites clutching pearls in their tiny bubble. They're the weak and stupid ones, and that is why the stupid shit said about them, day in and day out, isn't actually stupid. It all checks out.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
And I guess something like encyclopaedia would be good example of book they scan. At one point popular, but with most copies destroyed as no one actually wants them anymore.
I have found that often LLMs are familiar with the contents of SF books.
But I feel poorer, almost deprived, by the fact that all the LLMs I've checked with have NOT been trained on the contents of Eon by Greg Bear.
I'm sure they have some repository available somewhere. They can even sell the digital copies down the line if they're done with them (only once, of course).
Not in the EU, UK, or US. "AI" companies were already forced to settle their piracy cases, but often they get a free pass by law enforcement via regulatory capture.
The problem is a book author contracted publisher does not assign legal rights of duplication to a company/individual that buys a legitimate print. It can take over 70 years in some places to become public domain.
The core issue is "AI" firms have so much borrowed cash around, that getting a $1.5B fine for being a pirate is taken as a cost of doing business. The law is simply not equipped to handle this type of hyper-scaling criminal act. =3
https://www.bbc.co.uk/news/articles/c5y4jpg922qo
Anthropic's crime wasn't stealing the contents of books and making a derivative work of it, but torrenting a shitload of books. Had they bought all the ebooks, I don't think the lawsuit would've gone anywhere.
That is not how copyright/trademark/contract laws treat similar works. Most LLM know about Disney Mickey Mouse, and LLM vector search space proximity results will gravitate more accurate reproductions of protected works regardless of granularity of data.
OpenAI simply canceled a popular service to avoid Disney wrath.
https://www.theglobeandmail.com/world/article-openai-sora-di...
Also, trying to escape directly ripping off notable famous people with nonunion talent:
https://www.youtube.com/watch?v=YhgYMH6n004
I would say the "AI" firms will keep buying time with all that borrowed cash. Everything that could be scraped has already been stolen, and thus the problem should begin to self-correct. The Shrek movie release market correction history correlation is funny, and a new film is due 2027 in July. =3
https://annas-archive.gl/faq
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
https://www.techbrew.com/stories/2026/01/28/anthropic-ai-boo...
-- Thos. Jefferson
A partial solution there is of course a larger budget and a "last copy" policy where the last copy of a text at least is stored away in deep storage against a future loan.
The local libraries also accept book donations for an annual fund-raising sale.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Biggest problems:
I work in this space. I now require that before scanning a video is made carefully flicking through every page of the item so it can be checked after scanning to ensure all the pages are present and in the original order.Even the big libraries fuck up. I wanted an intact copy of Harper's Weekly from 1900 that has a big fold-out map in it. It's not clear to the libraries scanning this issue that the map is missing from their copies. None of the copies for sale from dealers have the map. Even when it is still glued into the middle it gets missed by industrial scanners. Google's scan only includes the (blank) back of the folded map.
Luckily GPT was able to track down a copy in a university special collections and fired off an email asking them to scan it. I just got the scan today:
https://imgur.com/a/vgMkM7b
(preview size, they sent a 500MB TIFF)
Now I can reassemble the issue and upload it.
I spend a lot of tokens getting LLMs vision tools to find the missing pages in vintage items and then try to reassemble them from other scans where available.
I'm also splitting up volumes to reupload. A lot of periodicals are only available online as giant multi-gig volume PDFs with all the issues in one file. I have a separate app I wrote to scan all the pages looking for covers so they can be split into PDFs and then identifying the volume/issue/month/year data from the cover or title page.
It's that simple, these corps have once again broken the social contract, you must not reward them. OpenAI and Anthropic especially, both owned by schizo sociopathic elites. Just use Chinese open models on 3rd party providers or more ethical companies.
This is literally the only power you have outside of Luigi, you're not going to fix anything with a letter writing campaign. We are entering a fight for survival so you really need to step up your game and stop letting elites run over you.
Today it's just books and manipulating society, tomorrow they will track and punish your behaviour and the control will only get worse. These people are pure fucking evil and we need to start acting like it while we literally still have the freedom and privacy to organise.
More like the consequence of being unwilling to change stupid laws once the stupidity of them is discovered. Nope. Gotta double down on the stupidity instead...
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
And this is supposed to be concerning?