This, somewhat alarmist, 25 July article by Frank Landymore, for instance, suggests that AI companies are purchasing antique books in large volumes (this is in the article title: "AI Companies Are Buying Antique Books…") to use as high-quality training data, as these texts are free from modern digital AI content. Because it is cheaper, destructive, high-speed automated scanning processes are being used to digitize the books (i.e., book spines are guillotined off, the covers removed, page-stacks are fed into scanner hoppers, then thrown away). The practice is facilitated by book database services that keep the tech company buyers anonymous. Rare booksellers have ethical concerns about selling books to these services, warning that it permanently destroys some of the last surviving copies of historic literature.
The FUD here is the "antique books" and "historic literature" claim. Once you read through the parent and grandparent linked articles (3 April essay by Val Giordano, promoting the practice, here; 25 June [anonymous] essay on the NL Times, reporting concerns of European book dealers, here; a 26 June article by Benj Edwards here; a solid, 21 July essay by Emanuel Maiberg here) it becomes clear that the books being hoovered up by Anthropic (and possibly others) all post-date 1969 (when ISBNs were introduced) and most, it seems likely, post-date 2006 (when ISBN barcodes became ubiquitous)—about the same time, coincidently, that the arrival of modern smart-phones ended civilisation as we knew it.
That the AI-company proxies are exclusively interested in books with an ISBN is made very clear from the following anecdote reported by Maiberg:
The seller told me that … Bulk purchases … usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases. (emphasis added)
That the bulk-purchase companies are likely to prefer later, bar-coded books, seems probable from the scale they are working on—it is simply easier and cheaper to scan the bar codes of hundreds of thousands of books, than it is to key in the ISBN (and then check for errors) on the same number.
From the articles I mention above, it seems that second-hand book dealers are getting what I take to be in-fill requests, for any ISBN that has not shown up in the literal truck-loads of books donated to charities, that are sold off to bulk-book dealers / waste-paper merchants (on this, see my post here concerning Mubin and Raza Ahmed's "Wrap Ltd."—a "printed matter," "waste and scrap paper" import/export business with an annual turnover of £6.5M "or more.")
It simply wouldn’t be economical to buy a book via ABE, Alibris etc. if you could buy the same book at a fraction of the price from
For me, this is where it gets interesting. Anthropic et al. have more money than they know what to do with, and have just been stung to the tune of USD3K per book, for all the digital books they stole, and then got sued for stealing. That being the case, the following observation should come as no surprise:
there's a total disregard for the price of the book. I've had some books that sold through this way that were [...] greatly overpriced. That's kind of a tell for AI because they have just so much money.
* * * * *
Since I was born before ISBNs were a thing, it is hard for me to think of books after 1970 as "antique" or "historic" ("rare," especially, "rare on market" is another matter completely). Perhaps, this is easier for Frank Landymore—who seems to be in his early 30s. In any event, I asked one of the water-guzzling, book-eating monsters (ChatGPT) to estimate how many books have been published since 1970 (and 2006) and the answer should be reassuring.
It seems likely that something like 80 million English-language ISBN editions have been released, or which approximately 60 million editions were printed with a barcode. The number of unique works (i.e., with Pride and Prejudice counting once) is obviously much smaller—perhaps 10–20 million, but the number of physical copies printed is much larger, likely "well over 100 billion copies since the 1980s."
So, it seems Anthropic is attempting to gobble up millions of recent books (emphatically-not antique you baby-faced ageists!), mostly comprised of unwanted donations, fenced by the likes of Messers Ahmed, and otherwise unsold obscuriana, laying heavily on the hands of dealers who once supplied libraries, out of a pool of something like 100 billion copies printed.
If it were possible to pile this 100-billion copies one atop the other, it seems that the pile would likely extend 2.5 million kilometres, or about 6.5 times the distance from the Earth to the Moon. Indeed, only 14.5 billion books would be need to reach the moon, or 14.5% of all books printed with an ISBN.
I would wager as many books as I could carry, that far, far more than this 14.5%—perhaps 80% of these 100 billion books—have already been pulped, without once raising the interest of Landymore et al., making this story a pretty good parallel to widely-disseminated concerns about the water consumption of data centres. (Which is mostly FUD too, since most new data centres consume no water—they use closed-loop cooling—and the water usage of older AI data centres that do use evaporative cooling etc., is dwarfed by the water consumption of golf-course and intensive cattle farms. I.e., this is a gnat on the arse of a flee in the grand scheme of things.)
And while most AI tokens will doubless go to things like generating fake security footage of President John F. Kennedy and Marilyn Monroe in flagrante delicto, or self-insert versions of Terminator or Twilight, it is already clear that AI can, and probably soon will, cure every disease known to man and put a complete stop to the death and mutilation caused by badly-driven cars on the road—to name but two of life’s miseries. If they need to guillotine a copy of each my books to do this (**), they are welcome to do so.
(**) if AI can also restore my use of AdobeCaslonExpert in the later versions of Word, they can have all my essays too.


No comments:
Post a Comment