> But nobody was ever going to that Didn't Google have a long standing project t...

godelski · 2025-09-06T07:49:40 1757144980

From TFA

  The Google Books project also faced a copyright lawsuit, which was eventually decided in favor of Google.

  After contacting major publishers about possibly licensing their books, [former head of the Google Books project] bought physical books in bulk from distributors and retailers, according to court documents. He then hired outside organizations to dissemble the books, scan them and create digital copies that could be used to train the company’s AI. technologies.

  Judge Alsup ruled that this approach was fair use under the law. But he also found the company’s previous approach — downloading and storing books from shadow libraries like Library Genesis and Pirate Library Mirror — was illegal.

rchaud · 2025-09-06T17:37:01 1757180221

That wasn't done as a play for venture capital. The Google Books project began before eBooks existed; in the 2000s, they spent money on all kinds of projects that had no real strategy for monetization. I remember Google Books being a valuable resource as it digitized books that were out of print. Back when they actually cared about making information available widely.

Thorrez · 2025-09-07T05:55:23 1757224523

Yeah. Weird that rchaud said "But nobody was ever going to that" when the article talks about someone doing it.

imwm · 2025-09-06T17:31:58 1757179918

Disassemble*

miohtama · 2025-09-06T06:23:14 1757139794

This lawsuit also makes sure that only parties that can train an AI with good enough training material are now

- Google

- Anthropic

- Any Chinese company who do not care about copyright laws

What is the cost of buying and scanning books?

Copyright law needs to be fixed and its ridiculous hundred years tenure chopped away.

godelski · 2025-09-06T07:50:53 1757145053

From TFA

  > Anthropic also agreed to delete the pirated works it downloaded and stored.

Also

  > As part of the settlement, Anthropic said that it did not use any pirated works to build A.I. technologies that were publicly released.

Iolaum · 2025-09-06T08:36:03 1757147763

Reminds me when Facebook said to EU that they did not have the technology to merge FB and Whatsapp accounts when they bought Whatapp.

kelnos · 2025-09-06T10:44:26 1757155466

That's not really the point, though, is it? Now Anthropic can afford to buy books and get them scanned. They likely didn't have the money or time to do that before.

And even if they didn't use the illegally-obtained work to train any of the models they released, of course they used them to train unreleased prototypes and to make progress at improving their models and training methods.

By engaging in illegal activity, they advanced their business faster and more cheaply than they otherwise would have been able to. With this settlement, other new AI companies will see it on the record that they could face penalties if they do this, and will have to go the slower, more expensive route -- if they can even afford to do so.

It might not make it impossible, but it makes the moat around the current incumbents just that much wider.

DrillShopper · 2025-09-06T21:28:40 1757194120

> As part of the settlement, Anthropic said that it did not use any pirated works to build A.I. technologies that were publicly released.

Oh so now we're at "just trust me bro" levels of absurdity

slow_typist · 2025-09-06T07:40:43 1757144443

Training a Model on 100+ years old literature only could be an interesting experience though.

IshKebab · 2025-09-06T08:53:49 1757148829

It's been done.

https://github.com/haykgrigo3/TimeCapsuleLLM

rollcat · 2025-09-06T08:26:45 1757147205

’Twould wax yet more marvellous to ye beholders.

efskap · 2025-09-06T02:55:17 1757127317

Crazy to think we've been helping train AI through captchas long before the "click all squares containing" ones.

a2128 · 2025-09-06T05:03:51 1757135031

"stop spam. read books." is a very ironic phrase to look back on considering the amount of spam on the internet that LLMs have enabled