Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.

  • st0v
    cake
    link
    fedilink
    English
    39 months ago

    I have to assume that openAI also paid for the books. if yes then i consider it the same as me reciting passages from memory or coming up with derivative text.

    if no, then by all means, go after them and any model trainer for the cost of one book.

    Asking an LLM to recite an entire novel isn’t even vaguely a thing yet.

    • @[email protected]
      link
      fedilink
      English
      3
      edit-2
      9 months ago

      Well, here’s straight from one of the suits against them:

      “The OpenAI Books2 dataset can be estimated to contain about 294,000 titles. The only ‘internet-based books corpora’ that have ever offered that much material are notorious ‘shadow library’ websites like Library Genesis (aka LibGen), Z-Library (aka B-ok), Sci-Hub, and Bibliotik. The books aggregated by these websites have also been available in bulk via torrent systems.”

      I’m not even sure how they would have logistically gone about purchasing 294,000 books in bulk in digital form to be fed into training. Using the existing collections seems much more likely, but I suppose we’ll see what turns up in litigation.

      Also, the penalty for downloading copyrighted material if willful infringement is up to $250,000 per work. So it’s quite a bit more than the cost of one book on the line…