• bignose@programming.dev
    link
    fedilink
    English
    arrow-up
    9
    arrow-down
    2
    ·
    5 days ago

    An LLM, being a large database derived mechanically from its training data, seems to meet the definition of “derived work” as exists in most copyright legislation.

    The output from an LLM, through the inference process, is also mechanically derived from that training data.

    Lawmakers might try writing exceptions to please a stock market absolutely baying for hyperscaler growth. Despite that, it seems entirely viable that the copyright, held by humans on the training data fed to the LLM, continue to hold in the output derived from that input.

    People saying it’s “obvious the output has no copyright”, or people who act as though it’s all free from any claim by the humans who wrote the works used as training data, are just showing their ignorance, it seems to me.

    • encelado748@feddit.org
      link
      fedilink
      arrow-up
      5
      arrow-down
      1
      ·
      4 days ago

      People learn from books and existing code. Does that means that anything we write is derived work?

      An LLM knowledge matrix is a neural network, not a database. It does not contains an exact copy of any given text. It may remember it, but human brain can do the same.

      You have to prove LLM code is derived work in the same way you prove human work is derived work: check the difference between the original work and the derived work.

      • bignose@programming.dev
        link
        fedilink
        English
        arrow-up
        1
        ·
        1 day ago

        Copyright grounds its justification in the creative expression of people. By creative transformation, new works can be created and some of those are deemed not restricted by copyright in the earlier work.

        An LLM is not a person. An LLM has no mind and no creative expression. The training of an LLM and the inference from the LLM are, by definition, non-creative mechanical transformation from the input training data.

        The copyright of the input work should, by this logic, survive whole in the output from that mechanical non-creative process. That is substantially different from the process by which humans learn and create.

        • encelado748@feddit.org
          link
          fedilink
          arrow-up
          1
          ·
          1 day ago

          You can make the case that an LLM is not a person, you cannot make the case that LLM work is not creative. The grounding justification of copyright lacked a foundational example of a non human creative process. If each single token emitted by the LLM is the result of math on the entire knowledge matrix, then each single token is derived from all the copyrighted documents used in training in a percentage you cannot even estimate. The mechanical transformation of an LLM is unknowable, unquantifiable and non deterministic. The copyright claim makes no sense in this case. A new shared mechanism must be put into place if we want to translate copyright through such mechanism. Lot of the training data used by LLM is synthetic so it pass through multiple layers of unknowable, unquantifiable, non deterministic “mechanisms”.

      • bignose@programming.dev
        link
        fedilink
        English
        arrow-up
        2
        ·
        1 day ago

        What creative transformation has that company done? It seems to me rather that the LLM is an entirely mechanical process with no creativity.

        No more creative than what my ISP does when it conveys data from one machine to another.

        We don’t allow the ISP to claim copyright in the data that it mechanically transforms. I don’t see why the company running an LLM would have any stronger claim of copyright in the output.

        • MonkderVierte@lemmy.zip
          link
          fedilink
          arrow-up
          1
          ·
          4 days ago

          No, i don’t think so, i said “if so”. A piece of code can’t own something, unlike a person. Would be a different matter, if the piece of code was a person, but alas.

  • Dejected Warp Core@lemmy.world
    link
    fedilink
    arrow-up
    3
    arrow-down
    4
    ·
    5 days ago

    Ethics aside yes… Right up until an llm directly quotes something from its training set. Then all hell breaks loose. Good luck figuring out when that happens.

    The rest of the time, I suspect most models act like “distinct from copyrighted material generation engines” and might even enjoy parody rights. That said, I’m not a lawyer and this is not legal advice.