Some argue that bots should be entitled to ingest any content they see, because people can.

Avram Piltch is the editor in chief of Tom’s Hardware, and he’s written a thoroughly researched article breaking down the promises and failures of LLM AIs.

RickRussell_CA
creator
link
fedilink
English
31
edit-2
1Y

Two things:

  1. Many of these LLMs – perhaps all of them – have been trained on datasets that include books that were absolutely NOT released into the public domain.

  2. Ethically, we would ask any author who parrots the work of others to provide citations to original references. That rarely happens with AI language models, and if they do provide citations, they often do it wrong.

I’m sick and tired of this “parrots the works of others” narrative. Here’s a challenge for you: go to https://huggingface.co/chat/, input some prompt (for example, “Write a three paragraphs scene about Jason and Carol playing hide and seek with some other kids. Jason gets injured, and Carol has to help him.”). And when you get the response, try to find the author that it “parroted”. You won’t be able to - because it wouldn’t just reproduce someone else’s already made scene. It’ll mesh maaany things from all over the training data in such a way that none of them will be even remotely recognizable.

RickRussell_CA
creator
link
fedilink
English
171Y

And yet, we know that the work is mechanically derivative.

@lily33@lemm.ee
link
fedilink
6
edit-2
1Y

From Wikipedia, “a derivative work is an expressive creation that includes major copyrightable elements of a first, previously created original work”.

You can probably can the output of an LLM ‘derived’, in the same way that if I counted the number of 'Q’s in Harry Potter the result derived from Rowling’s work.

But it’s not ‘derivative’.

Technically it’s possible for an LLM to output a derivative work if you prompt it to do so. But most of its outputs aren’t.

RickRussell_CA
creator
link
fedilink
English
41Y

a derivative work is an expressive creation that includes major copyrightable elements of a first, previously created original work

What was fed into the algorithm? A human decided which major copyrighted elements of previously created original work would seed the algorithm. That’s how we know it’s derivative.

If I take somebody’s copyrighted artwork, and apply Photoshop filters that change the color of every single pixel, have I made an expressive creation that does not include copyrightable elements of a previously created original work? The courts have said “no”, and I think the burden is on AI proponents to show how they fed copyrighted work into an mechanical algorithm, and produced a new expressive creation free of copyrightable elements.

@lily33@lemm.ee
link
fedilink
4
edit-2
1Y

I think the test for “free of copyrightable elements” is pretty simple - can you look at the new creation and recognize any copyrightable elements in it? The process by which it was created doesn’t matter. Maybe I made this post entirely by copy-pasting phrases from other people, who knows (well, I didn’t, only because it would be too much work), but it does not infringe either way…

keegomatic
link
fedilink
19
edit-2
1Y

So is your comment. And mine. What do you think our brains do? Magic?

edit: This may sound inflammatory but I mean no offense

RickRussell_CA
creator
link
fedilink
English
31Y

No, I get it. I’m not really arguing that what separates humans from machines is “libertarian free will” or some such.

But we can properly argue that LLM output is derivative because we know it’s derivative, because we designed it. As humans, we have the privilege of recognizing transformative human creativity in our laws as a separate entity from derivative algorithmic output.

conciselyverbose
link
fedilink
11
edit-2
1Y

So is literally every human work in the last 1000 years in every context.

Nothing is “original”. It’s all derivative. Feeding copyrighted work into an algorithm does not in any way violate any copyright law, and anyone telling you otherwise is a liar and a piece of shit. There is no valid interpretation anywhere close.

Every human work isn’t mechanically derivative. The entire point of the article is that the way LLMs learn and create derivative text isn’t equivalent to the way humans do the same thing.

It’s complete and utter nonsense and they’re bad people for writing it. The complexity of the AI does not matter and if it did, they’re setting themselves up to lose again in the very near future when companies make shit arbitrarily complex to meet their unhinged fake definitions.

But none of it matters because literally no part of this in any way violates copyright law. Processing data is not and does not in any way resemble copyright infringement.

RickRussell_CA
creator
link
fedilink
English
31Y

This issue is easily resolved. Create the AI that produces useful output without using copyrighted works, and we don’t have a problem.

If you take the copyrighted work out of the input training set, and the algorithm can no longer produce the output, then I’m confident saying that the output was derived from the inputs.

There is literally not one single piece of art that is not derived from prior art in the past thousand years. There is no theoretical possibility for any human exposed to human culture to make a work that is not derived from prior work. It can’t be done.

Derivative work is not copyright infringement. Straight up copying someone else’s work directly and distributing that is.

RickRussell_CA
creator
link
fedilink
English
31Y

There is literally not one single piece of art that is not derived from prior art in the past thousand years.

This is false. Somebody who looks at a landscape, for example, and renders that scene in visual media is not deriving anything important from prior art. Taking a video of a cat is an original creation. This kind of creation happens every day.

Their output may seem similar to prior art, perhaps their methods were developed previously. But the inputs are original and clean. They’re not using some existing art as the sole inputs.

AI only uses existing art as sole inputs. This is a crucial distinction. I would have no problem at all with AI that worked exclusively from verified public domain/copyright not enforced and original inputs, although I don’t know if I’d consider the outputs themselves to be copyrightable (as that is a right attached to a human author).

Straight up copying someone else’s work directly

And that’s what the training set is. Verbatim copies, often including copyrighted works.

That’s ultimately the question that we’re faced with. If there is no useful output without the copyrighted inputs, how can the output be non-infringing? Copyright defines transformative work as the product of human creativity, so we have to make some decisions about AI.

Well, I think that these models learn in a way similar to humans as in it’s basically impossible to tell where parts of the model came from. And as such the copyright claims are ridiculous. We need less copyright, not more. But, on the other hand, LLMs are not humans, they are tools created by and owned by corporations and I hate to see them profiting off of other people’s work without proper compensation.

I am fine with public domain models being trained on anything and being used for noncommercial purposes without being taken down by copyright claims.

RickRussell_CA
creator
link
fedilink
English
11Y

it’s basically impossible to tell where parts of the model came from

AIs are deterministic.

  1. Train the AI on data without the copyrighted work.

  2. Train the same AI on data with the copyrighted work.

  3. Ask the two instances the same question.

  4. The difference is the contribution of the copyrighted work.

There may be larger questions of precisely how an AI produces one answer when trained with a copyrighted work, and another answer when not trained with the copyrighted work. But we know why the answers are different, and we can show precisely what contribution the copyrighted work makes to the response to any prompt, just by running the AI twice.

Is there a meaningful difference between reproducing the work and giving a summary? Because I’ll absolutely be using AI to filter all the editorial garbage out of news, setup and trained myself to surface what is meaningful to me stripped of all advertising, sponsorships, and detectable bias

Tarte
link
fedilink
5
edit-2
1Y

I have yet to find an LLM that can summarize a text without errors. I already mentioned this in another post a few days back, but Google‘s new search preview is driving me mad with all the hidden factual errors. They make me click only to realize that the LLM told me what I wanted to find, not what is there (wrong names, wrong dates, etc.).

I greatly prefer the old excerpt summaries over the new imaginary ones (they‘re currently A/B testing).

RickRussell_CA
creator
link
fedilink
English
101Y

When you figure out how to train an AI without bias, let us know.

You’re confusing ai with chatgpt, but to answer your question: if it’s my own bias, why would I care that it’s in my personal ai? That’s kind of the point: using my personal lens (bias) to determine what info I would be interested in being alerted of

RaleighEnt
link
fedilink
61Y

oooh I dunno man having an AI feed you shit based on what fits your personal biases is basically what social media already does and I do not think that’s something we need more of.

You’re confusing ai with chatgpt

???

Create a post

A nice place to discuss rumors, happenings, innovations, and challenges in the technology sphere. We also welcome discussions on the intersections of technology and society. If it’s technological news or discussion of technology, it probably belongs here.

Remember the overriding ethos on Beehaw: Be(e) Nice. Each user you encounter here is a person, and should be treated with kindness (even if they’re wrong, or use a Linux distro you don’t like). Personal attacks will not be tolerated.

Subcommunities on Beehaw:


This community’s icon was made by Aaron Schneider, under the CC-BY-NC-SA 4.0 license.

  • 1 user online
  • 144 users / day
  • 275 users / week
  • 709 users / month
  • 2.87K users / 6 months
  • 1 subscriber
  • 3.09K Posts
  • 64.9K Comments
  • Modlog