3 ms·
There's a few options, OpenAI's list of legal risks is extensive and horrifying. * For the data sources, we know there's a lot of CSAM in there, and that OpenA
by ADeerAppeared 2y ago
There's a few options, OpenAI's list of legal risks is extensive and horrifying.
* For the data sources, we know there's a lot of CSAM in there, and that OpenAI knowingly shipped this data to 3rd party companies to tag it. Time ran an article about their Kenyan parters, who quit on them because "Look at this image. If the image is CSAM, push the button" is a horrifying job.
* There's copyright. AI relies extensively on hiding the true scale of the copyright problem by filtering infringing content out of the AI responses. But the engineers of those systems know the dirty details. (It's dubious that such filtering would bring them back into compliance with fair use; LLMs are paraphrasing machines and copyright filters do not cope well with paraphrasing)
* Discrimination is a shitshow. LLMs discriminate as their dataset isn't reflective of reality, but of what and how we choose to record reality. And as Google's "Diverse Nazis" image generation shows, messing with prompts doesn't work to fix this. It'll be discrimination lawsuits either way. All major AI firms know of this problem, and internally investigate it to avoid gaffes like Google's.
In practice, the boring answer to the question is "all of the above and then some". A big problem for OpenAI is that these are massive problems, for which their engineers could be dragged into a courtroom to testify about.