2 ms·
I've seen: 1. AI models frequently output large chunks of code which are plagiarised. In one specific case it was code for walking the stack, that was a clear
by 20k 2mo ago
I've seen:
1. AI models frequently output large chunks of code which are plagiarised. In one specific case it was code for walking the stack, that was a clear mix of two original sources that I was able to find with changed variable names, but the structure was identical and switched from the first to the second halfway through
2. AI models plagiarising stack overflow answers word for word, quite recently about the rotation rate of smoothbore cannons in the age of sail
3. AI misspelling answers because the physics papers its trained on made the same typos, which is how I discovered that it had plagiarised the answer
4. Misconceptions/wrong answers that can be traced back to specific papers due to the oddly specific nature of the language used
There's been a lot of research about getting AI models to output their training data, and it turns out they store huge amounts of it. You can use this to get people's personal information if you really want to, and that's very low occurance information
- seanmcdirmid 2mo agoI just want to make sure that you saw those things with pure context isolation and AI wasn’t just doing a search to grab the content directly and throw it into the context of your prompt. You’d be surprised how often I’ve heard these claims and it turns out they were just using an agent with RAG enabled.
- bonoboTP 2mo agoIf it never regurgitated the exact same thing, would you accept AI then?
- 20k 2mo agoFor me there's two separate problems: 1. The plagiarism aspect, and that most of the training data was used without permission 2. I haven't found it terribly useful in my personal work, as the data it was trained on was heavily polluted by incorrect information (at least in the field I'm using it)
- bonoboTP 2mo ago1. I said in my hypothetical it would not reproduce exact content. 2. Okay, others find it useful. So what?
- seanmcdirmid 2mo agoYou shouldn’t be relying on worlds knowledge accidentally captured in model weights, that’s a bug not a feature. Instead you should be stuffing relevant context into your prompt so it has the correct information. You obviously are really confused about how LLMs work and how they should be used.