4 ms·
From what I can tell, this has nothing to do with LLMs at all. In the example in the article, the user is asking Bing to go fetch the contents of an article dir
by px43 3y ago
From what I can tell, this has nothing to do with LLMs at all. In the example in the article, the user is asking Bing to go fetch the contents of an article directly from the website, and print it out, which it dutifully does.
Seems like the "problem" is that NYT etc gives privileged access to search engines for indexing their content, but then get upset when snippets of the indexed content is being shown to users without the users having to fight the paywall or whatever.
This article also claims that the screenshot is coming from ChatGPT when it clearly is not.
- rich_sasha 3y agoI suppose that's a relatively easy thing to fix, technically. It proves, however, that th underlying LLM is trained on copyrighted data. I'm not sure the problem goes away simply if the LLM in question (or any other one) gets some "no verbose regurgitation" filter.
- exitb 3y agoIn that case, the language model calls a search function and just repeats the result out its conversation context, not its training data. With that in mind it's not clear why it's ok for Bing itself to quote the source, but it stops being ok, when a chatbot does it.
- kolinko 3y agoThe example from the article doesn't show that LLM is trained on copyrighted data - it's just Bing fetching the source article, providing it to GPT, and GPT rephrasing the article. An agent trained on entirely copyright-free data would provide exactly the same output.