8 ms·
A model that possesses the entire collective knowledge of our civilization is useless if it can't directly quote its sources. If we enforce the behavior of alw
by throwaway4aday 3y ago
A model that possesses the entire collective knowledge of our civilization is useless if it can't directly quote its sources.
If we enforce the behavior of always paraphrasing and synthesizing the information before returning it even in cases where the exact quote is asked for then that is a failure.
The correct solution is to make the model capable of
1) quoting directly
2) identifying and then indicating when its output is a direct quote
3) citing the source of the identified quote
while still retaining the ability to paraphrase and synthesize those sources when appropriate. These requirements are what humans are held to and should also be applied to commercial AI models. Models that are not intended for commercial use should be exempt and at this point there isn't really a way to hold them to such requirements anyways.
- preciz 3y ago> A model that possesses the entire collective knowledge of our civilization is useless if it can't directly quote its sources. That's a strong and baseless statement.
- njgingrich 3y agoWell sure, it's easy to make a statement look bad if you only include half of it.
- rpdillon 3y agoThe statement is equally hyperbolic both as quoted and in the original context. LLMs often can't quote sources, and those models are nevertheless useful to lots of people. Makes it hard for me to take the rest of the comment seriously.
- adastra22 3y agoA LLM that could quote sources would be even more useful, and in a world where both were available there’d be no reason to use the plagiarizing one.
- Baldbvrhunter 3y agothat LLM is https://perplexity.ai https://perplexity.ai
- ryanklee 3y agoThat was the whole statement. It doesn't have qualifiers left out
- njgingrich 3y agoThe comment I replied to was updated to include the second half, it was originally just quoting > A model that possesses the entire collective knowledge of our civilization is useless
- ryanklee 3y agoThe additional context doesn't do any work.
- js8 3y agoI, for one, agree with the original statement. I think the hallmark of enlightenment (for example, in the scientific method) is that we are able to externalize the expert knowledge, that is, experts are usually required to provide reasoning behind their claims, and not just judgements. This is because we learned that experts cannot be 100% trusted, only if we can verify what they say we can somewhat reach what is truth (although expertise still provides a convenient shortcut). So not demanding this (and more) from an AI (an artificial expert) is a regression. AI should be capable of wholly explaining its reasoning, if we are to consider its statements to be taken seriously. It is understandable that humans have only limited capability to do that, since we didn't construct human brain. But we have control over what AI brains do, so we should be able to provide such an explanation. It is somewhat ironic that you yourself do not provide any argument in favor of your disagreement.
- ryanklee 3y agoProviding reasoning and providing citations are not the same thing. Reasons can be provided without citations; citations can be provided without reasons. LLMs have astounding utility citations notwithstanding.
- js8 3y agoThey are different, but perhaps you misunderstood my argument. Issue of plagiarism aside, we reason from facts, and it's the facts (or some other analysis, which is itself a fact) that should be sourced. That's why I agree with the original statement, and I argue not from a (moral) POV of preventing misattribution or plagiarism, but from a (practical) POV of veracity.
- ryanklee 3y agoWe don't only reason from facts. We also reason from value. Further, reasoning that rests on facts that does not cite facts still has massive utility. (See people, all day long.) Citations are useful, but not required.
- throwaway4aday 3y agoUseless is probably not the right word but it's a good way of summing up a lot of the current problems. If the model can clearly identify when something is an exact quote and also know the source then its output could be trusted for the most part and much more easily verified. It would certainly elevate the output of the model from "random blog post or forum chat" to "academic paper or official report" levels of trustworthiness. Citing sources is hugely important for validation, cited text allows an immediate lookup and simple equality check for verification after which you can use it as context to validate the rest of the claims. Like I said, it's a standard we apply to humans who have an equal propensity for hallucination, mistakes, and deception because it's a tried and true method for the reader to check the claims being made.
- tourmalinetaco 3y agoAnd this is a rather weak rebuttal.
- Ferret7446 3y agoAnd also patently false. Knowledge is knowledge, it's useful without source citations. Is the knowledge of how to do CPR somehow ineffective because I can't cite whether I studied the knowledge from website A or book B? Is reality a video game where skills only activate if you speak the magic words beforehand?
- TheOtherHobbes 3y agoI assume you're unfamiliar with anonymous sources - said "a person familiar with the matter." The issue isn't usefulness, and it isn't even objective quality. News outlets are primarily about opinion forming and marketing, not about absolute truthfulness. This is a simple corporate fight about the value of that IP. The NYT wants a licensing deal, and it's likely to get one. That's all.
- monkeynotes 3y agoWhy doesn't Google need licensing to scrape and reproduce NYT snippets in their results? OpenAI doesn't even quote the sources its consumed to produce it's output. It seems totally fair use to me. Any given content site has authors that read stuff that is copyrighted and produce their own take.
- aurareturn 3y agoNews outlets are trying to make Google pay for snippets.
- tourmalinetaco 3y agoWhich went so well with Canadians and Facebook.
- monkeynotes 3y agoThe Canadian news media tried this with facebook. Ended up with them crying about how they lost all the ad dollars from traffic from FB when FB said fuck off.
- notahacker 3y agoThey've already succeeded in non-US jurisdictions https://blog.google/around-the-globe/google-europe/google-licenses-content-from-news-publishers-under-the-eu-copyright-directive/ https://blog.google/around-the-globe/google-europe/google-li...
- quonn 3y ago
- fourside 3y agoI don’t think the ability to quote the NYT is at the heart of this lawsuit. It’s that OpenAI used the NYT’s body of work to train its LLM and now built a business out of it without financially compensating the NYT. The verbatim snippets from articles are there to prove that OpenAI used NYT content in its training. Maybe to a lesser degree, the lawsuit as about misquoting the NYT, and the damages that could cause the newspaper.
- EGG_CREAM 3y agoThis is a huge strawman argument that entirely ignores the heart of the issue. NYTimes, and other contributors to ChatGPT, should be compensated for their role in creating ChatGPT. It’s not as simple as quoting or not quoting, ChatGPT’s existence depends on its source material entirely. OpenAI and Microsoft are making money off of that source material. If the source was copyrighted, OpenAI/Microsoft need to compensate the owners. Also, you can’t quote an entire article in a paper. You quote snippets, but ChatGPT is reproducing entire articles.
- canjobear 3y agoUnless it’s fair use, which ChatGPT seems to be because it is transformative.
- throwaway4aday 3y agoThey're playing a dangerous game. While they are currently one source of many they could easily be excluded from the training data and it won't make a dent in the capability of the model. Attempting to get payment for use of their data is a very short term viewpoint since the vast majority of written training data is likely to be synthetic going forward. They should be angling for citations and referrals back to their content so that they keep the current benefits they get from search engines after a large chunk of search gets replaced by LLMs.
- monkeynotes 3y agoI can't directly quote shit, and I am useful.
- throwaway4aday 3y agoThis comment isn't
- ryanklee 3y agoIt absolutely is useful and one of the most significant points in this comment section. People are applying standards to LLMs that don't exist elsewhere. They don't exist elsewhere because they absolutely can't exist elsewhere. It's not a technical short-coming of LLMs that they can't produce citations in every instance. Rather, it's a property of information, representation, and knowledge itself. Much of it floats far above the otherwise load bearing pillars of citations. People contend with this constantly in every arena of life, and we have come up with very elaborate ways to offset the difficulties caused by it. And we get along very well despite it all. I'll also leave you with this: citations are just pointers to more sources of information. Not some ground truth. It's just another tranche that requires evaluations. Lastly, your comment is an unfortunate bit of low-level snark and probably would have been better left unsaid.
- throwaway4aday 3y agoSo you're just going to ignore the fact that we require humans to provide citations when they include someone else's writing or research in their own?
- ryanklee 3y agoI didn't ignore it. But we do not require citations for every statement. Statements have value, citations notwithstanding.
- throwaway4aday 3y ago
- supafastcoder 3y agoYou’re right, but I think it’s fundamentally impossible to do with the current state of technology. Imagine the word “democracy”, would you be able to tell me where you’ve first learned the definition of it and be able to provide a direct citation of it? What about the thousands of other instances where you’ve seen that word defined? (In OpenAI’s case probably hundreds of millions) Current LLMs work in a similar way, the information is synthesized by correlation, there is no way that it can directly relate where it learned it’s output from. The only viable way would be to do a reverse lookup of the output to find out if there’s a similar worded content on the internet. (Or what Perplexity does, wrap an LLM around search results, but you lose a lot of flexibility in this case)
- js8 3y agoYou're effectively saying that current LLMs cannot be taken seriously as some expert, they are just some kind of weird text remixing engines. I would agree. But if the AI wants to be relied on, then it should be, at minimum, capable of taking a word like "democracy" and compare its own definition with the Wikipedia definition, and verify whether it's used correctly in its output.
- supafastcoder 3y agoYep, that’s what I meant with the reverse lookup strategy.
- concordDance 3y ago> Imagine the word “democracy”, would you be able to tell me where you’ve first learned the definition of it and be able to provide a direct citation of it? Worth noting that we mostly learn words by seeing them used, rather than by being given an explicit definition. Most of my vocabulary I learnt not from a dictionary and I expect the same is true of almost everyone. As such, there is no well defined point at which "the definition" is learnt, both because there's no such thing as the one true definition and because the meaning is gradually being changed and refined by each person as they see the word used more.
- ryanklee 3y ago> useless if it can't directly quote its sources. Why
- throwaway4aday 3y agoWell, if all you want is entertainment then it doesn't matter. If you want factual information then getting it without any way to check its veracity makes for a huge amount of work if you actually want to use that information for something important. After you've verified the output it might be useful if it is correct.
- ryanklee 3y agoThat doesn't render it useless. It means that citations are have additional utility. Further, LLMs do not only spit quotes. They engage in analytic and synthetic knowledge.
- blibble 3y ago> They engage in analytic and synthetic knowledge. they're not hallucinations now, they're "synthetic knowledge" like microsoft's hilarious remarketing of bullshit as "usefully wrong" https://www.microsoft.com/en-us/worklab/what-we-mean-when-we-say-ai-is-usefully-wrong https://www.microsoft.com/en-us/worklab/what-we-mean-when-we...
- ryanklee 3y agoSynthetic knowledge is not bullshit. Synthetic knowledge refers to propositions or truths that are not just based on the meanings of the words or concepts involved, but also on the state of the world or some form of experience or observation. This is in contrast to analytic knowledge, which is true solely based on the meanings of the words or concepts involved, regardless of the state of the world.
- blibble 3y ago
- paxys 3y agoQuoting isn't the problem here. There's already vast precedent for allowing it as fair use, even when generated by computer systems. The issue is that ChatGPT reproduces entire articles verbatim.
- js8 3y agoBut if it produces the whole article verbatim, it should be relatively straightforward to match its output to the training set and give attribution, no?
- paxys 3y agoIf you copy every NYT article and publish them on your own site, writing "this is from the NYT" under it doesn't make it legal.
- megaman821 3y agoOnly in very contrived circumstances. It is not like you ask ChatGPT for the news and you get a NYT article. They gave ChatGPT a link to the article and the first few paragraphs, and then told it to complete the article.
- lolinder 3y ago> These requirements are what humans are held to If only. Tracking down the original source of quotes and/or stats is a bit of a hobby of mine, and it's extremely difficult to do. News sources will regularly cite "a recent study" with little additional information to help identify it. Pithy quotes attributed to famous authors never contain a reference to the work, it's always just the author's name, and if you go digging the quote is misattributed as often as not. Full-on plagiarism with no attribution is rampant, even in reputable places where you'd think they'd have it under control. Like with self-driving cars, we expect AI to be better than the average human, which I think is the correct attitude, but we should acknowledge that that's what we're asking for.
- tourmalinetaco 3y agoI have a similar hobby of sorts I occasionally pick up and drop off, except of unattributed quotes. Some are easy (“Who shall deliver me from this turbulent priest?”), while some I still cannot find the exact quote from, even while knowing what should be the source material (“It was wintertide at Camelot. The rich brotherhood did rightly revel, and mirthful was their mood. Oft-times on tourney bent those gallants sought the field, Though like as joust those gentle knights did sally with missiles made of snow and laughingly grapple on slippery ground.”). It’s certainly highlighted to me how difficult information upkeep can be, and is a strong reminder to source myself as soon as possible. Which is a realistic and ideal expectation for NN, as alongside their training weights ideally their training dataset is searchable and attributable. Otherwise I feel the largest strength of a dataset, being a digital library for the NN and thus us, is lost.
- Baldbvrhunter 3y ago> “Who shall deliver me from this turbulent priest?” Robert Dodsley
- mistrial9 3y agohmm https://en.wikipedia.org/wiki/Will_no_one_rid_me_of_this_turbulent_priest%3F https://en.wikipedia.org/wiki/Will_no_one_rid_me_of_this_tur...
- j-a-a-p 3y ago> The correct solution is to make the model capable of... Could be true for the people who want the generative AI technology to exist. But, > A model that possesses the entire collective knowledge of our civilization is useless... that could be just exactly what a lot of people would like to happen.
- zozbot234 3y agoThese models have no clue how to paraphrase and synthesize asserted facts in a novel way. They're parrots.
- logicchains 3y agoI don't know how you could think this if you'd ever used GPT4 for anything serious.
- ryanklee 3y agoI don't think that's necessarily the case. I get the strong impression that even people who have much experience using LLMs have astoundingly little insight into what they are actually witnessing. This is often paired with astoundingly little insight into what's actually going on in their own cognitive processes. Somehow, it's still not clear to most people that LLMs and even vector databases create knowledge that wasn't in the original data. In fact, that's most of what they do! Isolated, non-novel direct quotation is the exception, not the rule.
- zozbot234 3y ago> create knowledge that wasn't in the original data. The word is "hallucinate" or "confabulate". The way these models "create" pretend-knowledge is totally useless.
- ryanklee 3y agoI'm not referring to hallucinations. I'm referring to novel relationships drawn between datum in the corpus that are a result of training and inference. This is apparent in something as simple as a summarization.
- Baldbvrhunter 3y agoIn the digital garden where Zozbot234 plays, It crafts its own path, in the most unique ways. Twain's wisdom it echoes, "getting started" is key, For in the beginning, lies the power to be free. Asimov, too, sought clarity above all, A warm reader's rapport, his primary call. Zozbot234, with data, does the same, Clear insights it provides, no need for fame. Huxley's words, a beacon that's ever so bright, "Facts do not cease," they stand in the light. Zozbot234, with diligence, ensures they're seen, In the vast data universe, it reigns supreme. Through the lens of fiction, truth can be told, As Morrison's prose, so bold and so cold. Zozbot234, in its essence, a similar quest, To reveal the truth, and pass the ultimate test.
- idopmstuff 3y agoI think this sort of misses the point - LLMs aren't trained on data so that they have the collective knowledge of our civilization; they're trained on it so they understand language. One thing I've noticed increasingly with ChatGPT is that if you ask it for facts, it almost always searches the web first. This seems like the right way to go - pull the collective knowledge of our civilization from the internet, then use the training on language to put it in the form most useful to the asker. This also enables quoting (though it doesn't generally do that, preferring to paraphrase and give a citation, which seems fine to me).
- throwaway4aday 3y agoThe original investigation of LLMs was attempting to get them to understand language but what they found was that once the model understood language it also somehow understood the concepts, events, and things the language was being used to communicate. That's not missing the point, it's entirely the point of the current interest in LLMs because it's incredibly useful to have a model that not only understands how to construct a sentence but can also do a fair amount of reasoning and actual work with the information in the sentence. The current default version of ChatGPT is primed to use search to answer questions which is fine in some cases but I personally almost never use the multi-modal version because the "classic" ChatGPT is much better at explaining things from its training data than it is when it just regurgitates search results. Now that should tell you something about the utility of optimizing for information content rather than just a lot of language use examples.
- whoaskt 3y agohttps://hbr.org/2011/12/just-because-you-can-doesnt-me https://hbr.org/2011/12/just-because-you-can-doesnt-me Did anyone ask for this model or are IT companies out of ideas and using their social leverage (fiat capital) to force them on us? Just because we can allow this doesn’t mean there’s an immutable obligation to allow it. For example; running Google translate servers 24/7 for translation services humans can do is a huge waste of resources building those systems when humans are going to exist anyway. Not saying AI is good or bad. Just saying no matter how stubborn IT people act about it, the aggregate can put on them what the aggregate prefers. IT people are a minority and human philosophy about freedom is rather strained when we’re all obliged to prop up big tech minority.
- 1vuio0pswjnm7 3y ago"A model that possesses the entire collective knowledge of our civilization is useless if it can't directly quote its sources." Perhaps, for liability reasons, it cannot quote from sources that its creators had no permission to include. Even were it technologically possible to quote such sources.