5 ms·
"Why can't AI at least cite its source" each article seen alters the weights a tiny, non-human understandable amount. it doesn't have a source, unless you think
by make3 3y ago
"Why can't AI at least cite its source" each article seen alters the weights a tiny, non-human understandable amount. it doesn't have a source, unless you think of the whole humongous corpus that it is trained on
- aantix 3y agoWe're trying to solve AGI but can't solve sources/citations?
- deleted 3y ago[deleted]
- pxoe 3y agothat just sounds like "we didn't even try to build those systems in that way, and we're all out of ideas, so it basically will never work" which is really just a very, very common story with ai problems, be it sources/citations/licenses/usage tracking/etc., it's all just 'too complex if not impossible to solve', which just seems like a facade for intentionally ignoring those problems for benefit at this point. those problems definitely exist, why not try to solve them? because well...actually trying to solve them would entail having to use data properly and pay creators, and that'd just cut into bottom line. the point is free data use without having to pay, so why would they try to ruin that for themselves?
- simonw 3y agoWhat makes you think AI researchers (including the big labs like OpenAI and Anthropic) aren't trying to solve these problems?
- pxoe 3y agothe solutions haven't arrived. neither have changes in lieu of having solutions. "trying" isn't an actual, present, functional change. and it just gets passed around as an excuse for companies to keep doing whatever they're doing.
- pama 3y agoPlease recall how much the world changed in just the last year. What would be your expected timescale for the solution of this particular problem and why is it more important than instilling models with the ability to logically plan and answer correctly?
- pxoe 3y agothe timeline for LLMs and image generation has been 6+ years. it is not a thing where it "arrived just this year, and only just changing". it's been in a development for a long time. and yet.
- KHRZ 3y agoJust a question, do you remember a source for all the knowledge in your mind, or did you at least try to remember?
- pxoe 3y agoa computer isn't a human. aren't computers good at storing data? why can't they just store that data? they literally have sources in datasets. why can't they just reference those sources? human analogies are cute, but they're completely irrelevant. it doesn't change that it's specifically about computers, and doesn't change or excuse how computers work.
- umvi 3y agoYes, computers are good at storing data. But there's a big difference between information stored in a database and information stored in a neural network. The former is well defined, the latter is a giant list of numbers - literally a black box. So in this case, the analogy to a human brain is fairly on-point because just as you can't perfectly cite every source that comes out of your (black box) brain, other black boxes have similar challenges.
- wrs 3y agoThe analogy to a database is also irrelevant. LLMs aren’t databases.
- qup 3y agoWhen all the legal precedents we have are about humans, human analogies are incredibly relevant.
- jazzyjackson 3y agoThere is a hundred years of legal precedents in the realm of technology upsetting the assumptions of copyright law. Humans use tools - radios, xerox machines, home video tape. AI is another tool that just makes making copies way easier. The law will be updated, hopefully without comparing an LLM to a man.
- Foobar8568 3y agoSo why my employer implementation version of azure chatgpt on our document systems can successfully cite its sourced documents?
- layer8 3y agoBecause the model proper wasn’t trained on those documents, it’s just RAG being employed with the documents as external sources. It’s a fundamentally different setup.
- Tao3300 3y agoMy understanding is that this lawsuit is about the training corpus. This is on the level of asking it to cite its sources for a/an/the.