5 ms·
What's interesting is that sometimes it does a great job at something like telling you the holdings of a case, but then other times it gives you a completely in
by mountainb 4y ago
What's interesting is that sometimes it does a great job at something like telling you the holdings of a case, but then other times it gives you a completely incorrect response. If you ask it for things like "the X factor test from Johnson v. Smith" sometimes it will dutifully report the correct test in bullets, but other times will say the completely wrong thing.
The issue I think is that it's pulling from too many sources. There are plenty of sources that are pretty machine readable that will give it good answers. There's a lot of training that can be eked out from the legal databases that already exist that could make it a lot better. If it takes in too much information from too many sources, it tends to get garbled.
There are also a lot of areas where it will confuse concepts from different areas of law, like mixing up criminal battery with civil battery, but that's not the worst of the problems.
- lolinder 4y ago> The issue I think is that it's pulling from too many sources. There are plenty of sources that are pretty machine readable that will give it good answers. There's a lot of training that can be eked out from the legal databases that already exist that could make it a lot better. If it takes in too much information from too many sources, it tends to get garbled. No, this is a common misunderstanding about the way these things work. A LLM is not really pulling from any sources specifically. It has no concept of a source. It has a bunch of weights that were trained to predict the next likely word, and those weights were tuned by feeding in a large amount of text from the internet. Improving the quality of the sources used to train the weights would likely help, but would not solve the fundamental problem that this isn't actually a lossless knowledge compression algorithm. It's a statistical machine designed to guess the next word. That makes it fundamentally non-deterministic and unsuitable for any task where factual correctness matters (and there's no knowledgeable human in the loop to issue corrections).
- mountainb 4y agoIf that's true, then it'll never get to the point that it would need to, and its reliability would always be too low to use.
- lolinder 4y agoExactly. The paradigm is wrong, and in order to get past the problems we need a new paradigm, not incremental improvement on LLMs.
- jamesdwilson 4y agoIf it is never pulling from a source, then why is it able to provide citations? check the chat bot on you.com: https://you.com/search?q=list%20top%205%20best%20selling%20cars%20of%20all%20time%20and%20cite&tbm=youchat https://you.com/search?q=list%20top%205%20best%20selling%20c...
- dragonwriter 4y ago> If it is never pulling from a source, then why is it able to provide citations? If you have the training set, and models that summarize text and/or assess similarity, and some basic search engine style tools to reduce or prioritize the problem space, it seems intuitively possible to synthesize probably-credible citations from a draft response without the response being drawn from citations the way a human author would. Kind of a variant of how plagiarism detection works.
- Aune 4y agoAsking "Can you cite some legal precedence for lemon law cases?" gives an answer containing "In California, for example, the California Supreme Court in the case of Lemon v. Kurtzman (1941) held that a vehicle which did not meet the manufacturer's express warranty was a "lemon" and the manufacturer was liable for damages." I dont think that case exist, there is a first amendment case Lemon v. Kurtzman, 403 U.S. 602 (1971) though. I can't find any reference to Kurtzman or 1941 in any of the references. I think the answer is that the AI generating the text, and the code supplying the references are distinct and do not interact.
- lolinder 4y agoYou.com is hugged to death right now, but from what I can see it's a different kind of chatbot. It looks closer to Google's featured snippets than it is to ChatGPT. That kind of chatbot has different limitations that would make it unsuitable to be an unsupervised legal advice generator.
- retrac 4y agoOne useful way to think of language models is that they are statistical completion engines. It attempts to create a completion to the prompt, and then evaluates the likelihood, in a statistical sense, that the completion would follow the prompt, based on the patterns in the training data. A citation in legalese is very common. A citation that is similar or identical to actual citations, in similar contexts, is therefore an excellent candidate for the completion. A fake citation that looks like a real citation is also a rather good candidate, and will sometimes squeak past the "is this real or fake?" metric used to evaluate generated potential responses. This may seem like "pulling from a source" but there is no token, semantic information, or even any information in the model about where and when the citation was encountered. There is no identifiable structure or object (so far as anyone can tell anyway) in the model that is a token related to and containing the citation. It just learns to create fake citations so convincingly, that most of the time they're actually real citations.
- skissane 4y agoWhat if we prompt the LLM to generate a response with citations, and then we have program which looks up the citations in a citation database to validate their correctness? Responses with hallucinated citations are thrown away, and the LLM is asked to try again. Then, we could retrieve the text of the cited article, and get another LLM to opine on whether the article text supports the point it is being cited in favour of. I think a lot of these problems with LLMs could be worked around with a few more moving parts.
- theLiminator 4y agoDefinitely, no one is arguing that an AI lawyer will be the near future, but I can totally see it being good enough for the vast majority of small scale lawsuits within 10-20 years.