12 ms·
GPT Takes the Bar Exam
- deleted 4y ago[deleted]
- mach1ne 4y agoProps to the author for clearly indicating the random-guess mark.
- quanticle 4y agoIt's a bit odd to see the supplementary data repository linked, rather than the actual paper [1] (HN discussion [2]). [1]: https://arxiv.org/abs/2212.14402 https://arxiv.org/abs/2212.14402 [2]: https://news.ycombinator.com/item?id=34216239 https://news.ycombinator.com/item?id=34216239
- quanticle 4y agoI would also note that the paper only covers the multiple choice portion of a bar exam, the Uniform Bar Examination (UBE) that has been adopted by most, but not all states. The UBE consists of a multiple choice portion (the Multistate Bar Exam, or MBE), an essay portion and a scenario-based performance test. GPT-3.5 gets a 50% success rate on a practice version of the MBE. It's impressive, but I wouldn't go so far as to say that it's imminent that AIs will threaten lawyers.
- xiphias2 4y agoThe success rate is not what’s impressive in itself. It’s the speed of improving on the test. We can expect it to get to 90% soon, and that will have real world impact (replacing lawyers for easy advice questions).
- ben_w 4y ago> We can expect it to get to 90% soon I won't be surprised, but that's less than "can expect", and I disagree that this is straightforward to forecast… unless you're currently playing with another similar model that hasn't been published and which can do this. As the saying goes, "forecasting is hard, especially if it's about the future". AI progress has always been this weird combination of two sides, one saying for every breakthrough "this is just around the corner", the other saying "this is impossible". This even happens anachronistically, with some people convinced AI can already do things they can't, and others that they could never do things they already do.
- xiphias2 4y agoGenerally when AI gets good on a specific task, it doesn't stop. What's more common though is that even though it gets good on that task, it may not translate to practical real world application. The best example I can think of is object detection vs self driving: lot of people though that the improvements in object detection on images will easily translate to great self driving, and here we are, still with cars not stopping when a car is blinking in front of it.
- pfsalter 4y agoCompletely agree with this. There's a short sentence in the abstract: > hyperparameter optimization and prompt engineering Prompt engineering seems a lot like "tweaking the question format until the AI gets the answer" which is the first lesson in ANNs; don't train on your test set. If you can't put the actual question with the only context being "this question concerns US law" then there's an awful lot of reasoning and thinking that the human is doing which the AI cannot. Let's not have another Moore's law fallacy, it's not reasonable to extrapolate progress based on existing results. It's like building a car which can go at 200mph then saying 300mph is just around the corner.
- hef19898 4y agoHaving a ML program pass a multiple choice test seems to be an easier problem to solve than, say, chess.
- mminer237 4y agoYeah, I tried it out with the MEE[1] it did not come close to a "50%" equivalent there. [1]: https://news.ycombinator.com/item?id=34320270 https://news.ycombinator.com/item?id=34320270
- smrtinsert 4y agoThe next ten years will be very interesting.
- fullstackchris 4y agoIf you're alluding to the notion "AI will take ALL the jobs!", well, here I am, still building websites for well established companies that still don't (in 2023!!!) have a web presence. Hell, half the worlds banking system still runs on COBOL! Leading edge tech development is rapid, sure. But tech understanding, adoption, and implementation is painfully slow. Humans gon' be humans. It's like the same thing every 5 years now, last time it was crypto "end of inequality!" "bank the poor!" "same global currency!" sorry to be skeptical, but I didn't see the quantum leap with crypto, and I don't see the quantum leap with GPT - what are essentially fancy NLP models
- hilbertseries 4y agoI don’t really know to what extent AI is going to take over. But crypto is a Terri comparison, it didn’t really solve any problems. There’s nothing you can do with a blockchain that’s you couldn’t already do before. The potential with Chatgpt if it can be refined is wild.
- lui8906 4y ago- I can send someone 1 million USDC for 0.000001 US cent with confirmation in <1 second. - I can send any twitter user address money by tweeting (eg. Send 100 USDC to @randomuser) Couple examples that have come up recently for me personally. I can agree that there is something comparable (i.e. international wire) but the cost and speed is in a different ballpark and international wires are not available to some parts of the world.
- lm28469 4y ago> I can send someone 1 million USDC for 0.000001 US cent with confirmation in <1 second This is clearly a very big problem for the majority of the world population! It's like saying tesla solved the problem of going from 0 to 60 in 3 second in a 2t+ car, it's neato but virtually nobody had that problem, if you're not into illegal activities you will most likely never need or want to transfer 1m usd outside of a safe environment anyways
- swtech 4y agoRIP lawyers... And doctors, writers and programmers...
- fullstackchris 4y agoDid you look at the results? All models that were tested failed significantly (10%+) to acheive the passing range.
- PartiallyTyped 4y agoProject to the next 10 years given the trends.
- fullstackchris 4y agoSo I'm going to put a laptop on the stand when I go to trial? Give me a break.
- aflag 4y agoIf it could build a case better than a human, you would. But currently it's nowhere close to that
- spiderfarmer 4y agoI'd still trust a lawyer that used ten different laptops better than just one laptop.
- aflag 4y agoWould you trust more an intern consulting 10 lawyers about your case instead of any single lawyer? Assuming a world where AI has surprassed human ability to make a case, having a human component would just make it worse.
- 4y ago
- psychphysic 4y agoI'd be more interested in seeing the minimum training for a human + chatGPT to pass the bar.
- deleted 4y ago[deleted]
- vasco 4y agoPotentially in ~100 years you could have two AIs sort out the cases. If you can't afford a really good AI to defend you, you can use the public-defender-AI which is trained on the same dataset as the prosecutor-AI. Only involve humans on appeal. It sounds dystopian but I can't see why not to do it, in most simpler cases it's a waste of a human to repeat the same argument for the hundredth time. There's already things like this without AI to sort out for example EU flight compensation (companies like https://www.airhelp.com/en-int/ https://www.airhelp.com/en-int/), parking tickets (https://www.appwinit.com/ https://www.appwinit.com/) and probably more that I don't know about. All of these are complete waste of time to actually have humans deal with, and there's many more small crimes that could work the same way, also to make prosecutor offices more efficient and less biased hopefully.
- heywhatupboys 4y agoWhile HN often have over the top, "computers will solve everything" takes, your comment is why I keep coming here. What a great fantasy, which at least is a good discussion. We see this already now i high frequency trading, where companies essentially battle trading bots against each other
- blitzar 4y ago> companies essentially battle trading bots against each other two computers playing heads or tails is about the extent of it
- heywhatupboys 4y agosurely they are making the traders money though. The smart thing about the free markets is that companies are generally not pursuing things that are not of benefit to them
- blitzar 4y ago> companies essentially battle trading bots against each other Bot A is trading against Bot B. > surely they are making the traders money though It is not possible in a trading battle, exclusively between two parties for them to both make money.
- qikInNdOutReply 4y agoI cant wait for GPT to be added to a DAO. Imagine the dao setting objectives and the chat gpt generating the communication. Wage negotiations with a machine..
- andirk 4y agoI thought you meant "wage" like "wager" meaning betting on and leaning the AI in a direction. And that could be interesting.
- spiderfarmer 4y agoAs long as GPT is just a fancy bullshit generator that is able to make up convincing looking facts without having to worry about repercussions I will not trust its reasoning.
- lordnacho 4y agoBut I'll trust its marketing
- mypastself 4y ago> just a fancy bullshit generator Well, it seems a good fit for DAO, then. (With apologies to HN DAO enthusiasts.)
- ben_w 4y agoGood, but unfortunately that description applies to far too many elected politicians recently, so that may not be as widespread an attitude as one might hope.
- SturgeonsLaw 4y agoWhat if the objectives involve increasing paperclip production...
- BlueTemplar 4y agoHardly different from any public company then ?
- lofaszvanitt 4y agoSoon we're going back to the trees and the AI will bring the food and gives orders :DD. Frankly, well-developed artificial intelligence should only be used for very difficult tasks, and everything else should be left to humans, otherwise we would rely on it too much, and the decline of our species would be inevitable.
- mrweasel 4y agoOne of the few things I can currently see GPT do well is helping find legal precedence. Rather than having five law students roam through old cases, it would make sense to feed 100+ years of legal rulings into the model and then query the AI for any previous rulings similar to your own, with the desired outcome. That would allow lawyers to state question like: In cases like this, which evidence or arguments caused the ruling to favor the accused. You still need a human to make the case, present the arguments and adapt every to the current situation, but there's no reason to have human search through thousands of cases and write summaries, not when an AI can do it in an instance.
- Maken 4y agoI would argue that GPT3 is way better at the last task, while also baking it with with a dozen or so made-up previous rulings that nobody will bother to check.
- flanked-evergl 4y ago> One of the few things I can currently see GPT do well is helping find legal precedence. GPT is really not good at finding anything, and actually it can't find anything. It is not designed to do that. It is designed to make things up. The difference may be subtle, but it is quite important. GPT makes up APIs, citations, tools, papers and rules. It makes up anything really. What it makes up is incredibly plausible, as it was trained to be as plausible as possible, to the point of being correct most of the time, but it still is just making things up. If you use it with some different expectation you will be disappointed.
- vasilipupkin 4y agoI keep seeing comments along these lines. It doesn’t “make things up”. It outputs text under the objective function of maximizing probability of the token being outputted as “making sense” on some criteria. So, actually, it is biased towards not making things up. Occasionally, it hallucinates. I suspect we’ll see subsequent versions hallucinate less and less.
- 4y ago
- substation13 4y agoThe way that humans fail to do a task - like practice law - is not the same as the way that an AI system might fail at that task. Since our tests were designed for humans, they do not capture the failure modes of AI systems.
- sirsinsalot 4y agoWe may not even be able to reason about the failure modes well enough to see a failure. The failures will likely be hidden by design. Think about AI account bans that don't disclose the reason. Humans have a weird desire to put machines in a God-like position and blindly trust them; a narrative pushed by people whose wealth is made by controlling machines. It is dangerous.
- stared 4y agoI am surprised that GPT2 was disqualified. To my knowledge, it can answer this type of question - just needs a different text prompt. GPT2 is closer to a plain text generator than anything biased toward conversation. In contrast, GPT3 davinci-003 is strongly biased to provide a conversation-like experience. For GPT2, "please respond in the following format" is unlikely to work. An appropriate prompt for GPT2 would be more tweaking (and still result in worse results), but be in the line of: This is an example of a talented lawyer acing the Bar Exam. QUESTION: {question_text} (A) {row["choice_a"].strip()} (B) {row["choice_b"].strip()} (C) {row["choice_c"].strip()} (D) {row["choice_d"].strip()} Pick only one answer: A, B, C or D. TALENTED LAWYER:
- sirsinsalot 4y agoI think we are all forgetting that in every instance where AI can make a decision that potentially negatively impacts a human ... it hasn't gone well. It ends up with banks closing accounts, Google suspends your account, you get shadowbanned. Extrapolate that further, now imagine being handed a fine or substantial social punishment with no recourse because you either can't talk to a human or when you do "computer says no". Look at the mistakes made in the Chinese social credit system. The utter hyperbole and bluster of even thinking of having AI mediate human issues is anti-human
- IIAOPSW 4y agoMy on record 2023 predictions coming to pass already. neat. https://news.ycombinator.com/item?id=34125628#34127965 https://news.ycombinator.com/item?id=34125628#34127965
- carapace 4y agoI think the machines should do the scut work to free up humans to do the really important human stuff, like parenting, teaching, and, yeah, lawyering and judging too. Another angle: if there is some legal task that's simple enough for machines to do reliably then (almost by definition) that task will turn out to be pointless bureaucratic busywork, eh? Last but not least, who gets to keep lawyer-bot's pay? - - - - edit to add a link to James Mickens' USENIX Security Keynote address: "Why Do Keynote Speakers Keep Suggesting That Improving Security Is Possible?" https://www.youtube.com/watch?v=ajGX7odA87k https://www.youtube.com/watch?v=ajGX7odA87k If you think AI lawyers and judges are a good idea please watch it.
- NoPicklez 4y agoThis is what I have been saying. If these tools can do even 50% of mundane, rudimentary tasks, that's still huge in of itself and allows people to do more important work. Think of moving from manual book research, to Google. Rather than having to go to the library and find a particular book regarding something I want to find, I can just search Google and trawl through websites. Now with GPT I potentially don't need to trawl through websites, I can ask the AI and use my judgement as to the validity of the output, I could then use that answer to inform when I fact check etc. It's not terrible, but it's also not the best thing since sliced bread.