13 ms·
GPT-4 could pass bar exam, AI researchers say
- dragonwriter 4y ago> By passing this exam, lawyers are admitted to the bar of a U.S. state. No, they aren't. Meeting certain preparatory requirements (the details vary but in most US jurisdiction an accredited/approved law school program or, in some, what amounts to an apprenticeship with a licensed practitioner of certain duration and standards is required) and then passing the bar exam allows this. The difference is important, the bar exam is not seen, standing alone, aa adequate proof of readiness.
- nafeen 4y agoBar exam down. Medical next? While GPT-3 wasn't advanced enough for cracking medical exam, it was used for notable contributions. For e.g. this is an interesting 2021 paper about "Medically Aware GPT-3 as a Data Generator" - https://aclanthology.org/2021.nlpmc-1.9.pdf https://aclanthology.org/2021.nlpmc-1.9.pdf Would love to see if GPT-4 is advanced enough to take medical exam.
- ben_w 4y agoI want to perform some research of my own on which exams chatGPT can and can't do. It's multilingual, so can people from outside the UK (I already know where to get those) point me at some example exams and marking schemes? Any level, not just top. Currently have Polish school maths: https://news.ycombinator.com/item?id=34205732 https://news.ycombinator.com/item?id=34205732
- preommr 4y agoHow soon before this qualifies as a public defender? Gonna put this on my dystopia bingo.
- ben_w 4y agoDefender is probably good. Prosecutor is what would worry me, given I don't know better than to blindly trust the meme that the average person commits 6 felonies before breakfast.
- jacquesm 4y agoAnd if the cost of prosecution falls then more and more of those 6 felonies will end up prosecuted. The same happened with speed cameras, initially it was to reduce accidents, now it is just another income stream (which I'm sure still reduces accidents, but that's no longer the main reason they are out there).
- microtherion 4y agoDefender is a TERRIBLE idea. I can already see the Supreme Court cases down the line: Defendant was provided a state of the art, 50 trillion parameter, neural network for their defense. The internals of this network are not auditable, but it does not tire, engage in substance abuse, or get distracted, so it will by definition represent effective assistance of counsel, even if for some unfathomable reason it decides to raise the Chewbacca Defense in a Death Penalty habeas corpus petition.
- linuxftw 4y agoI'm willing to bet a chat bot would perform better than most public defenders.
- sebzim4500 4y agoOk? This is like the arguments that self driving cars are bad if they crash even once. The question isn't "is the AI giving me the perfect legal defence?" or even "is the AI giving me a defence as good as the best lawyer money can buy?". It's "is the AI better than the public defender that I otherwise would have been given?". As soon as the answer to that last question is yes (and I have absolutely no idea when that will be), it will be extremely difficult to justify not using it.
- secondcoming 4y agoHow long until it's smart enough to be a judge?
- kneebonian 4y agoI don't know about judge but it could probably outperform most of Congress at this point.
- mdp2021 4y ago...After sharpness and judgement will be implemented. -- Incidentally: there is an interesting video interview to Noam Chomsky and Gary Marcus on limits of current attempts at https://www.youtube.com/watch?v=PBdZi_JtV4c https://www.youtube.com/watch?v=PBdZi_JtV4c ...And Gary Marcus saying just before 7:00 that "something is missing" (understatement): ontology. Gray Marcus: «...and these systems fall apart left and right». Nice summary from Gary Marcus: «What they do is, they perpetuate past data - they don't really understand the world».
- klntsky 4y agoNever, assuming current legislation. First of all, it is not formalized (despite being written with the use of bureaucratic language). So, there's no way to validate the output. Secondly, juridical system is based on authority of the state (which manifests clearly in their ability to alter the rules). Why would any sovereign ruler(s) want to get rid of their authority? The only use cases would be automatic fines for speeding or inappropriate parking - but it's already there.
- ldh0011 4y agofwiw I had my dad ask ChatGPT relatively high-level questions about his field of practice in the state he is licensed in. Some were very good answers but that some were wildly off. The ones that seemed to be better were questions about a concept (ie "What is x concept in law") while the incorrect ones were the ones asking for specifics ("What is the statute of limitations for x in y state").
- DisjointedHunt 4y agoBecause it doesn't presently have memory or look things up in a table or the internet. You will notice that both are very easy fixes that computers have perfected in retrieval over the past 5 or so decades.
- jacquesm 4y agoJust stick Google's pre-search tools in front of the current version and it would solve a large chunk of those problems. The right tool for the job, essentially. After all, you wouldn't ask your English professor to solve a math problem either.
- fnordpiglet 4y agoI’m by far a layman in this respect but I feel like it’s the difference between conceptualizing and information retrieval. Further it feels like IR is a well researched area and by allowing the conceptualizing part access to a modern IR system would allow it to form searches, pull the IR results, sift them, and summarize them.
- jerf 4y agoThe next frontier for GPT-esque technologies is building one that is capable of saying "I don't know". GPT as it stands now is essentially incapable of it. (The cases of that you see in the current ChatGPT preview are, as near as I can tell, all rules-based overlays run by OpenAI for various reasons. When it declines to comment, and then more-or-less scolds you for even asking, you got caught before even getting to the model itself.)
- pyb 4y agoSounds like they didn't have access to GPT-4, but "Based on anecdotal evidence"... they still predict this.
- naillo 4y agoSource: "I have a hunch"
- danenania 4y agoMy knee always gets achey right before a technological singularity hits.
- michpoch 4y agoMaybe they asked chatGPT.
- swyx 4y agoyeah this is really low quality for HN. source is basically "trust me i heard a guy who knows a guy"
- munchler 4y agoThey’re extrapolating from the performance of GPT-3.5. It’s speculative, but not anecdotal. GPT has improved rapidly over time, so it's not a huge leap to predict that GPT-4 will be even better.
- mtgx 4y ago[dead]
- minimaxir 4y agoFor some reason, there's a thought-leader sect of Twitter talking about how good GPT-4 is, despite OpenAI having provided zero hints of what GPT-4 could entail or be differentiated from GPT-3/chatGPT.
- booleandilemma 4y ago
- moneywoes 4y agoWonder how data biases will surface
- sdenton4 4y agoThe fun part here is that most humans in the legal profession carry pretty extreme biases, judges included... The hope for legal ai is that you could progressively improve the biases, instead of waiting for N years for a bad judge to retire same maaaaybe get replaced by someone better.
- tetris11 4y agowho though, who has access to the resources to push the boundaries of next-gen AI except the rich who already have their own biases? The AI that the public will get will be just as useful as the tech that public get now: limited, isolating, and designed to restrict their freedoms I exchange for easy entertainment
- sdenton4 4y agoI'm confident that these things will get easier. It is approximately ten thousand times easier to train a decent classifier in 2023 than it was in 2013... We're also now living in a world with foundation models and fine tuning, which makes it /very/ possible to improve and specialize publicly released models. We see a lot of that with stable diffusion already.
- SV_BubbleTime 4y agoThis is what I found immediately interesting about ChatGPT. I asked about controversial topics. Its answers didn’t seem like biases that were programmed in, but rather it took traditional media and gave it more weight than what turned out to be the truth only accepted much later on and still against a media retelling. I lost a lot of faith in it knowing it was more CNN than careful deliberating AI.
- anononaut 4y ago
- nigerianbrince 4y agoYou passed the bar!
- DisjointedHunt 4y agoIf you're not actively building it or related tech, you shouldn't carry the label "Researcher" in the press. It's like : "I'm a doctor of homeopathy so i can write a headline for a story about a neural chip implant"
- charcircuit 4y agoHow do they know GPT-4 will be enough to let it pass? Is there even a big enough difference in the training data for it to improve in the areas it was struggling with?
- sebzim4500 4y agoRumours are that GPT-4 is a significant improvement over GPT-3.5. Given how big an improvement GPT-3.5 is over GPT-3 I am inclined to believe them. Probably we will find out for sure in a few months.
- nopinsight 4y agoRelated: "Large Language Models Encode Clinical Knowledge" https://arxiv.org/abs/2212.13138 https://arxiv.org/abs/2212.13138 "On the MedQA dataset consisting of USMLE style questions with 4 options, our Flan-PaLM 540B model achieved a multiple-choice question (MCQ) accuracy of 67.6%..." "The percentages of correctly answered items required to pass varies by Step and from form to form within each Step. However, examinees typically must answer approximately 60 percent of items correctly to achieve a passing score." -- https://www.usmle.org/bulletin-information/scoring-and-score-reporting https://www.usmle.org/bulletin-information/scoring-and-score... . It seems like the models in the paper could pass USMLE already. Some tests suggest that Med-PaLM is close to human clinicians in many aspects, incl reasoning (Figures 6-7). Other tests show that Med-PaLM still returns inappropriate/incorrect results much more often than clinicians do, however (Figure 8).
- lukko 4y agoI'm kind of surprised the model doesn't score higher as there is clear pattern to questions + answers and there would a huge amount of training data for USMLE. But as stated elsewhere, there is an enormous gap between passing exams and treating real patients as a doctor. It's rarely about making obscure diagnoses found in exam questions, but about managing illness in the context of a patient and their lifestyle, with many very human aspects - difficult communication, ethics & assessing family dynamics. Written exams are just to assess whether a medical student has the minimum required knowledge to practice, but also there are lots of practical exams and communication scenarios required too. It may well be the same for lawyers - passing the bar does not really relate to actual day-to-day practice.
- Aardwolf 4y agoI envisioned a cocktail shaking robot, but apparently Bar Exam is an exam for US lawyers
- microtherion 4y agoIF it could (I wouldn't know one way or the other), I'd consider that a damning indictment of the Bar Exam failing to test for sentience, rather than evidence of GPT-4 having attained the same.
- xyzelement 4y agoBar exam is not a test of sentience but of the ability to recall, interpret, and apply the law. Because law is an entirely textual thing, I would expect GPT to be exceedingly well suited for it. I've said for a long time that most doctors and lawyers are just databases with quick and imperfect retrieval.
- forgetfulness 4y agoMaybe the "talk about your issue and get a diagnosis" area of practice (internal medicine?); since far less sophisticated manual labor can't yet be automated, surgeons are going to be irreplaceable for longer than, say, BI, and many backend or frontend developers.
- phpisthebest 4y agoIf House taught me anything it is that People Lie, and you do not have to talk to patients to diagnose them /s
- VBprogrammer 4y agoI wonder if people could be more honest with a sub-sentient AI than they could be with a real life doctor. I bet they currently are more honest in their Google searches than they are with the doctor.
- dumbfounder 4y agoIf that's all a lawyer needs to do then AI should be able to take over large portions of the law process. I saw a dystopian short recently that explored this: https://tvtropes.org/pmwiki/pmwiki.php/Film/PleaseHold https://tvtropes.org/pmwiki/pmwiki.php/Film/PleaseHold
- Workaccount2 4y agoI feel like I can now see the event horizon of commoditized intelligence. No idea what society (is "society" even the right word? Who knows) is going to look like on the other side of it, but it is going to be wildly different. Perhaps a brief period where everyone is using an AI to do their job, uh, I mean, assist their work, but beyond that it's unknowable. Moreover, this looks like it is going to be happening sooner rather than later.
- forgetfulness 4y agoWe were perhaps a bit too enamored with the idea that it was intellect that made us unique, and thus knowledge workers would be the last to be replaced. Pouring our brains out by the Petabytes for neural networks to pick them up made the economics just work for an AI industrial revolution to start from there.
- mdp2021 4y agoNo, we were enamored with the idea that intelligence was well distributed between people, as if following Descartes' massive incipit "Good sense must be the best distributed thing in the world, given that nobody seems to be asking for more". Inability to recognize intelligence is and will be devastating.
- forgetfulness 4y ago> Inability to recognize intelligence is and will be devastating. It's a pop-culture quote from a movie that was no masterpiece, I know, but "I, Robot" presented in two sentences an argument for having more sober expectations on what machine intelligence could be capable of, and of our own > Detective Del Spooner: "Can a robot write a symphony? Can a robot turn a… canvas into a beautiful masterpiece?" > Sonny: "Can you?" We're discrediting the capabilities of current machine learning models for being unable of producing the thoughts that many, many people are unable to either. Alright, so the models are not at the level that us HN philosopher kings hold ourselves to be, and they won't be Senior Architects of distributed systems or what have you very soon, but what does it say about Average Joe, slightly-above-Average Joe, and their economic prospects? Specially since in the West and much of the developing world, we were taking solace in the idea that a service economy comprised of knowledge workers would provide plenty of opportunities on a political and economic landscape where manufacture was gone, or had never arrived.
- deleted 4y ago[deleted]
- notwokeno 4y agoThe bar exam answer key can pass the bar exam, that doesn't mean that it would be a good lawyer.
- criddell 4y agoWe don't ask students to calculate sin(1.234) by hand these days. Exams for mechanical engineering students assume they will have a calculator with SIN and EXP buttons. It may soon be time to update the bar exam and assume law students have access to AI tools.
- BeFlatXIII 4y agoHow much of the bar exam consists of confident rhetoric using deductive logic? That seems to be right up the alley for GPT models.
- ss108 4y agoA minority. It's mostly about having stored legal rules in long term memory.
- elicksaur 4y agoSuch a milestone would say more about the Bar Exam (and other standardized tests) being a poor proxy for wisdom, than the advancement of computers.
- micromacrofoot 4y agoYou know how hard it can be to talk to an actual support person at some companies? Imagine that for everything.
- czzr 4y agoThere’s a gap between passing the bar exam and actually practicing law - I’m pretty certain that I (someone with no legal training whatsoever) could pass the bar exam if you gave me unlimited access to the internet and a couple of additional hours to write the test. However, I don’t think that would make me an effective lawyer. Ultimately standardised tests are proxy measurements of legal ability - it’s easy to see how a LLM could subvert the proxy without being sufficiently reliable in real life. I do expect that even unreliable versions will be very useful tools for practicing lawyers, though.
- allochthon 4y ago> I do expect that even unreliable versions will be very useful tools for practicing lawyers, though. Agreed. It's like being able to call up a map on Google Maps for an area that you're already familiar with. The map can help you remember things about the area and terrain that you might not have recalled right away. A kind of cognitive aid.
- softwaredoug 4y agoThe Bar Exam is multiple choice, right? This isn't grading some freeform essay or generating arbitrary legal opinion. It's answering from a limited set of answers. IMO it's cool, but not THAT shocking given what we've seen from ChatGPT? Especially given GPT 3.5 is only 17% below human test takers?
- post-it 4y agoNo, you're thinking of the LSAT.
- deleted 4y ago[deleted]
- morsecodist 4y agoFrom the article it looks like there are multiple choice and written sections but they only ran the model on the multiple choice portion.
- chrismcb 4y agoI would think that it could post most tests, as the tests are generally based on factual information and not creativity.
- morsecodist 4y agoI am sorry but this title is click bait. These researchers ran GPT-3.5 on only the multiple choice sections of the Bar and it passed 2/7 sections. Is this really impressive? Absolutely. But the only element of the article that is about GPT-4 potentially passing the Bar is one paragraph near the end: > According to the researchers, the history of large language model development strongly suggests that such models could soon pass all categories of the MBE portion of the Bar Exam. Based on anecdotal evidence related to GPT-4 and LAION’s Bloom family of models, the researchers believe this could happen within the next 18 months. GPT-4 could potentially pass the Bar, it could potentially do a lot of things. But by their own admission the researchers have no hard evidence for this.
- Iwan-Zotow 4y agoSo, how new knowledge would be created? GPT has no reasoning capability. So, as time goes on, information massive(s) will be filled with GPT-X made up answers. It means GPT-X+1 will be trained on GPT-X generated data. So, without reasoning, how this thing will work in perspective?
- criddell 4y agoI wouldn't assume that future versions are going to work the same way past versions did.
- Iwan-Zotow 4y agoMAybe, maybe not. Problem is with data/content creation. If all new data are created with GPT-3, how it will help GPT-4? No new original content -> no new model
- machiaweliczny 4y agoHow is baseline 50% in 4 choices exam?
- w1nst0nsm1th 4y agoAI seems to be the next financial buzzword, after crypto, gig economy, CDO, dotcom, and so on. I have seen a video a few days ago saying we are coming out of data era and entering the 'Knowledge Era' thank to AI where knowledge is following a logarythmic path. A 'revolution', a 'paradigm shift', and other bubblebabble. Who was telling that ? A 30 years old startup CEO wearing... a t-shirt and a jeans... You see the pattern. I'm not an AI specialist, but for what I know, current AI are nothing more than fine tuned statistic algorythm. Here a is a short french video with english subtitles from arte, the german-french public cultural television, about a painting coming from Midjourney : https://www.arte.tv/en/videos/110342-003-A/the-world-in-images/ https://www.arte.tv/en/videos/110342-003-A/the-world-in-imag... The video explain very well what AI are able to do (and consequently what they can't do) if you listen (read) carefully what the art historian say about the painting, which received the first price of 2022 collorado art festival. In short, the painting is nothing new by itself but a patchwork of elements from different period of art history. In other word a statistic average of previous painting, photography, drawing, etc... based on the artist prompts in midjourney. Not to say the painting is aweful, I personnaly find it's beautiful and could happily put it in my living room, but it definitively shows how current AI works, commented by an historian art specialist which has no ball in AI game.
- ben_w 4y ago> I'm not an AI specialist, but for what I know, current AI are nothing more than fine tuned statistic algorythm. Yes, but Google Page Rank is just fancy matrix multiplication and worth a fortune, while the finely tuned statistics that is machine learning (specifically chatGPT) knows more about CSS and German than I do: ~~~~ Me: Erzählen Sie bitte wie ich kann ein div tag mit schwarzes hintergrund, dreihundert px hoch und 24 pixels wide machen mit css chatGPT: Um ein div-Tag mit einem schwarzen Hintergrund, 300px hoch und 24px breit zu erstellen, könntest du den folgenden CSS-Code verwenden: div { background-color: black; height: 300px; width: 24px; } Um den CSS-Code anzuwenden, kannst du ihn entweder in einem style-Tag innerhalb deines HTML-Dokuments oder in einer externen CSS-Datei platzieren und dann per link-Tag in deinem HTML-Dokument einbinden. Beispiel: <!DOCTYPE html> <html> <head> <style> div { background-color: black; height: 300px; width: 24px; } </style> </head> <body> <div></div> </body> </html> ~~~~ The actual HTML and CSS it gave me is multiple lines and sensibly indented, don't know a convenient way to mark a block as pre-formatted. Note that chatGPT understood me correctly even though I forgot the German for "wide" and switched to English for one word only. (I do know more CSS than is in this example; I used chatGPT over the weekend to update my website, and it solved two problems that I didn't know pure CSS could even do, but that conversation is too big to bother putting into a comment here).
- deleted 4y ago[deleted]