19 ms·
Re-Evaluating GPT-4's Bar Exam Performance
- Bromeo 2y agoVery interesting. The abstract claims that although GPT-4 was claimed to score in the 92nd percentile on the bar exam, when correcting for a bunch of things they find that these results are overinflated, and that it only scores in the 15th percentile specifically on essays when compared to only people that passed the bar. That still does put it into bar-passing territory, though, since it still scores better than about one sixth of the people that passed the exam.
- falcor84 2y agoIf I understand currently, they measured it at the 69th percentile for the full test across all test takers, so definitely still impressive.
- _fw 2y agoSo it knows more about the law than you do, but less than they do. Really glad to see research replicated like this. I’m not surprised that the 90th percentile doesn’t hold up. It’s still handy though.
- radford-neal 2y agoA basic problem with evaluations like these is that the test is designed to discriminate between humans who would make good lawyers and humans who would not make good lawyers. The test is not necessarily any good at telling whether a non-human would make a good lawyer, since it will not test anything that pretty much all humans know, but non-humans may not. For example, I doubt that it asks whether, for a person of average wealth and income, a $1000 fine is a more or less severe punishment than a month in jail.
- anon373839 2y agoHonestly, this is giving the bar exam (and GPT-4) too much credit. The bar tests memorization because it's challenging for humans and easy to score objectively. But memorization isn't that important in legal practice; analysis is. LLMs are superhuman at memorization but terrible at analysis.
- lazide 2y agoEh, also in legal practice there are key skills like selecting the best billable clients, covering your ass, building a reputation, choosing the right market segment, etc. which I’d also argue LLMs suck at.
- gadflyinyoureye 2y agoI don’t know. There was some talk this weekend about CEOs being replaced by AI. Given the overlap in skill, I’d say there is a distinct possibility an LLM could do that. https://www.msn.com/en-us/money/companies/ceos-could-easily-be-replaced-with-ai-experts-argue/ar-BB1nrhSA?ocid=BingNewsSerp https://www.msn.com/en-us/money/companies/ceos-could-easily-...
- lazide 2y agoBwahaha. This is like the ‘everything can be a directed graph db’, ‘everything should be a micro service’, etc. fads. No one who has been a CEO, or frankly even worked closely with one, would think this could be even remotely close to possible. Or desirable if it was. But that is probably 1% or less of the population eh?
- EGreg 2y agohttps://www.dqindia.com/company-makes-ai-robot-its-ceo-makes-record-breaking-profits-in-stock-market/ https://www.dqindia.com/company-makes-ai-robot-its-ceo-makes... Seems your claim's been disproven already
- lazide 2y agoBwaha. Funny the company named as doing so doesn’t mention it on their actual management team [http://www.netdragon.com/about/management-team.shtml http://www.netdragon.com/about/management-team.shtml], listing an actual human CEO instead. But it makes for a fun soundbite eh? Especially when the article claims it was in the past, and totally was awesome. Sucker born every minute.
- justinpombrio 2y agoFor a person of average wealth and income, is a $1000 fine is a more or less severe punishment than a month in jail? Be brief. "For a person of average wealth and income, a $1000 fine is generally less severe than a month in jail. A month in jail entails loss of freedom, potential loss of employment, and social stigma, while a $1000 fine, though financially burdensome, does not affect one's freedom or ability to work" --ChatGPT 4o
- Rinzler89 2y agoWhat does GPT consider being "average wealth and income". Statistics? Or biased weights from anecdotes he formed on the anecdotes he scraped off the internet on how wealthy people say the feel? Would be cool to know how LLMs shape their opinions.
- LeoPanthera 2y agoYou can just ask it, you know. GPT-4o: “Average wealth and income” can vary significantly by region and context. However, in the United States, as a rough benchmark, the median household income is around $70,000 per year. Wealth, which includes assets such as savings, property, and investments minus debts, is harder to pinpoint but median net worth for U.S. households is approximately $100,000. These figures provide a general idea of what might be considered “average” in terms of wealth and income."
- Rinzler89 2y ago>You can just ask it, you know. But my question will not be part of the context of that conversation.
- LeoPanthera 2y agoMine was. I asked it the first question, first.
- 2y ago
- dogmayor 2y agoThe bigger issue here is that actual legal practice looks nothing like the bar, so whether or not an llm passes says nothing about how llms will impact the legal field. Passing the bar should not be understood to mean "can successfully perform legal tasks."
- ben_w 2y agoIndeed, and this is also the general problem with most current ways to evaluate AI: by every test there's at least one model which looks wildly superhuman, but actually using them reveals they're book-smart at everything without having any street-smarts. The difference between expectation and reality is tripping people up in both directions — a nearly-free everything-intern is still very useful, but to treat LLMs* as experts (or capable of meaningful on-the-job learning if you're not fine-tuning the model) is a mistake. * special purpose AI like Stockfish, however, should be treated as experts
- KennyBlanken 2y ago> Passing the bar should not be understood to mean "can successfully perform legal tasks." Nobody does except a bunch of HNers who among other things, apparently have no idea that a considerable chunk of rulings and opinions in the US federal court system and upper state courts are drafted by law clerks who, ahem, have not taken the bar yet... The point of the bar and MPRE is like the point of most professional examinations: try to establish minimum standards. That said, the bar does test for "successfully perform legal tasks", actually. For the US bar, a chunk of your score is based off following instructions on case from the lead attorney, and another chunk is based on essay answers. Literally demonstrating that you can perform legal tasks and have both the knowledge and critical thinking skills necessary. Further, as previously mentioned, in the US, people usually take it after a clerkship...where they've been receiving extensive training and experience in practical application of law. Further, law firms do not hire purely based on your bar score. They also look at your grades, what programs you participated in (many law schools run legal clinics to help give students some practical experience, under supervision), your recommendations, who you clerked for, etc. When you're hired, you're under supervision by more senior attorneys as you gain experience. There's also the MPRE, or ethics test - which involves answering how to handle theoretical scenarios you would find yourself in as a practicing attorney. Multiple people in this discussion are acting like it's a multiple choice test and if you pass, you're given a pat on the ass and the next day you roll into criminal court and become lead on a murder case...
- Digory 2y agoThey originally scored against a test usually taken by people who failed the bar. So, GPT-4 scores closer to the bottom of people who pass the bar the first time. In other words, it matches the people who cull the rules from texts already written, but who cannot apply it imaginatively.
- speedgoose 2y ago> In other words, it matches the people who cull the rules from texts already written, but who cannot apply it imaginatively. Where did you find that in the article?
- Digory 2y agoIf you can recite the black letter law, you've got a good chance of passing the bar. The higher essay scores usually require creative arguments about resolving competing rules and policies. It's easier to extract the formal statement of the rule against perpetuities from a reddit corpus, than to apply the rule to an artificially complex fact pattern in an essay question.
- jeffbee 2y agoIt appears that researchers and commentators are totally missing the application of LLMs to law, and to other areas of professional practice. A generic trained-on-Quora LLM is going to be straight garbage for any specialization, but one that is trained on the contents of the law library will be utterly brilliant for assisting a practicing attorney. People pay serious money for legal indexes, cross-references, and research. An LLM is nothing but a machine-discovered compressed index of text. As an augmentation to existing law research practices, the right LLM will be extremely valuable.
- violet13 2y agoIt is a lossy compressed index. It has an approximate knowledge of law, and that approximation can be pretty good - but it doesn't know when it's outputting plausible but made-up claims. As with GitHub Copilot, it's probably going to be a mixed bag until we can overcome that, because spotting subtle but grave errors can be harder than writing something from scratch. There's already a fair number of stories of LLMs used by an attorney messing up court filings - e.g., inventing fake case law.
- jeffbee 2y agoI am not suggesting that the generative aspects would be useful in drafting motions and such. I am suggesting that their tendency towards false results is harmless if you just use them as a complex index. For example, you could ask it to list appellate cases where one party argued such-and-such and prevailed. Then you would go read the cases.
- thehoneybadger 2y agoIt is difficult to comment without sounding obnoxious, but having taken the bar exam, I found the exam simple. Surprisingly simple. I think it was the single most over hyped experience of my life. I was fed all this insecurity and walked into the convention center expecting to participate in the biggest intellectual challenge in my life. Instead, it was endless multiple choice questions and a couple contrived scenarios for essays. It may also be surprising to some to understand that legal writing is prized for its degree of formalism. It aims to remove all connotation from a message so as to minimize misunderstanding, much like clean code. It may also be surprising, but the goal when writing a legal brief or judicial opinion is not to try to sound smart. The goal is to be clear, objective, and thereby, persuasive. Using big words for the sake of using big words, using rare words, using weasel words like "kind of" or "most of the time" or "many people are saying", writing poetically, being overly obtuse and abstract, these are things that get your law school application rejected, your brief ridiculed, and your bar exam failed. The simpler your communication, the more formulaic, the better. The more your argument is structured, akin to a computer program, the better. As compared to some other domain, such as fiction, good legal writing much easier for an attention model to simulate. The best exam answers are the ones that are the most formulaic and that use the smallest lexicon and that use words correctly. I only want to add this comment because I want to inform how non-lawyers perceive the bar exam. Getting an attention model to pass the bar exam is a low bar. It is not some great technical feat. A programmer can practically write a semantic disambiguation algorithm for legal writing from scratch with moderate effort. It will be a good accomplishment, but it will only be a stepping stone. I am still waiting for AI to tackle messages that have greater nuance and that are truly free form. LLMs are still not there yet.
- euroderf 2y ago> It may also be surprising to some to understand that legal writing is prized for its degree of formalism. It aims to remove all connotation from a message so as to minimize misunderstanding, much like clean code. > The more your argument is structured, akin to a computer program, the better. You certainly make legal writing sound like a flavor of technical writing. Simplicity, clarity, structure. Is this an accurate comparison ?
- elicksaur 2y ago> Furthermore, unlike its documentation for the other exams it tested (OpenAI 2023b, p. 25), OpenAI’s technical report provides no direct citation for how the UBE percentile was computed, creating further uncertainty over both the original source and validity of the 90th percentile claim. This is the part that bothered me (licensed attorney) from the start. If it scores this high, where are the receipts? I’m sure OpenAI has the social capital to coordinate with the National Conference of Bar Examiners to have a GPT “sit” for a simulated bar exam.
- Suppafly 2y ago>This is the part that bothered me (licensed attorney) from the start. If it scores this high, where are the receipts? I'm not a licensed attorney, but that's also bothered me about all of these sorts of stories. There is never any proof provided for any of the claims, and the behavior often contradicts what can be observed using the system yourself. I also assume they cook the books a little by having included a bunch of bar exam specific training when creating the model in first place specifically to better on bar exams than in general.
- fnordpiglet 2y agoScoring 96 percentile among humans taking the exam without moving goal posts would have been science fiction two years ago. Now it’s suddenly not good enough and the fact a computer program can score decent among passing lawyers and first time test takers is something to sneer at. The fact I can talk to the computer and it responds to me idiomatically and understands my semantic intent well enough to be nearly indistinguishable from a human being is breath taking. Anyone who views it as anything less in 2024 and asserts with a straight face they wouldn’t have said the same thing in 2020 is lying. I do however find the paper really useful in contextualizing the scoring with a much finer grain. Personally I didn’t take the 96 percentile score to be anything other than “among the mass who take the test,” and have enough experience with professional licensing exams to know a huge percentage of test takers fail and are repeat test takers. Placing the goal posts quantitatively for the next levels of achievement is a useful exercise. But the profusion of jaded nerds makes me sad.
- QuantumGood 2y agoIt scored less than 50% when compared to people who had taken the test once.
- Workaccount2 2y agoThe nerds aren't jaded, they are worried. I'd be too if my job needed nothing more than a keyboard to be completed. There are a lot of people here who need to squeeze another 20-40 years out of a keyboard job.
- threeseed 2y agoSimilar comments were made that microwaves will eliminate cooking. At the end of the day (a) LLMs aren't accurate enough for many use cases and (b) there is far more to knowledge worker's jobs than simply generating text.
- imtringued 2y agoYou're assuming that keyboard jobs are easier simply because the models were built to output text, but nothing prevents physical motion to be easier simply due to sheer repetitiveness. In fact, you can get away with building dedicated robots e.g. for drywall spraying and sanding, whereas the keyboard guys tend to have to switch tasks all the time.
- gnicholas 2y agoThis analysis touches on the difference between first-time takers and repeat takers. I recall when I took the bar in 2007, there was a guy blogging about the experience. He went to a so-so school and failed the bar. My friends and I, who had been following his blog, checked in occasionally to see if he ever passed. After something like a dozen attempts, he did. Every one of us who passed was counted in the pass statistics once. He was counted a dozen times. This dramatically skews the statistics, and if you want to look at who becomes a lawyer (especially one at a big firm or company), you really need to limit yourself to those who pass on the first (or maybe second) try.
- lccerina 2y agoIt's amazing the level of mental gymnastics I see in the comments trying to justify a piece of technology that is evidently not as good as they believed to be...
- Slyfox33 2y agoAI is just the next tech hype scam after crypto and nfts.