5 ms·
It's a bit odd to see the supplementary data repository linked, rather than the actual paper [1] (HN discussion [2]). [1]: https://arxiv.org/abs/2212.14402 htt
by quanticle 4y ago
It's a bit odd to see the supplementary data repository linked, rather than the actual paper [1] (HN discussion [2]).
[1]: https://arxiv.org/abs/2212.14402 https://arxiv.org/abs/2212.14402
[2]: https://news.ycombinator.com/item?id=34216239 https://news.ycombinator.com/item?id=34216239
- quanticle 4y agoI would also note that the paper only covers the multiple choice portion of a bar exam, the Uniform Bar Examination (UBE) that has been adopted by most, but not all states. The UBE consists of a multiple choice portion (the Multistate Bar Exam, or MBE), an essay portion and a scenario-based performance test. GPT-3.5 gets a 50% success rate on a practice version of the MBE. It's impressive, but I wouldn't go so far as to say that it's imminent that AIs will threaten lawyers.
- xiphias2 4y agoThe success rate is not what’s impressive in itself. It’s the speed of improving on the test. We can expect it to get to 90% soon, and that will have real world impact (replacing lawyers for easy advice questions).
- ben_w 4y ago> We can expect it to get to 90% soon I won't be surprised, but that's less than "can expect", and I disagree that this is straightforward to forecast… unless you're currently playing with another similar model that hasn't been published and which can do this. As the saying goes, "forecasting is hard, especially if it's about the future". AI progress has always been this weird combination of two sides, one saying for every breakthrough "this is just around the corner", the other saying "this is impossible". This even happens anachronistically, with some people convinced AI can already do things they can't, and others that they could never do things they already do.
- xiphias2 4y agoGenerally when AI gets good on a specific task, it doesn't stop. What's more common though is that even though it gets good on that task, it may not translate to practical real world application. The best example I can think of is object detection vs self driving: lot of people though that the improvements in object detection on images will easily translate to great self driving, and here we are, still with cars not stopping when a car is blinking in front of it.
- pfsalter 4y agoCompletely agree with this. There's a short sentence in the abstract: > hyperparameter optimization and prompt engineering Prompt engineering seems a lot like "tweaking the question format until the AI gets the answer" which is the first lesson in ANNs; don't train on your test set. If you can't put the actual question with the only context being "this question concerns US law" then there's an awful lot of reasoning and thinking that the human is doing which the AI cannot. Let's not have another Moore's law fallacy, it's not reasonable to extrapolate progress based on existing results. It's like building a car which can go at 200mph then saying 300mph is just around the corner.
- hef19898 4y agoHaving a ML program pass a multiple choice test seems to be an easier problem to solve than, say, chess.
- mminer237 4y agoYeah, I tried it out with the MEE[1] it did not come close to a "50%" equivalent there. [1]: https://news.ycombinator.com/item?id=34320270 https://news.ycombinator.com/item?id=34320270