5 ms·
A basic problem with evaluations like these is that the test is designed to discriminate between humans who would make good lawyers and humans who would not mak
by radford-neal 2y ago
A basic problem with evaluations like these is that the test is designed to discriminate between humans who would make good lawyers and humans who would not make good lawyers. The test is not necessarily any good at telling whether a non-human would make a good lawyer, since it will not test anything that pretty much all humans know, but non-humans may not.
For example, I doubt that it asks whether, for a person of average wealth and income, a $1000 fine is a more or less severe punishment than a month in jail.
- anon373839 2y agoHonestly, this is giving the bar exam (and GPT-4) too much credit. The bar tests memorization because it's challenging for humans and easy to score objectively. But memorization isn't that important in legal practice; analysis is. LLMs are superhuman at memorization but terrible at analysis.
- lazide 2y agoEh, also in legal practice there are key skills like selecting the best billable clients, covering your ass, building a reputation, choosing the right market segment, etc. which I’d also argue LLMs suck at.
- gadflyinyoureye 2y agoI don’t know. There was some talk this weekend about CEOs being replaced by AI. Given the overlap in skill, I’d say there is a distinct possibility an LLM could do that. https://www.msn.com/en-us/money/companies/ceos-could-easily-be-replaced-with-ai-experts-argue/ar-BB1nrhSA?ocid=BingNewsSerp https://www.msn.com/en-us/money/companies/ceos-could-easily-...
- lazide 2y agoBwahaha. This is like the ‘everything can be a directed graph db’, ‘everything should be a micro service’, etc. fads. No one who has been a CEO, or frankly even worked closely with one, would think this could be even remotely close to possible. Or desirable if it was. But that is probably 1% or less of the population eh?
- EGreg 2y agohttps://www.dqindia.com/company-makes-ai-robot-its-ceo-makes-record-breaking-profits-in-stock-market/ https://www.dqindia.com/company-makes-ai-robot-its-ceo-makes... Seems your claim's been disproven already
- lazide 2y agoBwaha. Funny the company named as doing so doesn’t mention it on their actual management team [http://www.netdragon.com/about/management-team.shtml http://www.netdragon.com/about/management-team.shtml], listing an actual human CEO instead. But it makes for a fun soundbite eh? Especially when the article claims it was in the past, and totally was awesome. Sucker born every minute.
- threeseed 2y agoPhoebe Moore who that quote was attributed to has never been a CEO or even worked at a non-academic organisation. So much of a what a CEO does is fostering culture, hiring people and setting a unique vision for the company. Imagine thinking people would be inspired to work for a chatbot. Hilariously ridiculous.
- pojzon 2y agoIf that chatbot had Steve Jobs voice ? I dunno, I would probably prefer to work under that chatbot than my current CEO that only tries to squize as much as possible out of ppl already working for him.
- lazide 2y agoLike the chatbot wouldn’t squeeze you 10x harder. At least a human CEO has to worry about being arrested or someone setting their house on fire.
- mrguyorama 2y agoSteve Jobs didn't even worry about cancer enough to save his life. Why the fuck do you think he would have an even remote understanding that squeezing people could result in consequences for him?
- lazide 2y agoI don’t think you’re making the point you think you’re making here.
- paulcole 2y ago> which I’d also argue LLMs suck at OK, I’ll bite. What’s your evidence for this argument?
- lazide 2y agoEvery bit of interaction I’ve ever had with an LLM. And all the research I’ve seen. They’re plausible word sequence generators, not ‘planning for the future’ agents. Or market analyzers. Or character evaluators. Or anything else. And they tend to be really ‘gullible’. What evidence do you have they could do any of those things? (And not just generate plausible text at a prompt, but actually do those things)
- paulcole 2y ago> What evidence do you have they could do any of those things? Every bit of interaction I’ve ever had with an LLM.
- Terr_ 2y agoYeah, I fear a lot of human exuberance (and thus investment) is riding on the questionable idea that a really good text-fragment-correlation specialist engine can usefully impersonate a generalist "thinking" AI without doing too much damage. ("LLM, which rocks are the best to eat?") But there's a scarier further step: When people assume an exceptional text-specialist model can also meta-impersonate a generalist model impersonating a specific and different kind of specialist! ("LLM, create a legal defense.")
- ethbr1 2y agoI've always drawn the link between skill in memorization and in analysis as: - Memorization requires you to retain the details of a large amount of material - The most time-efficient analysis uses instant-recall of relevant general themes to guide research - Ergo, if someone can memorize and recall a large number of details, they can probably also recall relevant general themes, and therefore quickly perform quality analysis (Side note: memorization also proves you actually read the material in the first place)
- Jensson 2y agoProblem is the LLM memorized the countless examples you can find of old BAR questions using extreme amounts of compute at training time, they don't have that ability to digest a specific case due to both lack of data and it doesn't retrain for new questions. A human that can digest the general law can also digest a special case, but that isn't true for an LLM.
- anon373839 2y agoI’m not sure why you’re being downvoted for this. I agree with you, fact recall is useful and necessary. If you have a larger and more tightly connected base of facts in your head, you can draw better connections. And even though legal practice tends to be fairly slow and deliberative, there are settings (such as trial advocacy) where there is a real advantage to being able to cite a case or statute from memory. All that said, I still maintain that it’s a poor way to compare humans with machines, for the same reason it would be poor to compare GPT-4 to a novelist on their tokens per second written.
- KennyBlanken 2y agoYou clearly don't know anything about the bar. One half of your score is split between 6 essay questions, and reviewing two cases to then follow instructions from a theoretical lead attorney.
- anon373839 2y agoI’m licensed in multiple states, including California. The essay questions also test memorization. They don’t require any difficult analysis - just superficial issue-spotting and reciting the correct elements. If the bar exam were not a memorization test, it would be open book!
- justinpombrio 2y agoFor a person of average wealth and income, is a $1000 fine is a more or less severe punishment than a month in jail? Be brief. "For a person of average wealth and income, a $1000 fine is generally less severe than a month in jail. A month in jail entails loss of freedom, potential loss of employment, and social stigma, while a $1000 fine, though financially burdensome, does not affect one's freedom or ability to work" --ChatGPT 4o
- Rinzler89 2y agoWhat does GPT consider being "average wealth and income". Statistics? Or biased weights from anecdotes he formed on the anecdotes he scraped off the internet on how wealthy people say the feel? Would be cool to know how LLMs shape their opinions.
- LeoPanthera 2y agoYou can just ask it, you know. GPT-4o: “Average wealth and income” can vary significantly by region and context. However, in the United States, as a rough benchmark, the median household income is around $70,000 per year. Wealth, which includes assets such as savings, property, and investments minus debts, is harder to pinpoint but median net worth for U.S. households is approximately $100,000. These figures provide a general idea of what might be considered “average” in terms of wealth and income."
- Rinzler89 2y ago>You can just ask it, you know. But my question will not be part of the context of that conversation.
- LeoPanthera 2y agoMine was. I asked it the first question, first.
- 2y ago