4 ms·
That's what so surprising to me - they data clearly shows the experiment had terrible results. And the write up is nothing but the author stating: "glowing succ
by BoiledCabbage 9mo ago
That's what so surprising to me - they data clearly shows the experiment had terrible results. And the write up is nothing but the author stating: "glowing success!".
And they didn't even bother to test the most important thing. Were the LLM evaluations even accurate! Have graders manually evaluate them and see if the LLMs were even close or were wildly off.
This is clearly someone who had a conclusion to promote regardless of what the data was going to show.
- knallfrosch 9mo agoI found "well, the LLMs converge when given each other's scores, so they agree and are correct" to be quite the jump to a conclusion.
- pooper 9mo agoaccuracy versus precision is something we learn in high school chemistry. https://i.imgur.com/EshEhls.png https://i.imgur.com/EshEhls.png When someone at that level pretends to not understand it, there is no way to mince words. This is malice.
- bsenftner 9mo agoI've got a long standing disagreement with an AI CEO that believes LLM convergence indicates greater accuracy. How to explain basic cause and effect in these AI use cases is a real challenge. The essential basic understanding of what an LLM is is not there, and that lack of comprehension is a civilization wide issue.
- leoc 9mo agoAt the risk of perhaps stating the obvious, there appears to be a whiff of aggression from this article. The "fighting fire with fire" language, the "haha, we love old FakeFoster, going to have to see if we change that" response to complaints that the voice was intimidating ... if there wasn't a specific desire to punish the class for LLM use by subjecting them to a robotic NKVD interrogation then the authors should have been more careful to avoid leaving that impression.
- Hnrobert42 9mo agoYou can try out the voice yourself. It's not that bad. https://elevenlabs.io/app/talk-to?agent_id=agent_8101k9d1pq41f3rs8d9g9143dvvr&vars=eyJzdHVkZW50IjoiS29uc3RhbnRpbm9zIFJpemFrb3MiLCAibmV0aWQiOiJrcjg4OCIsICJwcm9qZWN0aWQiOiJCIn0= https://elevenlabs.io/app/talk-to?agent_id=agent_8101k9d1pq4...
- yayitswei 9mo agoTried it in earnest. Definitely detect some aggression, and would feel stressed if this were an exam setting. I think it was pg who said that any stress you add in an interview situation is just noise, and dilutes the signal. Also, given that there's so many ways for LLMs to go off the rails (it just gave me the student id I was supposed to say, for example), it feels a bit unprofessional to be using this to administer real exams.
- iamthepieman 9mo agoAlso tried it and it could have been a lot better. If I had any type of interview with that voice (press interview, mentor interview, job interview) I would think I was being scammed, sold something, or had entered the wrong room.
- Drupon 9mo agoNot that bad? I gave it a random name and random net ID and it basically screamed at me to HANG UP RIGHT NOW AND FIGURE OUT THE CORRECT NET ID. Hahaha That does not resemble any good professor I've ever heard. It's very aggressive and stern, which is not generally how oral exams are conducted. Feels much more like I'm being cross examined in court.
- plagiarist 9mo agoThe belligerence about changing the voice is so weird. And it does sort of set a tone straight off. "We got feedback that the voice was frightening and intimidating. We're keeping it tho."
- malcolmgreaves 9mo ago
- skybrian 9mo agoI don't think they're terrible, but I'm grading on a curve because it's their first attempt and more of a trial run. It seems promising enough to fix the issues and try again.
- bjt 9mo agoThey did compare the automated grades to the author's own manual ones. It's in there if you read more closely.
- wanderingbit 9mo ago> And they didn't even bother to test the most important thing. Were the LLM evaluations even accurate! This is not true; the professor and the TAs graded every student submission. See this paragraph from the article: (Just in case you are wondering, I graded all exams myself and I asked the TA to also grade the exams; we mostly agreed with the LLM grades, and I aligned mostly with the softie Gemini. However, when examining the cases when my grades disagreed with the council, I found that the council was more consistent across students and I often thought that the council graded more strictly but more fairly.)
- chairmansteve 9mo agoAs far as I can tell, there is very little empirical evidence of efficacy for most modern educational "advances". Having said that, LLMs can be good tutors if used correctly.