3 ms·
> 2024 called and wants its talking points back. It's a classic example of the kinds of questions LLMs get wrong. There's plenty of others. Not sure your point
by TonyAlicea10 3mo ago
> 2024 called and wants its talking points back.
It's a classic example of the kinds of questions LLMs get wrong. There's plenty of others. Not sure your point here. We can easily find things a 4-9 year old will talk about that an LLM will get wrong, hallucinate, etc.
> a child tutored by a good quality state-of-the-art LLM with a good teaching-focused harness could have better learning outcomes than a child without it.
There's a lot of work being done in 'could'. And it's entirely ignoring the dangers.
I'm not saying "no LLMs in education". I'm a technical educator and I give students LLM prompts and agent skills that I've built to help them learn.
This isn't that. We're talking about giving an LLM to a 4-9 year old and saying "this is your teacher".
- blitzar 3mo ago> lot of work being done in 'could' if my aunt had wheels, she'd be a bicycle.
- cindyllm 3mo ago[dead]
- Marha01 3mo ago> It's a classic example of the kinds of questions LLMs get wrong. There's plenty of others. Not sure your point here. We can easily find things a 4-9 year old will talk about that an LLM will get wrong, hallucinate, etc. My point is that the hallucination rate of modern LLMs, while not zero, is so low that this is no longer an issue in practice. > There's a lot of work being done in 'could'. And it's entirely ignoring the dangers. I agree that there are some risks. But for many children, the alternative is not human teacher, but no teacher at all. This considerably changes the risk-benefit calculus, IMHO.
- TonyAlicea10 3mo ago> My point is that the hallucination rate of modern LLMs, while not zero, is so low that this is no longer an issue in practice. That's just not true. Talk to an LLM about a subject you know deeply. They make up things all the time.
- senordevnyc 3mo agoWith stuff that a 4-9 year old is going to be learning? Color me skeptical.
- TonyAlicea10 3mo agoYou’re applying human cognition expectations to an LLM. Just because the topic is simple to a human doesn’t mean it improves the likelihood the LLM will behave better. It’s all inference to the machine.
- senordevnyc 3mo agoI’m applying my experience with these models, as well as common sense. What are you applying? Would you even admit LLMs are useful for this if they’re perfect at topics that kids would be learning? Or are your objections actually ideological rather than pragmatic?
- TonyAlicea10 3mo agoI'm not sure how common sense says that LLMs will hallucinate less on topics for children. That's not how they work. And no, LLMs are not perfect. They're non-deterministic. And no, even if they were perfect my statement was that it is irresponsible to teach a child to trust an LLM. Even if it was perfect on these topics, as they move forward in education that trust would later fail them. That is a problem both ideologically and pragmatically. You have to think long-term. On a side note: ask an LLM to write a story at a 4-9 year old level. Work with it for awhile doing iterations and it will hallucinate a previous detail or change a character name as context rots. Lots of other examples of "basic" topics that still produce hallucinations.
- senordevnyc 3mo agoThe common sense part is not holding LLMs to a standard of absolute perfection that literally no other medium, or humans, can reach. You’re welcome to your views, but they don’t make a lot of sense, or fit with reality. Continue angrily tilting against windmills I guess.
- sjsdaiuasgdia 3mo ago> My point is that the hallucination rate of modern LLMs, while not zero, is so low that this is no longer an issue in practice. Ahahahahahahahahahahahahahahahahaha Do you actually use these things?
- Marha01 3mo ago> Do you actually use these things? I use them every day for SWE. I cannot remember any hallucination from the last 6 months. It pretty much stopped being an issue from the beginning of this year, if you use top paid models. Free models or lower-tier paid models can still hallucinate, though.
- sjsdaiuasgdia 3mo agoI can use Claude as much as I want at work, and I get hallucinations every day. I'm not even doing anything all that fancy with it, but I get frequent hallucinatory behavior regardless of the model or how I go about things. Every single time I have asked Claude or another "top tier" offering a detailed question about something I'm an expert in, there's something wrong in the answer. Usually many things. Do you actually look at what comes out?
- wizzwizz4 3mo agoI would suggest that your ability to discern when the model is wrong might be the thing that has changed.
- duskdozer 3mo agoThis is what it seems to be. I've had to review the LLM output from coworkers who have insisted that the "old" models are terrible but the new ones are very good. They are always discussing the latest models and their tweaks and things. Their output is just.... not very good. For example, it may not fail tests, but it does so for wrong reasons and will fail under different conditions. Or it checks or guards for things in trivially redundant ways, like doing something like `if (x > 5 && x > 3)` that betray its lack of "knowing" what it's doing. And I can't even get answers from them anymore because they just feed anything I say into the LLM and copy/paste the response. I'm basically being forced to code with LLMs via review through a person proxy. Or maybe they just have it hooked up to read it directly. It's maddening. It's like having a realistic enough chatbot just shuts people's brains off.
- gverrilla 3mo ago> It's a classic example of the kinds of questions LLMs get wrong. There's plenty of others. Not sure your point here. We can easily find things a 4-9 year old will talk about that an LLM will get wrong, hallucinate, etc. I'm not participating in the debate, but fyi I have just asked haiku, sonnet, opus and fable that question and they all got it right. Tried with the letter s in portuguese word "sussuarana" (mountain lion) too - same thing.