3 ms·
We already know the system is really bad at spelling. I have Claude configured to periodically remind me “By the way, I think there are ** n's in 'banana'”, so
by mjd 1y ago
We already know the system is really bad at spelling. I have Claude configured to periodically remind me “By the way, I think there are ** n's in 'banana'”, so I don't forget what I am dealing with. It has never gotten this right.
But that doesn't mean that it is not extremely useful. It only means I shouldn't ask it to spell stuff.
If a human is unable to count the n's in 'banana' we expect them to be barely functional. Articles like this one try to draw the same inference about the LLM: it can't count 'n's, so it must not be able to do anything else either.
But it's a bad argument, and I'm tired of hearing it.
- yogurtboy 1y agoI don't disagree with your first point, that it's not still extremely useful despite its flaws. I absolutely use it to build project outlines, write code snippets, etc. Your overall conclusion though seems a little free of context. Average people (i.e. my mom googling something) absolutely do not have the wherewithal to keep track of the various pros and cons of the underlying system that generates the magical giant blue box at the top of their search that has all the answers. They are being deliberately duped by the salesmen-in-chief of these giant companies, as are all of their investors.
- thomassmith65 1y agoIt's as much that LLMs are bad at counting letters in words as it is that humans are good at it. LLMs are also bad at many things that humans don't notice immediately. That is a problem because it leads humans to trust LLMs with tasks at which LLMs currently are bad, such as picking stocks, screening job applicants, providing life advice...
- dragonwriter 1y agoThe particular problem (and one that AI firms marketing approaches have actively leveraged and made worse[0]) is that correlations between capacities that humans are used to from observing other humans do not hold for LLMs, so assumptions about what an LLM should be able to do based what is observed to do and what a human ovserved to do that qould also be expected to be capable of do not hold even as loose rules of thumb. [0] e.g., by promoting AIs as having equivalent capacities of humans of various education levels because they could pass tests that were part of the standards for, and correlate for humans with other abilities of, people with that educational background.
- drweevil 1y agoIt's a reminder that LLMs are not reasoning machines. LLMs are very useful in many cases, but one should not treat them as if they can reason.
- cube00 1y agoI can't understand why all the AI services are allowed to get away with modes such as "deep thinking" and "deep research". OpenAI even claims "reasoning" is available. > Built-in agents – deep research, ChatGPT agent, and Codex can reason across your documents, tools, and codebases to save you hours https://openai.com/chatgpt/pricing/ https://openai.com/chatgpt/pricing/
- snypher 1y agoChatGPT 5 just 'thought for a couple of seconds' and then output '2.'. Seems like we have to update our expectations as the technology improves.
- yakz 1y agoYou can tell gpt to write a program to count the n's in 'banana', and then run the program to find the answer, and it can do that.