6 ms·
> we usually say that we don't know I think this is one of the distinguishing attributes of human failures. Human failures have some degree of predictability.
by 13years 1y ago
> we usually say that we don't know
I think this is one of the distinguishing attributes of human failures. Human failures have some degree of predictability. We know when we aren't good at something, we then devise processes to close that gap. Which can be consultations, training, process reviews, use of tools etc.
The failures we see in LLMs are distinctly of a different nature. They often appear far more nonsensical and have more of a degree of randomness.
The LLMs as a tool would be far more useful if they could indicate what they are good at, but since they cannot self reflect over their knowledge, it is not possible. So they are equally confident in everything regardless of its correctness.
- ragequittah 1y agoI think the last few years are a good example of how this isn't really true. Covid came around and everyone became an epidemiologist and public health expert. The people in charge of the US government right now are also a perfect example. RFK Jr. is going to get to the bottom of autism. Trump is ruining the world economy seemingly by himself. Hegseth is in charge of the most powerful military in the world. Humans pretending they know what they're doing is a giant problem.
- 13years 1y agoThey are different contexts of errors. Take any of these humans in your example, and give them an objective task, such as take any piece of literal text and reliably interpret its meaning and they can do so. LLMs cannot do this. There are many types of human failures, but we somewhat know the parameters and context of those failures. Political/emotional/fear domains etc have their own issues, but we are aware of them. However, LLMs cannot perform purely objective tasks like simple math reliably.
- SpicyLemonZest 1y ago> Take any of these humans in your example, and give them an objective task, such as take any piece of literal text and reliably interpret its meaning and they can do so. I’m not confident that this is so. Adult literacy surveys (see e.g. https://nces.ed.gov/use-work/resource-library/report/statistical-analysis-report/literacy-everyday-life-results-2003-national-assessment-adult-literacy https://nces.ed.gov/use-work/resource-library/report/statist...) consistently show that most people can’t reliably interpret the meaning of complex or unfamiliar text. It wouldn’t surprise me at all if RFK Jr. is antivax because he misunderstands all the information he sees about the benefits of vaccines.
- 13years 1y ago> most people can’t reliably interpret the meaning of complex or unfamiliar text But LLMs fail the most basic tests of understanding that don't require complexity. They have read everything that exists. What would even be considered unfamiliar in that context? > RFK Jr. is antivax because he misunderstands all the information he sees about the benefits of vaccines. These are areas where information can be contradictory. Even this statement is questionable in its most literal interpretation. Has he made such a statement? Is that a correct interpretation of his position? The errors we are criticizing in LLMs are not areas of conflicting information or difficult to discern truths. We are told LLMs are operating at PhD level. Yet, when asked to perform simpler everyday tasks, they often fail in ways no human normally would.
- SpicyLemonZest 1y ago> But LLMs fail the most basic tests of understanding that don't require complexity. Which basic tests of understanding do state-of-the-art LLMs fail? Perhaps there's something I don't know here, but in my experience they seem to have basic understanding, and I routinely see people claim LLMs can't do things they can in fact do.
- 13years 1y agoTake a look at this vision test - https://www.mindprison.cc/i/143785200/the-impossible-llm-vision-test https://www.mindprison.cc/i/143785200/the-impossible-llm-vis... It is an example that shows the difference between understanding and patterns. No model actually understands the most fundamental concept of length. LLMs can seem to do almost anything for which there are sufficient patterns to train on. However, there aren't infinite patterns available to train on. So, edge cases are everywhere. Such as this one.
- SpicyLemonZest 1y agoI don't see how this shows that models don't understand the concept of length. As you say, it's a vision test, and the author describes how he had to adversarially construct it to "move slightly outside the training patterns" before LLMs failed. Doesn't it just show that LLMs are more susceptible to optical illusions than humans? (Not terribly surprising that a language model would have subpar vision.)