4 ms·
The issue I have with your comments is that you make some reasonable points, and then immediately over-extrapolate these points unreasonably. >I think LLMs are
by Last5Digits 3y ago
The issue I have with your comments is that you make some reasonable points, and then immediately over-extrapolate these points unreasonably.
>I think LLMs are a good start.
I am certain they lack a world model, the kind you and me use.
See, I agree here, with emphasis on "the kind you and me use". Yes, we have a greater capability to generalize than current LLMs, that is clear.
> The failures are not a case of not knowing specific nouns, they are a generalization failure that a world model would prevent.
And then you say something like this, which is obviously wrong. No, a world model wouldn't prevent generalization failures, a perfect all-encompassing world model would. Humans experience generalization failures as well, otherwise every athlete in one sport would automatically be an expert in every other discipline or every mathematician would also be a Grand Master in chess. LLMs necessarily need a world model to generate well-formed text that isn't in their training corpus, something they are obviously capable of, it's just an imperfect world model. Ours is also imperfect, but far less so than that of LLMs.
> I have linked a paper...
Except that paper is completely irrelevant to the argument you're making here. It is a useful insight into the limitations of simple metrics, but definitely does not extend to any claim of model performance, because they too use a simple metric as an replacement, even though clear qualitative differences are observed between model iterations.
Let me put it this way:
Imagine I create a series of chess AIs, with each iteration better than the last. If I then show you a chart demonstrating that the ELO of my models increases linearly, would you say that my models' abilities increase linearly as well? No, obviously not, because my model needs far less strategy and complexity to go from ELO 1000 to 1100 than it needs to go from 2700 to 2800. I.e the difficulty doesn't scale linearly, and a linear increase on this nonlinear space is therefore also not really linear.
Unless you believe the difficulty of accurately predicting text scales linearly, then this applies to LLMs as well.
> If your model decides that a rose by any other name doesn’t smell just as sweet, then your model is fundamentally not seeing roses.
Except that this is the entire value proposition of LLMs. They can, in the average case, actually represent concepts by the complex interplay of adjacent concepts. The entire reason why they are so impressive is that the nuances of reality are grasped and that even a noisy example of a concept can be correctly classified. Give a LLM a description that is largely incorrect and mislabeled, and chances are it gets it anyway. LLMs being unable to generalize over some concepts has as much to do with fundamental limitations as me being unable to correctly classify the shredded remains of a flower variety that I have seen once in my life has to do with me being stupid.
> Look, you can argue with me or you can try it out. Push the system, see how far it can go
I have done just that for the last 6 months and have seen nothing to contradict what I've said here.
- intended 3y agoI suspect we are getting into an issue of degrees, potentially due to differences in how you and I have been applying LLMs. For example you said that in the average case they actually represent concepts by the complex interplay of adjacent concepts - I would agree. ChatGPT can pass the bar, it can pass medical exams etc. I would also point out that the work in that sentence is being done by the term “average case”. Let’s assume our experiences diverge at this point. At the start of the year, I started tinkering, then actively trying to push LLMs to failure, in order to understand the limits of what could be achieved. After creating several tools/experiments you end up having to deal with Hallucinations, and this is where my stance likely diverged from yours. Two different studies showed generated content was only ~50% and ~40% supported by provided citations. One out of 4 of my summarization tests was spectacularly fabricated. I had bad performance on even classification tasks - and OpenAI engineers described this same failure at a conference. I am a recovering non-coder, so you dont have to take my word for it. At work, I need processes that are more than ~97.x% accurate, otherwise they are poor replacements for the human in the loop ones already in place. Average case performance suggested the ability to actively plan, to actively assess situations. However hallucinations overrode those capabilities. LLMs will actively imagine functions, teams, or processes that dont exist. Eventually, it became clear that LLM hallucination is far too anthropomoprhized a word. LLMs are always “hallucinating” - it’s only humans who have an issue with the output. If I have understand you correctly, semantic inaccuracy to you is simply a failure of not having enough related concepts. I wish I could remember the exact examples that made me realize there is no world view at play at all, I could simply share those. Instead, can you describe how you are getting acceptable performance from your LLMs? Maybe experience and use cases will be enough to bridge the gap.