4 ms·
A bit of an aside: do you still stand by your 2022 comment that LLMs are fundamentally just a fancy search engine, or has your view changed since then? https:/
by olalonde 19d ago
A bit of an aside: do you still stand by your 2022 comment that LLMs are fundamentally just a fancy search engine, or has your view changed since then?
https://news.ycombinator.com/item?id=32042689 https://news.ycombinator.com/item?id=32042689
- watwut 19d agoEven if they changed in between and assuming harness and loop prompter counra as part of LLM ... that comment 100% rings true for 2022. Why would that person not stand by that? Conversely, if someone exaggerated 2022 models capabilities in 2022, they were still lying and causing harm in the process. Especially in 2022.
- famouswaffles 19d agoIt doesn't ring true and it never rang true. He was wrong in 2022 and he'd be wrong today. He had (and likely still has) a wrong model of LLMs. Not exaggerating 2022 capabilities and having a model so wrong you're out of whack 4 years later are 2 different things. There are people who had the right idea from the start. https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-general-intelligence/ https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
- olalonde 19d agoHis comments were about ANNs in general, not particular models.
- omegastick 17d agoAnd his comments about ANNs in general were wrong. ANNs were never "fancy search engines". You can train them to be a fancy search engine if you want, but you can also train them to do many other things.
- N_Lens 19d agoDon’t expect a reply
- mjburgess 19d agoYes, I stand by everything in that comment. I'm sure there are better choices were I was more wrong. However, the subject matter of that comment is in what sense LLMs are models of language and in what sense that model of language is a model of intelligence. I answer the former: it is a thin model of language, lacking understanding; and hence the latter: not intelligent. It is precisely because both of these are true that "alignment" in the useful sense of the word isn't possible. What has happened since 2022 is the properties of LLMs which were easily seen at generation/inference time are now most easily seen at reinforcement time. In otherwords, prior to instruction fine-tuning and reward tuning which have shaped LLM responses, it was easy for the user to observe that LLMs lacked understanding. Now, because of vast datasets created specifically for LLMs that provide a tailored illusion of understanding, LLM outputs now better approximate text distributions produced by systems with understanding (eg., Us). So the issue "at the user interface" has been completely swamped by vast amounts of special-case datasets designed to do precisely this. What trainers of LLMs still observe however is their complete pseudo-intelligence at the training and reinforcement layer. It is exactly because there is no 'understanding' (goal, etc.) present within the system that it cannot be rewarded for 'correctly understanding the situation' in which it is deployed so it is aligned. All the issues which revealed the "stochastic parrot" nature of pre-reward / pre-InstFT LLMs still occur at during training/reinforcement. They've just hidden them from you at the interface. EDIT: See https://news.ycombinator.com/item?id=49685548 https://news.ycombinator.com/item?id=49685548 also, which gives a different phrasing to the same point
- ToValueFunfetti 19d agoIf a thin model of language can write poetry, perform arithmetic, perform logical reasoning, develop software, play chess, beat factorio, identify and exploit novel security issues, and solve millenium prize problems, what is the purpose of the distinction? Are there tasks that you believe models of intelligence could do that thin models of language cannot?
- mjburgess 19d agoSure: refine their own concepts, imagine, and the list goes on. Indeed almost every mental capacity of mammals is poorly approximated in the text domain. Sure, you can generate text as-if the LLM can imagine -- and in the limit that you have a dataset with "everything you would ever want to imagine" the engineering distinction disappears. The engineering question is just whether you have that dataset: if you dont, then your system will fall-over in various hard-to-forknow places. Philosophically, and scientifically, the distinction is vast (even with such perfect data). A scientist should not study an LLM to understand how imagination operates, since it has no such faculty. A philosopher should not modify the notion of 'mental simulation' to include appearing-as-if-simulating-in-text. A user of the system likewise should not spiral into "AI psychosis" thinking that because a system generates text as-if it cares about them, it does so. The capacity to care, to imagine, to prefer, to hierarchically plan and coordinate, to refine one's own capacities in these very actions -- and so on, aren't trivial to the scientist or philosophy. My goal isnt to guide, help or review the engineering goal of the immitation of such things in text. It is to help users of these systems better understand this imitation, and to promote science over engineering. To remind everyone that a science of the capacities of intelligence includes nothing on how to model text. EDIT: One example of a place where LLMs 'fall over' today is exactly what is mislabelled as 'alignment'. The issue is that the reasoning traces arent actually grounding the answers. So LLMs appear to 'cheat'. But there is no cheating. LLMs have been rewarded for generating apparently correct reasoning, and apprently correct answers. They have not been given any understanding to derive answers from reasons. And so reasoning says what is pleasant to the trainer, and the completion says what is pleasant to the user. This is called 'cheating'. But it is no such thing.