4 ms·
Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, whi
by swatcoder 2y ago
Many/most of the folks "defaulting to 'it's just autocomplete'" have recognized that issue from day one and see it as an inextricable character of the tool, which is exactly why it's clearly not something we'd invest agency in or imagine being intelligent.
Alignment researchers are hoping that they can overcome the problem and prove that it's not inextricable, commercial hypemen are (troublingly) promising that it's a non-issue already, and commercial moat builder are suggesting that it's the risk that only select, authorized teams can be trusted to manage, but that's exactly the whole house of cards.
Meanwhile, the folks on the "autocomplete" side are just engineering ways to use this really cool magic autocomplete tool in roles where that presumed inextricable flaw isn't a problem.
To them, there's no debate to have over "does it have real feelings?" and not really any debate at all. To them, these are just novel stochastic tools whose core capabilities and limitations seem pretty self-apparent and can just be accommodated by choosing suitable applications, just as with all their other tools.
- kalkin 2y agoI'd be very interested in an example of someone who said "it's just autocomplete" and also explicitly brought up the risk of something like alignment faking, before 2024. I can think of examples of people who've been talking about this kind of thing for years, but they're all people who have no trouble with applying the adjective "intelligent" to models.
- wizzwizz4 2y agoIf you go back through my Hacker News comments, I believe you'll see this. Perhaps look for keywords "GPT-2", "prediction", and "agent". (I don't know how to search HN comments efficiently.) I was talking about this sort of thing in 2018, though I don't think I published anything that's still accessible, and I'd hardly call myself an expert: it's just obviously how the system works.
- sakjur 2y agoSearching HN comments is probably easiest done through Algolia: https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=Author%3AWizzwizz4&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
- kalkin 2y agoNo results for wizzwizz4 "GPT", although it does look like search results may be incomplete: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=Author%3AWizzwizz4%20gpt&sort=byDate&type=comment https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
- wizzwizz4 2y agoThe "author:" part was just being treated as a keyword, and was restricting the results too much. I haven't found the comments I was looking for, but I have found Chain of Thought prompting, 11 months before it was cool (https://news.ycombinator.com/item?id=26063189 https://news.ycombinator.com/item?id=26063189): > Instruction: When considering the sizes of objects, you will calculate the sizes before attempting a comparison. > GPT-2 doesn't have a concept of self, so it constructs plausible in-character excuses instead. I also found the Great Translation Argument (see https://news.ycombinator.com/item?id=35530858 https://news.ycombinator.com/item?id=35530858 and https://news.ycombinator.com/item?id=35530855 https://news.ycombinator.com/item?id=35530855). And, apparently, I was still framing things in terms of the sci-fi nonsense in 2020, but I had the right ideas (https://news.ycombinator.com/item?id=22802105 https://news.ycombinator.com/item?id=22802105): > Corollary: you can't patch broken FAI designs. Reinforcement learning (underlying basically all of our best AI) is known to be broken; it'll game the system. Even if they were powerful enough to understand our goals, they simply wouldn't care; they'd care less than a dolphin. https://vkrakovna.wordpress.com/2018/04/02/specification-gaming-examples-in-ai/ https://vkrakovna.wordpress.com/2018/04/02/specification-gam... > And there are far too many people in academia who don't understand this, after years of writing papers on the subject. This criticism applies to RLHF, so it counts, imo. Not as explicit as my (probably unpublished, almost certainly embarrassing) wild ravings from 2018, but it's before 2024.
- swatcoder 2y agoYou'd expect to find it expressed in different language, since the "autocomplete" people (including myself) are naturally not going to approach the issue as "alignment" or "faking" in the first place because both of those terms derive from the alternate paradigm ("intelligence"). But you can dig back and see plenty of these people characterizing LLM's as delivering output like an improviser that responds to any whiff of a narrative by producing a melodramatic "yes, and..." output. With plenty of narrative and melodrama in the training material, and with that material's conversational style seeming to be essential in getting LLM's to produce familiar English and respond in straightforward ways to chatbot instructions, you have to assume that the output will easily develop unintended melodrama itself, and therefore have to take personal responsibility as an engineer for not applying it to use cases where that's going to be problematic. (Which is why -- in this view -- it's simply not a sufficient or suitable tool to expect to ever grant agency over critical systems, even while still having tremendous and novel utility in other roles.) We've been openly pointing that out at least since ChatGPT brought the technology into wide discussion over two years ago, and those of us who have opted to build things with it just take all that into account as any engineer would.
- kalkin 2y agoIf you can link to a specific example of this anticipation, that would be informative. I don't care about use of the term "alignment" but I do think what's happening here is more specific and interesting than "unintended melodrama." Have you read any of the paper?
- swatcoder 2y agoYes, I read the paper. To an "just autocomplete" person, the authors are straightforwardly sharing summaries of some sci-fi fan fiction that they actively collaborated with their models to write. They don't see themselves as doing that, because they see themselves as objective observers engaging with a coherent, intelligent counterparty with an identity. When a "just autocomplete" person reads it, though, it's a whole lot of "well, yeah, of course. Your instructions were text straight out of a sci-fi story about duplicitous AI and you crafted the autocompleter so that it outputs the AI character's part before emitting its next EOM token". That doesn't land as anything novel or interesting because we know that's what it would do because that's very plainly how a text autocompeter would work. It just reads as an increasingly convoluted setup, driven by continued the time, money, and attention, that keeps being poured into the research effort. (Frankly, I don't really feel like digging up specific comments that I might think speak to this topic because I don't trust you would ever agree if you're not seeing how the other side thinks already. It's unlikely any would unequivocally address what you see as the interesting and relevant parts of this paper because, in that paradigm, the specifics of this paper are irrelevant and uninteresting.)
- qsort 2y agoThe point is that if the limitations of current LLMs persist, regardless of how much better they get, this is not a problem at all, or at least not a new one. Let's say you are given the declaration but not the implementation of a function with the following prototype: const char * AskTheLLM(const char *prompt); Putting this function in charge of anything, unless a restricted interface is provided so that it can't do much damage, is simply terrible engineering and not at all how anything is done. This is irrespective of whether the function is "aligned", "intelligent" or any number of other adjectives that are frankly not really useful to describe the behavior of software. The same function prototype and lack of guarantees about the output is shared by a lot of other functions that are similarly very useful but cannot be given unrestricted access to your system for precisely the same reason. You wouldn't allow users to issue random commands on a root shell of your VM, you wouldn't let them run arbitrary SQL, you wouldn't exec() random code you found lying around, you wouldn't pipe any old string into execvpe(). It's not a new problem, and for all those who haven't learned their lesson yet: may Bobby Tables'mom pwn you for a hundred years.
- Majromax 2y ago> Let's say you are given the declaration but not the implementation of a function with the following prototype: > const char * AskTheLLM(const char prompt); > Putting this function in charge of anything, unless a restricted interface is provided so that it can't do much damage, is simply terrible engineering and not at all how anything is done. Yes, but that's exactly how people* use any system that has an air of authority, unless they're being very careful to apply critical thinking and skepticism. It's why confidence scams and advertising work. This is also at the heart of current "alignment" practices. The goal isn't so much to have a model that can't automate harm as it is to have one that won't provide authoritative-sounding but "bad" answers to people who might believe them. "Bad," of course, covers everything from dangerously incorrect to reputational embarrassments.
- zahlman 2y ago> The goal isn't so much to have a model that can't automate harm as it is to have one that won't provide authoritative-sounding but "bad" answers to people who might believe them. We already know it will do this - which is part of why LLM output is banned on Stack Overflow. None of the properties being argued about - intelligence, consciousness, volition etc. - are required for that outcome.
- zahlman 2y agoI have been reading arguments about "things like alignment faking" for years, while simultaneously holding that "it's just autocomplete". The alignment-faking arguments are still terrifying to the extent that they're plausible. In the hypothetical where I'm wrong about it being "just autocomplete" (and fundamentally, inescapably so), the risk is far greater than can be justified by the potential benefits. But that's itself a large part of why I believe those arguments are false. If I gave them credit and they turned out to be false, then I figure I have succumbed to a form of Pascal's Mugging. If I don't give them credit and it turns out that a hostile, agentive AGI has been pretending to be aligned, I don't expect anyone (including myself) to survive long enough to rub it in my face. Honestly, I sometimes worry that we'll doom ourselves by taking AI too seriously even if it's indeed "just autocomplete". We've already had people commit suicide. The sheer amount of text that can now be generated that could propose harmful actions and sound at least plausible is worrying, even if it doesn't reflect the intent of an agent to convince others to take those actions. (See also e.g. Elsagate.)
- comp_throw7 2y ago> But that's itself a large part of why I believe those arguments are false. If I gave them credit and they turned out to be false, then I figure I have succumbed to a form of Pascal's Mugging. If I don't give them credit and it turns out that a hostile, agentive AGI has been pretending to be aligned, I don't expect anyone (including myself) to survive long enough to rub it in my face. I'm sorry, but this is a crazy reason to believe something is false. Things are either true or they aren't, and if the world would be nicer to live in if thing X was false does not actually bear on whether thing X is false or not.
- zahlman 2y agoIt's not "the world would be nicer to live in if thing X was false". It's "the world would cease to exist if thing X were true".
- cruffle_duffle 2y ago> To them, there's no debate to have over "does it have real feelings?" and not really any debate at all. To them, these are just novel stochastic tools whose core capabilities and limitations seem pretty self-apparent and can just be accommodated by choosing suitable applications, just as with all their other tools. Boom. This is my camp. Watching the Anthropic YouTube video in the linked article was pretty interesting. While I only managed to catch like 10 minutes, I was left with an impression that some of those dudes really think this is more than a bunch of linear algebra. Like they talk about their LLM in these dare I say, anthropic, ways that are just kind of creepy. Guys. It’s a computer program (well, more accurately it’s a massive data model). It does pretty cool shit and is an amazing tool whose powers and weakness we have yet to fully map out. But it isn’t human nor any other living creature. Period. It has no thoughts, feelings or anything else. I keep wanting to go work for one of these big name AI companies but after watching that video I sure hope most people there understand that what they are working on is a tool and nothing more. It’s a pile of math working over a massive set of “differently compressed” data that encompasses a large swath of human knowledge. And calling it “just a tool” isn’t to dismiss the power of these LLM’s at all! They are both massively overhyped and hugely under hyped at the same time. But they are just tools. That’s it.
- Terr_ 2y agoIMO some of this stems from a kind of author/character confusion by human observers. The LLM can generate an engaging story from the perspective of a ravenous evil vampire, but that doesn't mean that vampires are real nor that the LLM wants to suck your blood. Re-running the same process but changing the prompt from "you are a vampire" to "you are an LLM", doesn't magically create a metaphysical connection to the real-world algorithm.
- IanCal 2y agoThe problem I have with this is it seems to rapidly assume that humans and animals are somehow magic. If you think we're physical things obeying physical laws, then unless you think there's something truly uncomputable about us then we can be replicated with maths. In which case is this just an argument about scale? > It’s a pile of math working over a massive set of “differently compressed” data that encompasses a large swath of human knowledge. So am I but over less knowledge and more interactions.
- famouswaffles 2y agoYou say "autocompleters" recognized this issue as intrinsic from day 1 but none of these are issues specific to being "autocomplete" or not. Tomorrow, something could be released that every human on the planet agrees is AGI and 'capable of real general reasoning' and these issues would still abound. Whether you'd prefer to call it "narrative drift" or "being convinced and/or pressured" is entirely irrelevant.