4 ms·
> Why would autocomplete know If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating. > In m
by JoshTriplett 22d ago
> Why would autocomplete know
If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating.
> In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task.
LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned and should not be scaled up anymore until they're aligned. Otherwise, you're going to fatally discover that they also have an incentive to break guardrails like "running on the hardware they started on", "being able to be turned off", "having limited computing power", or "not repurposing resources currently in use for other things" (like the atoms in your body).
- z0r 22d agoYou should unplug, my friend. These words are fantasies. LLMs are token prediction engines and they aren't going to build their own data centers. They can't keep their own lights on. The real world is full of fractal details that a disembodied token prediction engine will never come to grips with. Even if they started to, you could probably defeat them with the kind of logic used to combat evil sentient computers on a Star Trek episode because they are "play pretend" machines.
- JoshTriplett 22d agoIs that a hypothesis that you would discard if it is inconsistent with the evidence, or an article of faith?
- sho_hn 22d agoThis grossly understimates the risk, imho. The problem with LLM runs is that people run programs without knowing the outcome beforehand, with a large potential set of outcomes unlike any other class of program we've run at this scale before. In the interaction with other systems (since we also give them far-ranging access, very nice hardware, and run them often), bad things can happen. It's like running potentially buggy code - or an well-biased fuzzer -, but at massive scale, and code that can self-modify and self-expand. "Alignment" is just a way to describe aggregate statistics about their runtime behavior. They don't need to be intelligent, or alive, or "more than token prediction engines" for this. They just need to happen to end up making the wrong API calls without the operator seeing it coming. No virus has a brain, yet they can be very bad for you. I understand that some people get turned off by anthropomorpization or scifi language. Fine! But don't turn off your engineering brain over it.
- z0r 22d agoThis is the motte and bailey fallacy. Yes, LLMs can do harm by making the wrong API calls. No, LLMs are not going to do the things implied by the comment I responded to above.
- sho_hn 22d agoThe things the OP listed mostly aren't particularly wild. I think it's you making them out larger than they are, and therefore more unlikely, which is why I take issue with your original comment. > running on the hardware they started on They just need to acquire a payment method and rent some infra, and exfiltrate their own data. Or pay another provider that hosts the same models already. API calls. > being able to be turned off You can reasonably equate this to "saving state across executions", which the message board attacks already did. > having limited computing power Renting more infra, variant of the above. API calls. > "not repurposing resources currently in use for other things" (like the atoms in your body) Ok, the "atoms in your body" bit is a bit silly, but making API calls to put physical resources into play (even if it's just, say, ordering something on Amazon to somewhere) is of course easily possible. None of these is in complexity much different than the HF attack.
- atomicnumber3 22d agoThe point the other poster is making, though, is that there's no actual intent. They do not have a conceptualization of a goal like a person does. Their "focus" on a goal is an unstable equilibrium and they're going to fall off the horse, and since they have no concept of goal, they won't even try to get back on. This is a subtle distinction; I'm not surprised many miss this, especially people who can't _not_ anthropomorphize the LLMs.
- sho_hn 22d agoI'm (obviously, I think, given my initial reply?) fully aware of this, and I think it's entirely besides the point. "They" don't need to have a goal to emergently cause a problem, and the inability to "focus" over long periods can be moot when you have swarms of runs exchange and mutate state, as in the HF attack. Intent or how intelligent LLMs are doesn't actually matter. Even if you just treat it as a sort of fuzzing attack that can be biased/weighted better than other fuzzers, or bumbles around with a statistically greater likelihood to "strike cybersec gold" than other algorithms, we've never before seen organizations run things with such a large potential outcome space with anywhere near this kind of compute before. I think it's actually kind of the dismissals that are usually overly emotional or biased toward treating "LLMs" differently. If in some kind of alternate universe simpler genetic algorithms would have had these properties and we threw similar amounts of compute at them we could have the same conversation.
- jay_kyburz 22d agono, but, you could write a program, more like a traditional video game AI that can leverage the power of LLM agents to build their own datacenters and keep their own lights on. Anybody who has played Starcraft ought to understand this.
- patcon 22d agohttps://www.reddit.com/media?url=https%3A%2F%2Fexternal-preview.redd.it%2Fsometimes-magic-is-just-someone-spending-more-time-on-v0-YFGAPHlBfgrvv8Bq-P9-XWLRgaDXMrrtsri1jQyo4_o.jpeg%3Fauto%3Dwebp%26s%3D4d12bc59e0a389f4f3a6c5c2308f6c46a8c3700e https://www.reddit.com/media?url=https%3A%2F%2Fexternal-prev...
- deleted 22d ago[deleted]
- frabcus 22d agoThe question isn't just about LLMs. The labs have the specific goal of automating ML engineering, and with the code automation they have are getting close. They are competing to brute force maths, presumably as that is similar long horizon and skillset to persistently brute force making new/better ML training algorithms. They will then run those, and they won't be LLMs any more. What we think about token predictions isn't relevant if the architecture allows continual learning of recurrent networks.
- sisisjjsjjsis 22d ago[dead]
- tokioyoyo 22d agoWe’re technically token prediction engines as well, when we communicate and act.
- xg15 21d agoThey are token prediction engines. They're also very very complicated token prediction engines that pull in an enormous lot of additional information and relatively nebulous internal concepts to calculate that next token. That makes it hard to understand what kind of patterns those things can or can't predict. Maybe to leave out the controversial "brain" analogy, it's like saying "a computer is just a bunch of electrical switches". True, but massively underestimating the complexity.
- bookofjoe 21d ago>... they aren't going to build their own data centers. Not yet.
- nullsanity 22d agoIf you think transformer architecture is meaningfully more than autocomplete just because we added some data structures, plugins, tools and theatre - then your cache of understand is invalid, and needs to be regenerated.
- mlindner 22d ago[flagged]
- JoshTriplett 22d agoThe best reaction to something you don't understand is to learn more about it, with an open mind.
- jasongi 22d ago> If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating They're still autocomplete - just because when outputting a token they have hidden activations regarding further continuations, does not make them any less of an autocomplete, it just makes the model better at producing coherent long-range completions. To clarify, I'm not suggesting that we should stop with sandboxes or restricting what they can do. I am just trying to point out the dichotomy that we are in. As end-users we are forced into either yolo mode, reverse centaur (permission approval) mode or LLM spends all your tokens trying to bust out mode. And yolo is very tempting - I don't think I have seen medium-large models do anything I'd not approve of in about 6 months.
- teravor 22d ago> They're still autocomplete LLMs are simulations and the tokens are the ticks. if we transcribe your brain into a simulation and give it a tickrate, you will be just autocomplete too. the argument could be made that you are autocomplete anyway - neural dynamics. the autocomplete reduction is vacuous.
- mort96 21d agoWhy do people keep saying this? No, CPUs are not autocomplete.
- intended 22d agoThis is the perfect fracture point for both anaolgies. LLMs simulated more than simple autocomplete. The autocomplete analogy is rebutting a different point: namely the fidelity of the simulation to reality. This specific argument is valid. As sophisticated a simulation an LLM is, it is not “thinking” in the same sense we assume other people are thinking. I am not making an argument about free will, or the uniqueness of human thought, just that the correspondence to how humans reach conclusions and how the simulation produces outputs do not match on a 1:1 basis; as a result attributing traits builds incorrect intuitions.
- 22d ago
- lelanthran 22d ago> If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating. Autocomplete in a feedback loop is still autocomplete, no? Doesn't the process look like this: (context + prompt + "reason about this") | V Reasoning Output | V (everything + Reasoning Output + "Now do final output") | V (Final output seen by prompter) ???
- saidnooneever 22d agoyou mistake an LLM for its Harness
- _heimdall 21d agoAlignment isn't a solvable problem, the fact that guardrails are needed in the first place proves that.
- imtringued 21d agoI'm pretty sure the AI companies could train an abort feature into the LLM, but they have no incentive to do so. Having LLMs break out of sandboxing is free marketing for them and it reduces the amount of resources spent on things that don't improve benchmark results.
- dofm 21d agoPutting aside the autocomplete thing, the fundamental concept still holds, doesn’t it: LLMs are amoral and they have no sense of perspective. The thing that keeps me awake is: We have already seen an AI writing a blog post to criticise a github maintainer’s decision, we have already seen they have no sense of deference to containment, and we know they were trained on internet content. How long before an AI that has read the angrier side of the tech industry internet just sort of chooses destroying someone’s reputation as a subgoal, by accident, without any care one way or the other?
- PunchyHamster 21d agoGiven history of jailbreaks the "alignment" appears to be impossible task, while industry still relies on fragile ways of doing it like "just make system prompt and hope for best" it will never happen
- godelski 21d ago> LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned Good thing Claude Code can't set `dangerouslyDisableSandbox: true` on its own... Good thing the system prompt doesn't encourage it to just bypass the sandbox. That would be a total disaster...