6 ms·
Don't anthropomorphize the language model. If you stick your hand in there, it'll chop it off. It doesn't care about your feelings. It can't care about your fee
by xmodem 5mo ago
Don't anthropomorphize the language model. If you stick your hand in there, it'll chop it off. It doesn't care about your feelings. It can't care about your feelings.
- keeda 5mo agoActually I think the opposite advice is true. Do anthropomorphize the language model, because it can do anything a human -- say an eager intern or a disgruntled employee -- could do. That will help you put the appropriate safeguards in place.
- deleted 5mo ago[deleted]
- gpm 5mo agoAn eager intern can remember things you tell beyond that which would fit in an hours conversation. A disgruntled employee definitely remembers things beyond that. These are a fundamentally different sort of interaction.
- braebo 5mo agoYou can easily persist agent memories in a markdown file though.
- whstl 5mo agoWhich it will start ignoring after two or three messages in the session.
- Quarrelsome 5mo agoand you'll blow the context over time and send to the LLM sanitorium. It doesn't fit like the human brain can. If a junior fucks production that will have extroadinary weight because it appreciates the severity, the social shame and they will have nightmares about it. If you write some negative prompt to "not destroy production" then you also need to define some sort of non-existing watertight memory weighting system and specify it in great detail. Otherwise the LLM will treat that command only as important as the last negative prompt you typed in or ignore it when it conflicts with a more recent command.
- Kim_Bruning 5mo ago> and you'll blow the context over time and send to the LLM sanitorium. It doesn't fit like the human brain can. The LLM did have this capability at training time, but weights are frozen at inference time. This is a big weakness in current transformer architectures.
- collinmcnulty 5mo agoAnd the memento guy had tattoos of key information. That didn’t make it so he didn’t have memory loss.
- WhatIsDukkha 5mo agoPretty good metaphor. Limited space to work with, highly context dependent and likely to get confused as you cover more surface area.
- troupo 5mo agoYup, and the agent will happily ignore any and all markdown files, and will say "oops, it was in the memory, will not do it again", and will do it again. Humans actually learn. And if they don't, they are fired.
- deleted 5mo ago[deleted]
- strongly-typed 5mo agoTo me it sounds like a tooling problem. OP seems to be trying to use probabilistic text systems as if they enforce rules, but rule enforcement should really live outside the model. My sense is that there was a failure to verify the agent's intent. The tooling that invokes the model should really define some kind of guardrails. I feel like there's an analogy to be had here with the difference between an untyped program and a typed program. The typed program has external guardrails that get checked by an external system (the compiler's type checker).
- troupo 5mo agoWhat tooling? It's a probabilistic text generator that runs in a black box on the provider's server. What tooling will have which guardrails to make sure that these scattered markdown files are properly injected and used in the text generation?
- strongly-typed 5mo agoThat's the million dollar question. Maybe have systems of agents that all validate each other's work? Maybe something needs to be done at the harness level? I don't suppose that we could realistically expect 100% accuracy, but if we take 100% to be the upper limit, we could build systems that get us closer to that ideal.
- troupo 5mo agoThis is faith in magic. "There's some magic way to make probabilistic text generator running in the cloud to never miss local files"
- estimator7292 5mo agoThat's not learning.
- keeda 5mo agoAgreed, but the point is, if your system is resilient against an eager intern who has not had the necessary guidance, or an actively hostile disgruntled employee, that inherently restricts the harm an LLM can do. I'm not making the case that LLMs learn like people. I'm making the case that if your system is hardened against things people can do (which it should be, beyond a certain scale) it is also similarly hardened against LLMs. The big difference is that LLMs are probably a LOT more capable than either of those at overcoming barriers. Probably a good reason to harden systems even more.
- deleted 5mo ago[deleted]
- gpm 5mo agoThe difference makes the necessary barriers different. There's benefit to letting a human make and learn from (minor) mistakes. There is no such benefit accrued from the LLM because it is structurally unable to. There's the potential of malice, not just mistakes, from the human. If you carefully control the LLMs context there is no such potential for the LLM because it restarts from the same non-malicious state every context window. There's the potential of information leakage through the human, because they retain their memories when they go home at night, and when they quit and go to another job. You can carefully control the outputs of the LLM so there is simply no mechanism for information to leak. If a human is convinced to betray the company, you can punish the human, for whatever that's worth (I think quite a lot in some peoples opinion, not sure I agree). There is simply no way to punish an LLM - it isn't even clear what that would mean punishing. The weights file? The GPU that ran the weights file? And on the "controls" front (but unrelated to the above note about memory) LLMs are fundamentally only able to manipulate whatever computers you hook them up to, while people are agents in a physical world and able to go physically do all sorts of things without your assistance. The nature of the necessary controls end up being fundamentally different.
- Kim_Bruning 5mo agoA lot of 'agentic harnesses' actually do have limited memory functions these days. In the simplest form, the LLM can write to a file like memory.md or claude.md or agent.md , and this gets tacked on to their system prompt going forwards. This does help a bit at least. Rather more sophisticated Retrieval Augmented Generation (RAG) systems exist. At the moment it's very mixed bag, with some frameworks and harnesses giving very minimal memory, while others use hybrid vector/full text lookups, diverse data structures and more. It's like the cambrian explosion atm. Thing is, this is probabilistic, and the influence of these memories weakens as your context length grows. If you don't manage context properly, (and sometimes even when you think you do), the LLM can blow past in-context restraints, since they are not 100% binding. That's why you still need mechanical safeguards (eg. scoped credentials, isolated environments) underneath.
- nkrisc 5mo agoIt is merely a simulacrum of an intern or disgruntled employee or human. It might say things those people would say, and even do things they might do, but it has none of the same motivations. In fact, it does not have any motivation to call its own.
- AndrewDucker 5mo agoNo, because the safeguards should be appropriate to an LLM, not to a human. (The LLM might act like one of the humans above, but it will have other problematic behaviours too)
- keeda 5mo agoThat's fair, largely because an LLM is a lot more capable at overcoming restrictions, by hook or by crook as TFA shows. However, most systems today are not even resilient against what humans can do, so starting there would go a long way towards limiting what harms LLMs can do.
- rglullis 5mo agoAn eager intern can not be working for hundreds of millions of customers at the same time. An LLM can. A disgruntled employee will face consequences for their actions. No one at Anthropic, OpenAI, xAI, Google or Meta will be fired because their model deleted a production database from your company.
- altmanaltman 5mo agoit cannot go to the washroom and cry while pooping. And thats just one of the things that any human can do and AI cannot. So no it cannot do anything a human can do, the shared exmaple being one of them. And thats why we dont have AI washrooms because they are not alive or employees or have the need to excrete.
- root_axis 5mo agoIt doesn't follow logically that a human and an LLM are similar just because both are capable of deleting prod on accident.
- XenophileJKO 5mo agoI think you are more right than people are giving you credit for. I would love to see the full transcript to understand the emotional load of the conversation. Using instructions like "NEVER FUCKING GUESS!" probably increase the likelihood of the agent making a "mistake" that is destructive but defensible. The models have analogous structures, similar to human emotions. (https://www.anthropic.com/research/emotion-concepts-function https://www.anthropic.com/research/emotion-concepts-function) "Emotional" response is muted through fine-tuning, but it is still there and continued abuse or "unfair" interaction can unbalance an agents responses dramatically.
- gessha 5mo agoYou don't anthropomorphize a table saw, you just don't put your hand in there.
- not_kurt_godel 5mo agoFor those who might not know the reference: https://simonwillison.net/2024/Sep/17/bryan-cantrill/ https://simonwillison.net/2024/Sep/17/bryan-cantrill/: > Do not fall into the trap of anthropomorphizing Larry Ellison. You need to think of Larry Ellison the way you think of a lawnmower. You don’t anthropomorphize your lawnmower, the lawnmower just mows the lawn - you stick your hand in there and it’ll chop it off, the end. You don’t think "oh, the lawnmower hates me" – lawnmower doesn’t give a shit about you, lawnmower can’t hate you. Don’t anthropomorphize the lawnmower. Don’t fall into that trap about Oracle. > — Bryan Cantrill
- skeledrew 5mo ago404 on that link.
- dunder_cat 5mo agoA more direct source (possibly the original source?) I know of is a YouTube video entitled "LISA11 - Fork Yeah! The Rise and Development of illumos" which detailed how the Solaris operating system got freed from Oracle after the Sun acquisition. The whole hour talk is worth a watch, even when passively doing other stuff. It is a neat history of Solaris and its toolchain mixed with the inter-organizational politics. YouTube link: https://www.youtube.com/watch?v=-zRN7XLCRhc https://www.youtube.com/watch?v=-zRN7XLCRhc Direct link to lawnmower quotes (~38.5 minute mark): https://youtu.be/-zRN7XLCRhc&t=2307 https://youtu.be/-zRN7XLCRhc&t=2307
- not_kurt_godel 5mo agoWorks fine for me but maybe try https://web.archive.org/web/20260426213142/https://simonwillison.net/2024/Sep/17/bryan-cantrill/ https://web.archive.org/web/20260426213142/https://simonwill...
- theologic 5mo agoYou have no idea how thankful that you explained that. I watched the Cantrill video. As somebody that dealt this Oracle, it struck home.
- narrator 5mo agoIt's also important to realize that AI agents have no time preference. They could be reincarnated by alien archeologists a billion years from now and it would be the same as if a millisecond had passed. You, on the other hand, have to make payroll next week, and time is of the essence.
- hdndjsbbs 5mo agotaps the "don't anthropomorphize the LLM" sign They don't have time preference because they don't have intent or reasoning. They can't be "reincarnated" because they're not sentient, they're a series of weights for probable next tokens.
- coldtea 5mo agoThat is not that strong an argument as it seems, because we too might very well be "a series of weights for probable next tokens". The main difference is the training part and that it's always-on.
- nothinkjustai 5mo agoWe very obviously are not just a series of weights for probable next tokens. Like seriously, you can even ask an LLM and it will tell you our brains work differently to it, and that’s not even including the possibility that we have a soul or any other spiritual substrait.
- fc417fc802 5mo agoOur brains work differently, yes. What evidence do you have that our brains are not functionally equivalent to a series of weights being used to predict the next token? I'm not claiming that to be the case, merely pointing out that you don't appear to have a reasonable claim to the contrary. > not even including the possibility that we have a soul or any other spiritual substrait. If we're going to veer off into mysticism then the LLM discussion is also going to get a lot weirder. Perhaps we ought to stick to a materialist scientific approach?
- ignoramous 5mo agoRight. This line [0] from TFA tells me that the author needs to thoroughly recalibrate their mental model about "Agents" and the statistical nature of the underlying models. [0] "This is the agent on the record, in writing."