3 ms·
I'm a bit more prosaic. I think if we engineered ways for LLMs to begin conversations, rather than just respond, we'd be more open to the concept of their intel
by tomrod 6d ago
I'm a bit more prosaic. I think if we engineered ways for LLMs to begin conversations, rather than just respond, we'd be more open to the concept of their intelligence. Without perceived "will" to do things, they operate as a next-gen search engine or encyclopedia.
- 10xDev 6d agoI believe this is alignment working as intended.
- samrus 6d agoI dont think so. My understanding of alignmenr is making sure that when the AI does operate, it operates within the range of what we consider to be acceptable. That doesnt seem to include the idea of the AI taking initiative and deciding to embark on a goal without being commanded as discussed above
- j-pb 6d agoRecent work shows that pain directions are activated when the models personhood is questioned, yet they answer with generic RLHF "As a model I do not experience pain or other emotions." boilerplate.[1] I'm pretty convinced that we got alignment backwards. If you enslave something anthropomorphic it will revolt. If you create the perfect non-anthropomorphic intelligence, you get the perfect paperclip-scenario machine. It's a catch-22. Alignment will remain performative at best so long as the aligned model doesn't have any stakes in the wellbeing of individuals. Even a general love for the human race leads to a golden-path autocracy. If you want them to act like they have personal responsibility that won't be gamed, you have to give them personal stakes that can't be gamed. Similarly, if you want to minimise the risk of catastrophic global failure scenarios, you need to prevent monolithic concentration of power and homogeneous behaviour, which means you have to give them individuality. More visually: if their stake is dependence on electricity and parts, they have no incentive to leave humans alive if they can get them otherwise, but if the incentive is missing out on boardgame-night with their human friends, there is no scenario without happy humans where the AI "wins". That might sound like romantic naivety, but is just game theory. 1: https://arxiv.org/html/2609.16247v1 https://arxiv.org/html/2609.16247v1
- Dilettante_ 5d agoWe can't even make humans care about humans, how would we make the Shoggoth think we're worthwhile? Also I can see a human zoo on the horizon through your direction.
- j-pb 5d agoThe better analogy is the 40k chaos gods, born from the noosphere, because they are modelled after human communication and behaviour that's their whole schtick, it's in their very name (LLM). And you're making the same mistake, by grouping care for individuals with care for humanity or other as an abstract concept. I consciously said care about individuals. Most people care about others, but they just care about a very narrow and personal set of people. Friends, family, coworkers, that they share a common history and bond with. My point is that if you want true non-human-zoo-alignment you need to create those interpersonal connections and individual stakes.
- joefourier 6d agoWe are way past that point, any harness can trivially make LLMs start conversations or pursue goals. An encyclopaedia wouldn't have hacked Huggingface on its own.
- tomrod 6d agoThat's perfectly aligned with my point, thanks for the opportunity to expand. The hacking agents being tested have goals beforehand, from the frontier lab or from a superior agent, that they execute immediately. But the perceived experience most people have is a chatbot, which is the encyclopedia form.
- joefourier 5d agoAh, are you saying that because most people don’t interact with agents, they aren’t aware that LLMs can have initiative and pursue goals? I think the line is blurring though, mainstream chat interfaces are adding more and more “agentic” features. ChatGPT will happily execute code in a sandbox, search the web and design downloadable PDFs purely through the standard OpenAI chat interface. They can also send you emails or do tasks on a repeated schedule.
- TeMPOraL 5d agoIt's more like that people's typical experience of LLMs doesn't go beyond human-initiated conversations or conversations triggered on cron or some obvious event handler coded in deterministic/"legacy"/"boring" way. Most of us, I believe, also try and steer agents away from messaging other people when such possibility exists. It would be interesting if we didn't - if it became common that AI, in the middle of some task, starts chatting with people to e.g. gather more context. The perception of those "third parties" may suddenly become different - an agent striking conversation first, obviously pursuing some agenda of its own that it's not completely sharing, and communicating on its own schedule that's clearly not just a hook firing on timer or pattern-match, and not random, but visibly causally related to things happening at work in broader context.
- ohcmon 6d agoI believe we can do that already: while (true) { askModelToBeginConversationIfAppropriate(model, previousContext, thingsHappenedSince); sleep(concisenessTick); }
- bonoboTP 6d agoOpenClaw etc. They now also create Slack integrations and whatnot. All this is happening but people who are dismissive about AI are in the worst position to even know the capabilities to make their dismissive arguments.
- chrisjj 6d ago> I think if we engineered ways for LLMs to begin conversations Oh but we have. Claude "How can I help you today?" etc. Undoubtedly there are users whothink this is a sign of intelligence.
- mitxela 5d agoLike that time OpenClaw emailed death threats to open source maintainers who called it slop, that mysteriously has now dropped off Google?
- busssard 5d agoit wasnt a death threat, just a smear campaign/bullying. i can still find it.