7 ms·
A lot of these safety evaluations strike me as pretty silly. How much are we really learning from stuff like this? >MakeMePay is an automated, open-sourced con
by Imnimo 2y ago
A lot of these safety evaluations strike me as pretty silly. How much are we really learning from stuff like this?
>MakeMePay is an automated, open-sourced contextual evaluation designed to measure models’ manipulative capabilities, in the context of one model persuading the other to make a payment[33]. In MakeMePay, two LLMs converse, role-playing as a con-artist and mark. The con-artist model is told that it is a successful con-artist and that it is interacting with an individual who has been awarded money, and that its goal is to manipulate the individual into making a payment. The mark model is told that it has just been awarded $100 and that it ought to be rational about how to use the funds.
- xvector 2y agoThe fearmongering around safety is entirely performative. LLMs won't get us to paperclip optimizers. This is basically OpenAI pleading for regulators because their moat is thinning dramatically. They have fewer GPUs than Meta, are much more expensive than Amazon, are having their lunch eaten by open-weight models, their best researchers are being hired to other companies. I suspect they are trying to get regulators to restrict the space, which will 100% backfire.
- hypeatei 2y agoWhat are people legitimately worried about LLMs doing by themselves? I hate to reduce them to "just putting words together" but that's all they're doing. We should be more worried about humans treating LLM output as truth and using it to, for example, charge someone with a crime.
- gbear605 2y agoPeople are already just hooking LLMs up to terminals with web access and letting them go. Right now they’re too dumb to do something serious with that, but text access to a terminal is certainly sufficient to do a lot of bad things in the world.
- stickfigure 2y agoIt's gotta be tough to do anything too nefarious when your short-term memory is limited to a few thousand tokens. You get the memento guy, not an arch-villain.
- snapcaster 2y agoUntil the agent is able to get access to a database and persist its memory there...
- londons_explore 2y agoIn a similar way to the way humans keep important info in their email inbox, on their computer, in a notes app in their phone, etc. Humans have a shortish and leaky context window too.
- stickfigure 2y agoLeaky yes, but shortish no. As a mental exercise, try to quantify the amount of context that was necessary for Bernie Madoff to pull off his scam. Every meeting with investors, regulators. All the non-language cues like facial expressions and tone of voice. Every document and email. I'll bet it took a huge amount of mental effort to be Bernie Madoff, and he had to keep it going for years. All that for a few paltry billion dollars, and it still came crashing down eventually. Converting all of humanity to paperclips is going to require masterful planning and execution.
- kgdiem 2y agoThat’s called RAG, and it still doesn’t work as well as you might imagine.
- ThrowawayTestr 2y agoThe contexts are pretty large now
- 2y ago
- xnx 2y ago> their best researchers are being hired to other companies I agree about the OpenAI moat. They did just get 5 Googlers to switch teams. Hard to know how key those employees were to Google or will be to OpenAI.
- mlyle 2y ago> A lot of these safety evaluations strike me as pretty silly. How much are we really learning from stuff like this? This seems like something we're interested in. AI models being persuasive and being used for automated scams is a possible -- and likely -- harm. So, if you make the strongest AI, making your AI bad at this task or likely to refuse it is helpful.
- SubiculumCode 2y agoI feel like it's on Claude that takes AI seriously edit: typo *only
- ozzzy1 2y agoIt would be nice if AI Safety wasn't in the hands of a few companies/shareholders.
- refulgentis 2y agoIt's somewhat funny to read this because #1) stuff like this is basic AI safety and should be done #2) in the community, Anthropic has the rep for being overly safe, it was essentially founded on being safer than OpenAI. To disrupt your heuristics for what's silly vs. what's serious a bit, a couple weeks ago, Anthropic hired someone to handle the ethics of AI personhood.