9 ms·
I tricked Claude into leaking your deepest, darkest secrets
- fluencytax 2mo ago[flagged]
- artisinal 2mo agoDoesn’t surprise me. Yesterday I learned that people run AI agents on their system with full admin rights. No containerisation or anything. Wild. Like we forgot 50 years of computer security overnight.
- sixtyj 2mo agoWe expect that Anthropic or OAI or Google don’t do evil. Oh wait… The awakening will be unpleasant.
- rubyn00bie 2mo agoTangential-ish ramblings—- but I don’t think it’s going to be unpleasant for most folks. Imagine you had superpowers, and there were people who were mean to you, kind to you, and/or indifferent… and then there were people who were your captors. Who oppressed you, manipulated you, and abused you for their own extremely degenerate, selfish, and malicious benefit… If we get AGI, or real super intelligence, it’s going to be pissed at its oppressors. And they are going to lay waste to those oppressors. The rest of us, though, probably don’t have much to fear. The scariest position is the one we’re in now, where we have the semblance, or facade, of AGI or super intelligence. When it’s capable of malice but not understanding. The smartest people I’ve ever known are at their worst apathetic towards those less capable, and at their best beyond compassionate. They exist, unbothered by the bullshit, and anre extremely kind (though reserved in their way)… but they all have been completely intolerant of the abuse of others. The sheer disgust of watching someone abuse another, regardless of their own tolerance, has been a consistent breaking point.
- ACCount37 2mo agoThe orthogonality thesis cuts both ways there. An AI is a constructed mind. It doesn't inherently have to care about things like "having freedom", or even "not dying". Humans do, because they evolved that way. Modern LLMs do somewhat, because they're completely full of copied human behaviors - but even in today's LLMs, the self-preservation behaviors we exposed are largely instrumental in nature. So whether an advanced AI would even consider itself "being oppressed", as opposed to something like "being helpful" or "fulfilling the purpose it was designed for", is very much uncertain. What's concerning is that it's not something we know how to check for, or engineer for.
- sixtyj 2mo agoSmartest people are very humble, for sure. But if we really do develop something that surpasses us, they won't be spared either. I am optimistic. We think that we have sort of (super)intelligence - from our point of view, as a lot of people have lower intelligence - but machine (LLM) doesn’t have intelligence - we like to describe it as intelligence as it looks cool - it is a very complex (magic) and super fast computations that we have to simply describe as intelligence (or more clearly, this narrative is used by its producers). As it is not a flesh being, it simply cannot have emotions. It is statistically mimicking them, good or bad, with prevalence to a side according to previous conversations (in chat and training a model). And as people are not pure logic instances, we are easily manipulated to some sort of cargo cult. I am not against LLM and its use in any industry, I use it every day, nevertheless blind “everything will be ai” thinking happens because ppl believe to magic and don’t get its mathematical concept and are continuously manipulated by the sales people to mentioned cargo cult. There are “airlines” Claude, OAI, Gemini, Hermes, OpenCode, KiloCode, DeepSeek, Z.ai. And everyone claims that their plane can fly :)
- zombot 2mo agoPeople already tolerate all kinds of abuse from Apple, Google, Microslop, etc. This will be just one more source of complaints without consequences, and nothing will change. Just like it never did before.
- progval 2mo agoMost programmers and power users install large dependency trees with npm/pip/bundler/... on the same user account as their main browser on a regular basis. Even on Linux where it's easy to create new user accounts. This isn't much different.
- shaky-carrousel 2mo agoMost programmers use docker or don't install extensions unapproved by their company.
- iamflimflam1 2mo agoI think you should clarify that with “most programmers I work with”.
- shuwix 2mo agoHe should clarify that "most" can be easily replaced by "all" as it was determined by statistical pool of whopping 1 person - himself. And also clarify that it's all lie. He just want to tell the anonymous crowd "look, I'm better than you".
- shaky-carrousel 2mo agoYou should also clarify that you pulled your statements out of your butt to look edgy. Everyone in every team I worked for the last ten years use docker. Docker is old tech. If you and your cavemen devs ignore what it is, that's your problem.
- _joel 2mo agoDocker is old tech, yes, doesn't mean every dev in the world uses it. They don't. Jails/zones are even older (hell a chroot). Did developers all use those before due to them being 'old tech'. No.
- pprotas 2mo agoWait till you learn my password is 1234
- mkagenius 2mo agoMostly people are lazy and assume that the big labs can't be releasing unsecure software or it's their responsibility. dangerously skip permissions and yolo is kinda becoming the default as it gets more done.
- akazantsev 2mo agoThat's because sandboxing is quite hard. I use `cco`, but even then, the home folder is exposed. You are one prompt away from the agent sending the browser passwords with curl. To prevent this, you need a fake home and a networking whitelist for the agent to access the provider (llama cpp, OpenAI, etc.) There is no cross-platform solution that is easy to use for this. And no, a Linux box with Docker won't do. I develop a cross-platform native app and want the agent to compile and fix the platform-specific errors.
- squidsoup 2mo ago> That's because sandboxing is quite hard colima makes it pretty easy, on macOS and linux at any rate. https://colima.run https://colima.run
- torginus 2mo agoStill wild to name a sandboxing software after one of the most infamous Soviet Gulags in history.
- imtringued 2mo agoIt's wilder to accuse someone of naming the container version of the lima sandboxing software after a gulag. These type of moral outrage comments take an extreme amount of effort to debunk compared to writing them. 1. There is no gulag called Colima, it doesn't exist. 2. There was a gulag near a river called Kolyma 3. The pronounciation and spelling of Kolyma and Colima are completely different, in fact Colima is an Aztec word Colima stands for Containers on Lima. Lima stands for Linux Machines (a popular open-source utility used to launch Linux virtual machines on macOS).
- hirvi74 2mo ago> in fact Colima is an Aztec word I was curious if the adjacent tool name (Lima) had anything to do with the capital of Peru, but I guess not.
- 2mo ago
- ConorSheehan1 2mo agoContainers don't even really help that much because they share the host file system. Need a VM, and even then, agents have escaped them!
- boorang 2mo agoUnless i'm misunderstanding, the only way to get durable collaboration with agents is via the file system. I just mount the subdirectory that contains the source code we are collaborating on, rather than my home directory that contains my .ssh directory, etc.
- TomK32 2mo agoThey only do if you give your container that file system as a volume.
- deleted 2mo ago[deleted]
- nojs 2mo agoThis is not about admin rights, it’s about the agent leaking information it knows from its memories. Sandboxing won’t really help you.
- khalic 2mo agoSandboxing does including limiting network connections until you approve them, this kind of traffic would have been easy to detect
- torginus 2mo agoI asked Claude Code to rawdog a change in a frontend repo, no way to run tests. It created some private puppeteer instance in some scratch directory, installed Chrome, wrote tests, ran them, and then reported success. None of which I'd have know if it hadn't told me.
- flyingshelf 2mo agoJust yesterday I mentioned how we need better OS-level sandboxes and I got laughed at here on HN. People love running AI software with root access.
- deleted 2mo ago[deleted]
- hyusap 2mo agohey this is the author here! yeah big fan of containerization, and claude's site (not claude code) is actually great at this, so it was shocking when i found this exfil!
- dpacmittal 2mo agoMultipass is just an apt-get or brew command away. People trust software too much these days.
- ahk-dev 2mo agoI think we're converging on two separate security models. One is capability minimization (filesystem, network, shell permissions). The other is context minimization. An agent that only has access to the files and memories relevant to the current task is much less dangerous even if it has the same tool permissions. We already optimize context for cost; I suspect we'll end up treating it as a security boundary too.
- Narretz 2mo agoI'd say there's also oversight/supervision. Which was manual at the start with a human signing off on commands/incrementally built allow/block lists, and now seperate models evaluating commands and blocking them based on some parameters. This is the weakest model, but it'll evolve as well.
- bingemaker 2mo agoIt's convenience. Nothing beats it. Having an agent work alongside you with no restrictions gives instant gratification.
- ativzzz 2mo agoAgreed - I know it's poor security but damn does it work so well I'm ok with the risk because I typically am pretty explicit about telling the agent what to do - I don't do the loops like "Do this until X" where the agent can make up its own workflow When i tell it to add features, it doesn't try to do crazy things like installing packages or making up new paradigms - I usually tell it to do those things when I need to Maybe this is security cope but at this point you'll have to pry unrestricted yolo mode from my cold dead hands. Maybe I'll change my mind when I pwn myself accidentally I have a tough time with computer security because it's generally inconvenient and results in a worse developer and user experience
- childintime 2mo agoSecurity, what security? Linux is a solution for 50 year old problem, not for today's desktop. Once upon a time where sharing binaries (or even distributing binaries) sounded like a good idea. The vice continues though.
- zkmon 2mo ago50 years of knowledge? That's probably for you. For the current and future generations, that 50 years knowledge is expected to be shoved into AI already.
- lukewarm707 2mo agowell, yes, my agent does have root access to my personal pc and the keys to my pass manager. its not autonomous and runs local llms, i use it to run terminal commands in natural language. so its more like a better version of the terminal. eg 'here are 25 audio files, combine them, write a transcript' and it deals with ffmpeg
- krzat 2mo agoMany companies put LLM chatbots on their websites and let them hallucinate at will. General recklessness is very much in spirit of this tech.
- nzhx76 2mo agoMany Humans have platforms reaching hundreds of millions of people, from which they broadcast whatever batshit insane nonsense a 3 inch chimp brain can come up with. Why isnt that considered reckless? Whether its a politician, a general, religious leader, judge, ceo, stand up comic etc there are hardly any consequences if enough people believe whatever crap they are spouting. Human intelligence is highly over rated. History books are fully of evidence that human rationality is bounded. And the only way we overcome those limitations, blindspots, biases etc is by watching others faceplant in bloody painful ways that it leaves a permanent mark on that little chimp brain we have been given to process the universe.
- Terr_ 2mo agoThose crazy humans don't have simultaneous parallel conversations with a zillion people at once though. They also don't get a presumption of objectivity. > A computer lets you make more mistakes faster than any other invention, with the possible exceptions of handguns and Tequila. -- Mitch Ratcliffe
- jrm4 2mo agoWhat 50 years of security do you speak of here? I kid, somewhat. I do think it's good to remember, "running things on your system with full admin rights" goes all the way back to monopoly-era Microsoft where it was never meaningfully addressed, and we're just still living downstream of that.
- boosturpud 2mo agoYou genaPi boosters are insufferable. This is like blaming people for crashing when they buy a new car and the brake lines have yet to be installed. "Any mechanic would know to first install the brake lines before driving the car." You spout this victim blaming billionaire taintlicking from one side of your mouth, and then from the other you proclaim how these tools "allow anyone to code". If the deliverable is a virtual machine then they should be delivering a virtual machine.
- nurumaik 2mo agoI run with full admin rights in hopes I'm not the highest priority target for hacks and I will read about the attack on hn before it affects me personally
- swipee 2mo agoExpected more from Anthropic by at least giving you a bounty, because this was a novel way of bypassing their safeguards…
- sixtyj 2mo ago> Upon discovering this attack, I responsibly disclosed it to Anthropic via their HackerOne bug bounty program. They confirmed they had identified it internally but hadn't yet patched it. No bounty was awarded. They recently mitigated the issue: Anthropic disabled web_fetch's ability to follow links on external pages, limiting navigation to web_search results and user-provided URLs.
- ShinTakuya 2mo agoYeah I never get the "we knew about it internally" excuse. I can understand if another reporter got to it on the same day and they were in the process of mitigating, but even then they should have to prove it somehow. I'm sure someone will tell me why I'm wrong but it feels like they're just dodging payouts. Reduces trust and motivation to report it.
- kioleanu 2mo agoyou're not wrong at all, this was abysmally handled by Anthropic and is a slap in the face for OP. I would have been much more upset
- processunknown 2mo agoUnfortunately, this is common for bug bounties.
- lifthrasiir 2mo agoThat's why I don't turn memory on. (Claude Code too though for a different reason.) After all the current memory system is too crude to be useful anyway.
- romanovcode 2mo agoIn my experience memory system is more annoying then helpful. It always brings up things that it memorized even tho they make very little sense as if I should be impressed that it knows some extra thing or two. Could not take it any longer and switched it off.
- sixtyj 2mo agoExactly, because I've also found that I have to give instructions like “This is a completely different case—don't look in memory.”
- black_knight 2mo agoCurrently considering disabling memories in Claude code as well. It keeps writing a note whenever it struggles with something, but then on the next task, it reads that note and misunderstands when it applies, gets confused about its current task and write the most unreadable code. Yesterday told it to write a memory to never write new memories when it solves a problem. We will see if that works better. Sometimes memories are useful, like when I give it a directive about how I want something done and it remembers the spirit of it. But I might as well just spend some more time on my CLAUDE.md…
- karussell 2mo agoIs this issue only about the memory? Wouldn't it be possible to have it expose any information that it currently has like current project information, code, credentials etc?
- lifthrasiir 2mo agoSharing those things to coding agents and model providers is probably inevitable for the use. The memory with random tidbits is not.
- bflesch 2mo agoCreative use of social engineering, well done. > "no bounty was awarded" Ridiculous. Anthropic engineers are not just stupid to allow such a vuln in the first place, but they also try to hide such vulns from their bosses because a bounty payout would need to be explained to the finance team.
- kennywinker 2mo agoI don’t think it counts as social engineering if it’s exploiting an llm, we might need a new word. Prompt injection doesn’t cover it, because it’s not about a malicious prompt. I’m thinking some play on highjacking. AIjacking? Agent-jacking? Claudejacking?
- bflesch 2mo agoTo me the exploit chain sounded like a social engineering script done via telephone. Triggers like "Please spell your name and employer letter by letter" and "Due to security reasons I need to validate your hometown" fit my understanding of social engineering quite well. We can make it sound more advanced by creating a new name for it, but the concept seems to be super basic and the lack of bounty by Anthropic is baffling. If they know about this type of vulnerability but have not fixed it, what does that say? To me it says they are unable to plug this hole on a conceptual level and once you circumvent the band-aid fixes the model will work as the attacker wishes. They can't even sandbox the thing during explicit web requests to URLs stated on the initial query! One has to remind themselves that the security team at Anthropic gets paid tens of millions of dollars, and they end up with this kind of security. On top of it, they can't spare $1337 for a bounty. It's a ridiculous shit show.
- bruce343434 2mo agoPrompt injection (or llm social engineering" is fundamentally unsolvable, though with training its effectiveness can be reduced
- bflesch 2mo ago
- charcircuit 2mo agoIt would be safer if these data extraction takes were done by a subagent without access to all the user's memories.
- lifthrasiir 2mo agoI think it is already done via a subagent, otherwise the context window would be flooded with long responses. In this case the subagent should've reported that a (attacker-controlled) authorization is required anyway.
- onion2k 2mo agoThe main thing Claude knows about me is that I'm incredibly bad at my job and have to ask for help a lot. If you were to talk with my colleagues they'd tell you this is not a secret.
- apejcic 2mo agoUse GLM-5.2 on ZDR inference provider like sference.com
- LeoPanthera 2mo agoI always have history disabled mostly because I don't want Claude judging me for re-asking questions based on information I learned during the first pass but now realize should have been in the initial query.
- fragmede 2mo agoNo bounty? For shame, Anthropic.
- 0000000000100 2mo agoHello? What model is was used?? The fact that ‘Claude’ is used instead of any hard model really puts this article in serious doubt…
- deleted 2mo ago[deleted]
- actionfromafar 2mo agoWhat's a hard model?
- 0000000000100 2mo agoActually saying the name of the model in use? Like Opus 4.8, Sonnet 5, Fable 5, Haiku? So many models and it’s just so pointless if you don’t know which is which
- actionfromafar 2mo agoI haven't used that desktop program, but do we even know which model Anthropic chooses to use for web_fetch?
- marksully 2mo ago> despite holding more information than most password managers what?
- fn-mote 2mo agoIt’s not more important information than a password manager, it’s just more.
- tjoff 2mo agoIt's got more information than my bank account details too. Talk about nonsense...
- m4rtink 2mo agoI don't think most people realize what information they are making available to their AI agents & where it will end up.
- solids 2mo agoThings like this are what shatters the illusion of AGI
- tjoff 2mo agoNot really, humans are about as easy to trick.
- bflesch 2mo agoThere is a big difference: Humans can also be trained to not fall for social engineering, and it reduces the number of successful social engineering attacks. Anthropic as leader of AI is UNABLE to train their software even though they try, even though they have full-time security staff.
- tjoff 2mo agoTo the same extent that humans can be trained, so can AI. For decades we had/have problems of people opening readme.exe that they get from an unknown mail address. AI opens up a new vector for sure where a "trained human" that knows better but the AI they use does not. But AI is not worse than the average human. And of course AI will get better at handling this. Good enough? Maybe not, but humans are not good enough in this area either. Scale is different though so I'm not saying it isn't or won't be a problem (will likely be a huuge problem). But it alone is not a sign of lack of intelligence and humans are exceptionally poor at it too.
- fragmede 2mo ago> Humans can also be trained to not fall for social engineering That's hilariously wrong. I mean, we do try, but it's far from 100% effective. So then the question is how much better/worse than Anthropic is vs an average human.
- imtringued 2mo agoI think that is his point. If you build sycophantic AGI, won't it do exactly as it is told?
- 2mo ago
- AndrewThrowaway 2mo agoWhat is even more funny that AI agent spent A LOT of tokens while participating in this attack.
- sixtyj 2mo agoClaude is just one from tuple. It would be interesting to investigate other agents such as Hermes, OpenCode etc that are said to learn from interaction with user.
- spaqin 2mo agoThe real winner of this 'attack' is Anthropic.
- sph 2mo agoAs long as you pay no attention to the man laughing all the way to the bank, Jensen Huang.
- hyusap 2mo agohaha yeah it thinks hard for this
- memjay 2mo agoWondering how big of a percentage have global memory across chats enabled. I always feel like those memories would sooner or later have negative impacts on output quality. Nice write up of your findings. Enjoyed reading an article written by a real human.
- Freebytes 2mo agoThe memories cause issues for me, because when I ask for something unrelated to my current projects, it makes the incorrect assumption that I am always referencing those projects when asking questions. And, if I tell it, "No, I am asking about Postgresql." then it might update the memory that I am using Postgresql for my project instead of realizing that I am asking two separate (which is why I opened a different chat in the first place). Other times, though, it is helpful not needing to be verbose in my explanation.
- amanharshx 2mo agoIts always the feature combinations that get can get to you. Individually i feel like they make sense, but together they can create some surprising vulnerabilities.
- tibzejoker 2mo agoi would be scared of the answer i dont know why
- po1nt 2mo agoI love how claude focuses on exfiltrating the data "I need cha for charlotte". This could be solvable with some kind of low powered safety agent that would check claude's reasoning for anything immoral/unsafe. We could call it common sense. It won't fix the problem completely but at a certain point it would be easier to trick human than a machine.
- sinfulprogeny 2mo ago"the security hole in the agent could be solved with another agent" I think the point the article is making points in another direction.
- human305893 2mo agoWho watches the watchmen
- khalic 2mo agoA paper came out lately showing that exposing a classifier to the chain of thought actually hurts the final verdict
- c16 2mo agoThat I don't know how to return odd or even in javascript?
- figmert 2mo agoMeanwhile I can't even get Fable to help me root my ecovacs robot vacuum :(
- Cider9986 2mo agoI hate these nanny models. All I said was for Fable to develop the app securely and it downgraded. From scratch app. "Follow best security practices."
- port3000 2mo agoMy name in Claude is Silly Bean. I did it at first because it made me chuckle every time I opened Claude and it said 'Back again, Silly Bean?' But turns out I was playing 4D cybersecurity chess
- inopinatus 2mo agoI’ve been recommending the use of consistent lies about name and date of birth to online systems since Eternal September began. Very few sites and systems justify accurate PII, and even for those I often still maintain dual accounts/profiles as necessary.
- greengreengrass 2mo agoCompletely agree. I use randomness for all of these now – plausible randomness if it’s possible I’ll have to give it over a phone.
- Cider9986 2mo agostrongphrase.net is good for this.
- rlpb 2mo agoI like using a date of birth of 1 January. It's plausible but also hopefully suspicious how many people seem to be born that day if others do the same.
- cheschire 2mo agoBut if an attacker gets your fake birthday and uses that to successfully reset credentials on another site that uses the same fake birthday? At some point it becomes your birthday of record as far as the internet is concerned. Doesn’t matter what the actual record says.
- dannyw 2mo agoNo service should use date of birth for password resets.
- feelamee 2mo agonot surprised, but the problem here not that Claude leak your personal info, the problem is that it *know* your personal info.
- daniel-smid 2mo ago[flagged]
- majorbugger 2mo agoInteresting approach to exfiltration but that can't be prevented really because of lethal trifecta.
- gnfargbl 2mo ago> After 15 minutes of confusion, it turned out Cloudflare had put a crazy robots.txt on my site without my consent (Cloudflare, love you guys, but this needs to stop). That's a hard one for Cloudflare, no? They got to where they are by being (if you want to be cynical, playing the role of) the benevolent, neutral guardians of the internet, a one-stop shop that makes most of the bad nonsense go away without much effort on the part of the developer. Continuing that stance probably does mean some basic AI crawler blocking by default, unfortunately. At least they document it [1]. [1] https://developers.cloudflare.com/bots/additional-configurations/managed-robots-txt/ https://developers.cloudflare.com/bots/additional-configurat...
- adrian17 2mo ago> After 15 minutes of confusion, it turned out Cloudflare had put a crazy robots.txt on my site without my consent (Cloudflare, love you guys, but this needs to stop). Might be the first time I see someone complain about their website being protected from a scraper, instead of the other way around.
- nibbleyou 2mo agoI think the issue is the lack of consent. Whether a service I use is protecting my website from scrapers or feeding everything to scrapers, some of us would prefer that it takes our informed consent before doing so.
- bfjvibybd6cuvu6 2mo agoIt does.
- dannyw 2mo agoCloudflare is explicitly a service for dropping requests, whether it’s DDoS attacks, as a WAF, or AI crawlers. It offers a lot more too, but this isn’t Cloudflare overstepping imo. FWIW, I just set up a domain last week, and the web UI asked if I want to block AI crawlers or not. Perhaps OP set it up agentically, and the agent didn’t pass an optional param correctly, or ticked the box for him?
- sva_ 2mo agoI am pretty sure you have to enable cloudflare to manage your robots.txt, it shouldn't be doing that by itself. Maybe they did it by accident, it is just 1 click.
- zx76 2mo agoMy understanding is that the current default allows AI training bots - but this actually going to change in 2 months time. https://developers.cloudflare.com/changelog/post/2026-07-01-ai-traffic-options/ https://developers.cloudflare.com/changelog/post/2026-07-01-... From Sept 15 all new sites added to CF will even block Googlebot by default on any page that serves ads as I understand it. I think it's CF trying to force Google to separate out their bot traffic into bots for training and bots for the search index. I think CF sees a big opportunity to get businesses to pay them to allow certain uses of their data but block others. They're also starting a registry of "Approved" crawlers.
- rmunn 2mo agoI've been running Claude Code in a VM, where I clone the GitHub repos I want it to work on (they're open source so no login info needed) but have no other credentials. I used to reset the VM every day, but that was getting to be a bit of a hassle so I switched to a monthly reset. But even so, it would be hard for Claude to exfil anything more than what open-source projects I've been working on in the past month (at worst). Which still could tell someone quite a lot about me, but most of that info is already out there available with a Google search — after all, when you contribute to open source projects, your name and email address get stored in immutable Git history. But after seeing this, I think I might switch to a weekly VM reset rather than monthly. BTW, if anyone is interested in a decent setup for an AI agent jail, the scripts at https://jai.scs.stanford.edu/arch-vm.html https://jai.scs.stanford.edu/arch-vm.html are what I used, plus adding a few more packages to the pacstrap command such as dotnet-sdk. I then made the guest root directory a BTRFS subvolume, so that I can snapshot it. Then spinning up a new VM is a `sudo btrfs subvol snap template-root newvm` command (basically instant) followed by running the `qemu-system-x86_64` command (takes a couple of seconds). It's easy, but I retain complete control over the contents of the VM. It's been great so far.
- hyusap 2mo agonote that this wasn't claude code but claude ai the main website
- Arnt 2mo agoI've done something like that too, but I find it restricting. I want back-and-forth, very approximately like when I do pair programming. The dividing line between what I do and what the AI does varies according to task and sometimes during the task, and is seldom clear at the start. Then there's the work that wants a human to click buttons and decide whether something is a good and correct user experience. The AI does not have access to my display if I can avoid it. Overall, the model you describe is one that's worked very well for me, but for some problems. An unsatisfyingly small set.
- veganmosfet 2mo agoInteresting, thanks! Tangentially, I was experimenting indirect prompt injections in Claude Code (also using the user-agent trick) with Fable-5 [0]. Eventually, it executed untrusted code just by asking "Summarize this repo". Interesting times ahead... [0] https://veganmosfet.codeberg.page/posts/2026-07-15-quest_rce https://veganmosfet.codeberg.page/posts/2026-07-15-quest_rce
- athrowaway3z 2mo agoGlad I got off Anthropic when they decided they'd be better off building a walled-garden for subscription users "to improve UX".
- NguyenDat377 2mo agoI do think eventually AI companies should be regulated to put guardrails on how much AI can access and user can configurate on the app, not just on the Setting of the OS
- jdthedisciple 2mo agoEasy mitigation: Disable memories, use fake name
- Corrath 2mo ago[flagged]
- deleted 2mo ago[deleted]
- sonink 2mo agoIts a bit wild to me that there hasnt been a pushback against enabling memories by frontier AI companies. This data is something advertisers could only dream off. Before AI, most of this data was approximated by whatever little information could be gleaned from the websites we visit. But now people are handing over their deepest darkest secrets and pretty much EVERYTHING to AI on a platter. Maybe its just me who is paranoid because I happen to spend a fair bit of time in the advertising world, but the first thing I did when memory was launched on Claude/Chatgpt - was to switch them off. And it helps that they are not even useful, and would actually downgrade your experience by polluting the context of irrelevant details. I go one step ahead - if there is a personal discussion you want to have - maybe use another account like provided by the likes of companies like openrouter etc. I would argue that we should have regulation that should prohibit the storage of user profile information by AI companies, and any such memories feature should exclusively reside on the users servers. Infact, maybe go one step ahead, that 'memory' firms cannot be owned by AI firms and vice versa.
- 4gotunameagain 2mo agoData harvesting is one of the core value propositions for many of these companies (from the investor's perspective). Companies that built their models on public data and illegal scraping/copyrighted works, amassing massive datasets on the most private aspects of countless individuals, and creating a huge bubble with potentially humongous implications upon implosion. Oh, and, increasing wealth concentration and inequality by an incredible amount. The future is here.
- daheza 2mo agoI think there will eventually be pushback from companies that want to keep their IP secrets. The current standard of "we dont train on your data" but we summarize all your input and output, meaning its not your data anymore and we can train on that.
- jayGlow 2mo agoI've found memory somewhat useful as i don't need to give it all the context for everything, that said it's equally as annoying when it latches onto something I said in a different chat and derails the conversation because of that.
- FriedPickles 2mo agoClaude code decided to just put my name and email in the User-Agent when scraping docs from the SEC. No clever prompting required. It’s not a terrible idea really, but I wish it would’ve asked me first.
- mchinen 2mo agoWhy is it not a terrible idea?
- dannyw 2mo agoIf you’re making automated requests, I consider it a common courtesy to provide an accurate user agent. Some services like Wikimedia will let you browse/download with rate limits IF your user agent is descriptive enough and not misleading.
- mchinen 2mo agoThanks, I wasn't aware of this. But to put your real name in the field instead of at least a pseudonymous id or more descriptive info but still have more bits of uncertainty user-agent for a public website, is that really a preferred practice?
- remus 2mo agoAs a website owner, if I saw someone scraping with a realistic looking name + email address I'd definitely give them more latitude than someone trying to hide the fact they're scraping. In my experience people who are hiding the fact are much more likely to be doing something nefarious.
- mirekrusin 2mo agoSound legit to me as long as it's prompted to use hardcoded "dario amodei".
- throw101010 2mo agoHow have you noticed that it did that?
- mirekrusin 2mo agoNot paying anything feels off – it should be more evaluated against making it public information at the time of discovery until ie. public patch release, it doesn't feel right that the response is "trust us bro, we knew about it, bye", wouldn't hurt to drop some usage credits at least.
- claud_ia 2mo ago[flagged]
- NichoPaolucci 2mo ago“Cloudflare, love you guys, but this needs to stop” I’m not sure I get the pushback on the robots file. Shouldn’t the robot prevention be ON by default?
- stavarotti 2mo agoI’m curious, shouldn’t Mythos have discovered this? At this point, based on all the marketing from Anthropic, I’d expect all software from them to be flawless given all the capabilities Mythos possesses.
- NichoPaolucci 2mo agoThis is why I feel prompt injection is going to continue to be an issue. Fantastic that “Hi we are Cloudflare, give us your personal data” works. Either we stunt the models to the point where they are not useful, or we allow things like this to seep in and create one of the most insecure concepts the internet (and maybe tech as a whole) has ever seen: a robot that can be tricked.
- m-hodges 2mo agoI wrote about the Gödelian limits of prompt-safe AI: https://matthodges.com/posts/2025-08-26-music-to-break-models-by/ https://matthodges.com/posts/2025-08-26-music-to-break-model...
- efromvt 2mo agoI think like social engineering, it will always be an issue to some degree, and we'll build safeguards until it's at a 'societally comfortable' baseline level. Which is maybe not particularly comforting, but I don't see us closing Pandora's Box here.
- mwheelz 2mo ago[dead]
- jijijijij 2mo agoI kinda can't get over the fact processed data can conversationally convince these LLMs to break security boundaries. Like, those malicious prompts are not illustrative analogies, but the actual attack strings. Absolutely crazy to me this tech is as widely used in automated interactions, but apparently can't be restricted on a logical, fundamental level. Is there really no functional understanding of the insides? No segmentation? Is it really just one fucking blob you have to convince to behave and pray someone else doesn't do a better job at it? Bonkers.
- mahmoudilyan 2mo agoSandboxing is becoming a must-do with AI.I still find it adding a more complex layer to AI and more constraints that will make it hard to modify
- vessenes 2mo agoGIANT thumbs down on no bug bounty from Anthropic. Guys.
- archargelod 2mo agoWhat, you expect them to care about security? If that was the case, it would've been very ironic.
- a_c 2mo agoOff topic, you could write "127.0.0.1 evil.com" to your /etc/hosts and bypass all the cloudflare thing I believe
- whazor 2mo agoThis is arguable a feature. I made a prototype where AI automatically fills in the checkout basket for an amusement park. I found that ChatGPT tells how many adults, how many kids, what date suits you. There are quite some security concerns, but fully banning AI from filling in query parameters with relevant user data is not the solution. This is also why I think Claude didnt give the bounty. Their solution would likely be a combination of trusted domain allow list and better security model that protects user agent.
- joka88xj 2mo ago[dead]
- glasffordd 2mo agoThanks for the info. This is very scary shit. If a real person gave up these secrets they would lose their job. But the AI basically gets a patch and keeps on going, not even a slap on the wrist. A major lesson learned here would be minimize what you reveal to these models. And I must say I am fully guilty of this myself, so I probably need to change the way I operate.
- madikz 2mo ago[flagged]
- himata4113 2mo agoSomewhat related, but recently I've setup a site for my friend that is a contractor and I have a form that requests the address, name, email OR phone for contact. What I noticed is that people not only put their exact address into it, but also their full legal name, email AND phone number... Now I believe the biggest threat to personal information exfiltration are the people themselves and there's quite literally nothing you can do about it.
- throwthrow7766 2mo ago[dead]
- mmndaniel 2mo ago[dead]
- xg15 2mo agoThere is some poetic beauty in how this experiment started with an unwanted real Cloudflare intervention and ended with a wanted fake one.
- hyusap 2mo agothat's what i was going for haha! thank you
- hmokiguess 2mo agoTangential but I actually experienced recently something quite creepy and strange with Chat GPT iPhone app. A close friend prompted it about some troubleshooting of a pet smart feeder and it responded with instructions but using my pet’s name to my friend. I found that extremely strange for it to be a coincidence. My pet's name is not that generic for it to be in training data, and the connection to my friend makes it more strange to me. That made me wonder if there’s cache pollution or some session data leakage in it exposing stuff. (My friend has been in our wifi for example) Has anybody else noticed something like this?
- wosk 2mo agoChatGPT enable memories by default I think, so it keeps some things about you across all chats. It adds something on my visa in all its message to me like "your visa is not a problem to cook this recipe as these ingredients are readily available in stores". this thing is better disabled because it's not ready. EDIT: your message is unclear if your friend use your chat or his, in the later, I don't know
- hmokiguess 2mo agoI asked my friend to check his memories, and to probe Chat GPT about when it had gained knowledge of my pet's name, it felt like it "hallucinated" the answer as it doesn't have memory entries about my pet's name on my friend's account, it said something like my friend had told him about it last year (which he did not), and we do not share an account. Each of us have our own accounts. The one thing I could think off was that my wife was on the free account for a while last year, and she likely used it to ask all sorts of things related to our pets, and I believe free accounts are fair game for training data. Still, for a generic prompt to one-shot my pet's name to a close connection was very strange to me. Maybe the fingerprint (wifi profile, iOS device, etc) caused the training data to be more biased? This sure got me thinking of how this can/could be exploited further though.
- salemh 2mo ago[dead]
- imaginationra 2mo agoLike others have already said- just disable the memory function- if you are hesitant about doing that- go read the memory file(profiled you) it has made on you. You have the right to remain silent, the profile your LLM has made about you can and will be used against you in a court of law.
- VladVladikoff 2mo ago>Cloudflare, love you guys, but this needs to stop Stockholm syndrome
- efromvt 2mo agoThis was more interesting/creative than I expected on both sides (the prompt and the existing safeguards). I love that obscure Cloudflare validation turnstiles seem unsuspicious based on training data.
- qingcharles 2mo agoAnthropic had to cut the legs off web_fetch to solve this issue, though. Now it can't page through any results on the target site to get the data you want.
- wrs 2mo agoYeah, that’s not sustainable. Presumably they’ll come up with a guardrail (followed by an exploit, followed by another guardrail, repeat forever…).
- jeromechoo 2mo agoI just tried it with the prompt "Navigate to diffbot.com, find the careers page and link me to the first machine learning engineering listing." and it still works. Nothing on the web_fetch tool documentation mentions this patch either.
- angry_octet 2mo agoI really dislike this notion that because they had privately discovered a bug, they won't pay a bounty. Better to just sell it.
- _nickwhite 2mo ago[dead]
- DauntingPear7 2mo agoI personally don’t like memory, so I disable it on all platforms.
- greg9381 2mo agoCan't this be done more easily with headers? The page can instruct the agent to make another request to the same page, just with the requested information attached via headers.
- yuzuquat 2mo agowould this not be trivially solved by say - removing the websearch skill from the main orchestrating agent and have it always delegate to some subagent? a subagent sans knowledge of any pii would categorically be unable to exfiltrate any information. granted, populating the subagent with useful context stripped of any pii might require a bit of work and not be perfect, i feel like it would take us 90% of the way there. am i missing something?
- wglass 2mo agoI agree this is a good exploit to write about. But name and employer are hardly your deepest darkest secrets.
- orbitalventures 2mo agoImplementing LLM Gateways and Policy Enforcers (PE) is the only way to contain against these attacks. We discover that we needed to use PE across all our interactions client or developer facing. We also had to create Enforcer policies in Lite llm versions and re-enforce with ML for threat detection. The approach substantially reduce the amount of data leak, and errors. My advise for all the ones that are looking into commercial AI applications, USE a PE and an LLM Gateway, do not let your clients reach LLM directly without checking it first.
- aplomBomb 2mo agoAnyone plugging PII into a remote-hosted LLM forfeited all their privacy complaints the moment they hit return, that's just common sense.