9 ms·
Potential session/cache leakage between workspace instances or consumer accounts
- Tiberium 3mo agoSounds like a hallucination unless proven otherwise, even the leading LLMs can do those from time to time, and they will always appear plausible like that. Also could be the session having a lot previous context, like 800K+, which (I think) makes hallucinations more likely. Relevant comment from the OP which makes a hallucination more likely: > There is one tool call result that includes a string that printed a pathname including minecraft.py because it was listing the files in a Python virtual environment and the Pygments package has a lexer called minecraft.py
- xyzzy_plugh 3mo agoI don't disagree but this sort of thing has to be investigated regardless. It's unfortunate that there is so little transparency that even if they deny there was a leak we will never know for certain.
- macNchz 3mo agoThe person posting this claims to have reproduced in a separate context down the thread: > Same thing just happened on a Claude Mobile session in same Enterprise account. Common theme in both is Sonnet 5, first response after more than 5 minutes (cache miss).
- alserio 3mo agoWhy? what does make it more likely?
- andy99 3mo agoI realize hallucination has no precise definition but this doesn’t sound at all like anything I’ve ever heard called hallucination. Hallucination is usually plausible wrong answers or made up info that ends up fitting the most likely response (like a manufactured citation) and comes from the way LLMs work at predicting tokens. This example demonstrates completely implausible output, it’s not something that fits with hallucination. All that said, it doesn’t require cross session leakage, it could just be training data or like those nightingale (probably the wrong bird*) data generations where they just prompt an LLM with nothing and it starts spitting out conversations. I see a bunch of downstream comments about caching, sounds like maybe there’s an error where it loads nothing instead of the cache and so starts spitting out random generations. * edit: it’s magpie. Worth looking at the concept, I’m not sure people realize they LLMs generate random conversations when prompted with nothing, this seems at least as likely as sessions leaking: https://github.com/magpie-align/magpie https://github.com/magpie-align/magpie
- solenoid0937 3mo agoOne of his tool results mentioned the word minecraft.py, and the response was about Minecraft. It's a hallucination.
- Aurornis 3mo agoThe word “hallucination” has become overloaded, but it general means an LLM producing some output that isn’t plausible or grounded. When you have a very long context session where the context includes “minecraft.py” it’s not hard to extrapolate that Minecraft may have ended up in one of the reasoning traces and that distraction snowballed until it appeared in the output. These effects are becoming more rare as the SOTA models are improving so much. If you spent a lot of time with earlier LLMs or you experiment with smaller, quantized local LLM models this type of thing happens very frequently. When you see it happen so much on a model you’re running on your own hardware it becomes a reflex to chuckle and reset the session with a clean context. When it happens from a hosted provider it can be scarier because it’s not the type of failure mode most people are used to seeing.
- paulddraper 3mo agoExactly. If you've never had an LLM (all models) suddenly start spouting nonsense in a completely different language...you haven't been using LLMs that much. They will go absolutely insane some % of the time.
- andy99 3mo agoWorth looking at https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues https://www.anthropic.com/engineering/a-postmortem-of-three-... They can “go insane” but it seems often to be infra related as opposed to anything one would consider hallucination. Smaller models will often get stuck repeating a word or phrase forever but that’s a bit different and nobody would call it hallucination.
- ambicapter 3mo agoOne annoying one is we have an LLM-as-a-judge that is supposed to quote parts of a transcript to justify its reasoning, and sometimes it’ll get stuck on something short like “No.” and then just endlessly repeat it: “4. No. 5. No. […] 728. No. […] 1435. No. …”
- unknownfuture 3mo agoI've used LLMs all day five days a week plus my own free time for the last year or so (new job). I've seen plenty of hallucinations and context collapse behaviours. I've never seen that.
- shepherdjerred 3mo ago
- prima-facie 3mo ago[dead]
- acepl 3mo agoOh yes, we do not need programmers any more…
- emehex 3mo ago"Coding is largely solved"
- techpression 3mo agoI love that quote, especially considering the insane amount of bugs that are produced. It’s as easy to debunk as someone claiming ”I can jump to the moon”.
- consp 3mo agoWhile abused by LLM vendors, that phrase in one form or another I've been hearing since the early '00s and it's likely way older.
- ethagnawl 3mo agoSure but have you ever seen it actually play out in practice like it currently is? Whether or not it's true (of course it's not) people are currently behaving as if it is and firing/hiring accordingly.
- philipov 3mo agoWell, when was the last time you wrote machine code by hand? ... but then they went and changed what coding meant. We've always been layering abstractions on top of abstractions. If we get to an abstraction that works well enough that you no longer have to dive down into the previous layer, we say we've solved coding, and change what coding means. Obviously LLMs aren't there yet.
- supriyo-biswas 3mo agoThe funny thing is at my current employer, they mentioned that "coding is increasingly becoming a solved problem" and in the same breath, mentioned that one project was too hard for anyone to do so they're not doing it and would rather sell existing features...
- Avicebron 3mo agoIn order Fable 5 has rejected: "Recipe for red-braised pork, I have pork shoulder" "Write up a framework for MCP patterns I can give to claude code" "explain the biomechanics of motion in c. elegans" (I get this one, I mostly did it to test and it's related to my hobby project) Do we get an extra day of functional Fable 5 because it's down?
- HumanOstrich 3mo agoWhat does this have to do with anything? Who are you talking to? This is Hacker News, not Anthropic support.
- andy99 3mo agoNot sure the relevance of this comment, but normally if someone built a classifier that bad they’d be fired. Anthropic obviously thinks they have some monopoly power they can use to foist garbage on consumers, I think they don’t.
- ec109685 3mo agoCaching doesn’t work the way the bug reporter implies. Caches are shared (at least across the enterprise), but its key is always a function of the input before it. We achieved significant savings simply by moving everything that varies across individuals out of the system prompt so every session starts from a cache point. For example you never want your system prompt to start with the time that the session started. Move that to the first user message if needed.
- macNchz 3mo agoCaching is not supposed to work like that, but that doesn’t preclude the cache key computation function from having bugs.
- marginalia_nu 3mo agoYeah there's quite a lot of potential bugs that could have this shape. If I were to guess it could be a buffer in a buffer pool not being sized and zeroed correctly, allowing stale data to bleed between sessions.
- nok22kon 3mo agoor the cache retrieval function for a key retrieving the wrong entry
- supriyo-biswas 3mo agoThere could just also be a bug where the output tokens of session 1 were shared with session 2, due to a race condition or similar.
- Waterluvian 3mo agoThere is a massive incentive for optimization, so I expect they’re doing a ton of very clever tricks, all of which make this kind of bug more likely.
- estebarb 3mo agoHash functions necesarily have collisions. Also, it is perfectly possible to introduce bugs in the hash function (hash inputs, hash function itself) that allows cross account contamination.
- jstummbillig 3mo agoIs there anything particular about LLMs that would make separating customer data harder than in all SaaS cases?
- 27183 3mo agoIf I had to hazard a guess, doing anything in a multi-tenant way on a GPU is going to be hard mode compared to most SaaS due to lack of memory safe tooling. I've built multi-tenant SaaS systems, and I've done a little GPU programming (a long time ago), but I've never tried to combine the two disciplines.
- woadwarrior01 3mo agoIt'd be terribly compute inefficient to not share prefix caches (KV cache) across customers.
- acepl 3mo agoWhat is the probability that two customers will have exactly the same tokens in cache? Wouldnt it require using the exact same CLAUDE.md, skills, MCPs and context? After that it is even worse since the nondeterminism of LLMs and humans
- 27183 3mo agoI suspect what GP is getting at is there will be a strong incentive to implement some structural sharing across tenants to avoid redundantly storing the same tokens over and over. At least I'd be tempted to do this if I was working with a very precious, constrained resource (e.g. VRAM). Doing this correctly seems.. very difficult. [edit] To answer your question directly: the probability that the entire cache is identical between two different users is very low, but the probability that there exists identical chunks of cache between two different users is very high. Exploiting those commonalities successfully will significantly compress the data.
- weitendorf 3mo ago
- deleted 3mo ago[deleted]
- deleted 3mo ago[deleted]
- bix6 3mo agoSo the options are this amazing tech is so stupid it just randomly brings up Minecraft or it’s got a major security issue?
- 27183 3mo ago¿Por qué no los dos?
- paulddraper 3mo agoNot that different than people, amiright? --- Note that the author did have a minecraft.py file. So not quite 100% random.
- bee_rider 3mo agoIt’s the weekend so we’re allowed to anthropomorphize. I’ve known some brilliant engineers who would also just randomly bring up Minecraft (more likely Factorio these days) so this makes sense.
- Aurornis 3mo agoThe person had “minecraft.py” in their context and the session context was very long. Having an LLM session with very long context occasionally go off on a tangent is not uncommon. The people who expect absolute perfection out of every LLM interaction see this as some total indictment of the entire technology, but the people who use these tools daily have learned to treat the output as partially stochastic and to avoid extremely long context, even if the model offers it. It’s best to compact strategically or summarize next steps to hand off to a new session. Using sub-sessions can also reduce context pollution at the cost of additional token expenditure to summarize and transfer data to and from the sub-session.
- ShinyLeftPad 3mo agoTLDR: the first one > this amazing tech is so stupid it just randomly brings up Minecraft or it’s got a major security issue You can sugarcoat it but that's what it is. It's not slightly wrong like a junior engineer or weird like a junior engineer on LSD, it becomes like "your junior engineer suffered a stroke or sudden onset dementia completely forgetting the entire point". one trigger word and that's it we're building Minecraft castles now.
- Kapura 3mo agohappy fourth of july everybody!
- ofjcihen 3mo agoHappy fourth to you too :)
- ryantsuji 3mo agoNote the repro condition: first response after 5+ min, i.e. a cache miss. A cache leak would show up on hits (someone else's cached prefix), not on misses where everything is recomputed from your own tokens.
- ai_fry_ur_brain 3mo agoOpenrouters model providers give me urls people have given them quite frequently.
- dofm 3mo agoJust add a line in AGENTS.md that says "never talk about Minecraft unless you're explicitly asked", I'm sure it'll be fine after that.
- repeekad 3mo agoCLAUDE.md, Anthropic is too exclusive and next level to use a standard idiomatic pattern like AGENTS.md
- notnmeyer 3mo agoecho “read @AGENTS.md” > CLAUDE.md
- dofm 3mo agoYep that should work 100% of the time.
- folkrav 3mo agoWhen I still used Claude outside of work, my CLAUDE.md was just a symlink to my AGENTS.md.
- deleted 3mo ago[deleted]
- jasonjmcghee 3mo agoJust use a symbolic link
- pertymcpert 3mo agoProblem with that is that if the agent starts to browse the contents of the repo, it may read both AGENTS and CLAUDE.md.
- Frost1x 3mo agoI noticed you were linking a file vs creating a correct CLAUDE.md implementation. Would you like me to fix that for you?
- TZubiri 3mo ago0 evidence. If this were a real privacy leak, the author would ask their coworker if they talked about the unexpected topic instead of >"Maybe my coworker was talking about this in another session?" This would be a critical bug that would slash the market value of a T$ company significantly, go ask your coworker or close the ticket, why do you expect the devs to put an enormous amount of effort hunting a potentially inexistent if you can't make that minuscule debugging effort.
- bfeynman 3mo agofwiw, this could be a bug but the submitters level of arrogance places this rather high on the dunning-kruger side of things. There are multiple other plausible explanations, but this person is probably vibe coder who believes anything an llm says (including explaining its own hallucinations)
- dchest 3mo agoCan be malware? Something like https://news.ycombinator.com/item?id=48667495 https://news.ycombinator.com/item?id=48667495
- andy99 3mo agoInteresting to see the claudeslop reply as the first comment to the gh post and the reaction to it.
- mplappert 3mo agoSeems like a hallucination to me; note that the context contains “unmarkBlock” as the function name, which invites a connection to Minecraft. Still shouldn’t happen of course. The alternative explanation is that the inference engine, which batches several unrelated requests for parallel processing, messed up the unpacking and returned an unrelated user’s query. This one would be very scary as it will leak arbitrary content, but it seems much less likely here.
- dainiusse 3mo agoDon't worry. Mythos will fix that before release. Oh, wait...
- _def 3mo agoReminds me of a session I had recently (on web!) where claude insisted that i prefixed all my messages with statements about code execution or something, which was not the case. I interrogated it about that and it confirmed that it came from somewhere else, but could not get rid of it and each response mentioned that its gonna ignore those instructions. Eerie.
- andy99 3mo agoAnthropic injects text into the conversation triggered by certain conversation topics. This happened to me in relation to some red-teaming related discussion that was adjacent to something “sensitive”, I think sex, and Claude got confused about why I had said some kind of warning and mentioned it it’s response. After a back and forth it was clear that some extra warning to answer but avoid anything inappropriate had been inserted into the conversation.
- wongarsu 3mo agoClaude also sometimes mentions getting messages from classifiers, probably related to auto mode. Amusingly enough, when this happens to a subagent/fork, the orchestrator will call these " hallucinations by the subagent"
- deleted 3mo ago[deleted]
- throwaway260704 3mo agoUsing a throwaway account for obvious reasons, but I’m very involved in this space using LLMs from multiple providers. I’m aware of at least two instances in which the intermediate infrastructure “swapped” responses, once impacting Claude models and once impacting GPT models, from two different providers. One gave us a proper postmortem in which their API gateway was incorrectly handling HTTP 100 status codes, putting them into an error state where there was effectively an off by one error - you would receive the response to the prompt that came in before yours and would pay it forward (your response would go to the next caller). The other instance never had root cause explained to us, and we were just told to trust it wouldn’t happen again. Both of these are from $1T+ companies. ZDR wasn’t compromised in these cases since it was responses being swapped in flight. I wouldn’t be surprised if this is a similar issue - it’s not that data is being retained, it’s just not being safely isolated in intermediate infrastructure.
- theplumber 3mo ago[flagged]
- minhaz23 3mo agoCurious why you feel that way about Dario?
- solenoid0937 3mo agoHN thinks the safety crowd is dumb, and has never seriously engaged with the AI safety space. HN doesn't believe superintelligence will be a thing; while the AI safety crowd believes they are building it. So the decisionmaking of the safety crowd is incomprehensible to HN.
- DrewADesign 3mo agoReductionist. Many of us think they’re all dumb.
- pseudony 3mo ago
- noperator 3mo ago[dead]
- jonhohle 3mo agoI’ve been seeing this in Gemini in the past few days. Often during a prompt with a reasonably large input set, I’ll get answers that appear to belong to someone else. It may be trigger hallucination, but it seems like it may be cache collisions or something else. I’ve not seen anything to suggest private information is leaking, but it’s disconcerting to be researching something and then get what appears to be a math tutoring response.
- malfist 3mo agoMy whole company is doing mid year reviews and Gemini is the only allowed tool and its been flumoxing people with seemingly random unrelated responses. Often in different languages. That is when it bothers to respond instead of just sending back an 1099 error code
- weitendorf 3mo agoI’ve also had problems with Gemini when accessed through their UI in the past few weeks. That’s concerning that you are also seeing it several days later in a different context. I wonder if there could be a large security situation playing out behind the scenes right now. I’ve been working on using AI to assist me in writing meta parsing grammars. Fortunately I have not launched most of them yet. I know for a fact that the next generation of models represent a major step change in basic vulnerability identification and exploitation, especially if you know where to point them. They’ve found several bugs and at least one exploit in my parsing tools so far, I can’t imagine how many there still are waiting to be discovered across the entire modern tech ecosystem.
- solenoid0937 3mo ago> one tool call result that includes a string that printed a pathname including minecraft.py This seems like a hallucination.
- Trasmatta 3mo agoThe first reply clearly being a copy and paste from Claude made me want to vomit If people absolutely need to use AI to write replies, they NEED to start including a "everything after this was generated by AI" disclaimer
- mwnn 3mo agoI am facing a billing/subscription problem and there's nothing I can do or get help on. Their chatbot support shuts me down. Their email is also handled by the chatbot (not even sure whether it's the "same chatbot"). It has been a dead-end. I contacted my bank (credit card issuer) and finally a staffed said I am better off just marking the card lost and having it reissued and that's what I did in the end. I hope that works. I've never understood in what world this world decided it was okay to hand over these much unchecked power to such corporations. But this is how it has always been one way or the other.
- nullbio 3mo agoDon't worry guys, Anthropic are the experts at security and no one else should have access to bug fixing LLMs because that would be dangerous.
- jdw64 3mo agoThe biggest problem with AI agents is this. You can't debug what the AI is doing, so it's really hard to track down where something went wrong. What I know for sure: 1.Stuff that has nothing to do with the current session got mixed in. What guessing: 1.There's a minecraft.py file in the tool folder, and that might have triggered some hallucination. 2.Maybe data from some other project on the user's local machine got mixed in somehow. 3.Or it could be from another user's conversation. Honestly, if I think about how the system actually works, I don't think it's pulling from another user's data. But other people say they've had issues like that, so I can't completely rule it out. I saw this thing on YouTube once. When a bunch of users share the same system prompt, or prefix, the computation results get shared through something called a KV Cache. At least, that's what I understood. Not sure if I got it right. But if there's some bug in the hashmap that's supposed to keep those caches separate, then maybe multi-tenant memory management just broke down and that's what caused this. I mean, I can guess, but who knows. And honestly, even if that's exactly what happened, they'd never admit it. At the end of the day, LLMs are just word predictors, right? They build up some kind of semantic space inside. So maybe the user's question just happened to be near Minecraft in that space. That's kind of what I think.
- trq_ 3mo agoHi, it's Thariq from the Claude Code Team here. Thanks for the detailed report. We’re confident this is a hallucination but of course take these reports seriously and the team is looking into it. We’ll report back if anything turns up.
- jdw64 3mo agoI know it's the weekend, so thanks for working hard. Just a suggestion from a user: I wish we could manage Claude Code's memory more easily. Right now, when I go into the .claude folder and change a project folder name or something, sometimes it can't pull up the memory properly. It'd be nice if there were an easier way to import or export it. Thanks!
- MuffinFlavored 3mo agoPiggybacking onto a second/different "just a suggestion from a user": The VS Code extension needs love. I'm sure you guys are aware but it feels like it is neglected. GitHub Issues is a graveyard of 3-10 duplicates of really important issues with no activity getting closed. A few examples: * Lots of /commands missing in between the CLI harness and the VS Code extension * No way to monitor subagents/tasks/progress visually * No status bar/line
- deleted 3mo ago[deleted]
- codeduck 3mo agoAs a Butlerian, this is hilarious.
- impartshadow 3mo ago[flagged]
- ShinyLeftPad 3mo agoCan we acknowledge how it is sad that people get LLMs to basically play computer games for them. What's the point of fun?
- aberrahmane_b 3mo agoCould still be a hallucination, but the concerning part is that from the outside a hallucination, local context bleed, and an infra/routing bug can be very hard to distinguish.
- shard972 3mo ago[dead]
- cverinc 3mo ago[flagged]
- chris_explicare 3mo ago[flagged]
- beyondscaletech 3mo ago[flagged]