6 ms·
It is actually worse than that. It is at least 30 days. There is an "almost" that is doing a ton of heavy lifting here "deletion after 30 days in almost all cas
by pseudosavant 4mo ago
It is actually worse than that. It is at least 30 days. There is an "almost" that is doing a ton of heavy lifting here "deletion after 30 days in almost all cases". My read of that is they can hang onto data for as long as they want, even if they usually won't. And "all traffic" with an agentic harness is basically your entire codebase you work on.
> We will require 30-day retention for all traffic on Mythos-class models, on both first- and third-party surfaces. We won’t use this data to train new Claude models, or for any non-safety-related purpose, and we’ve instituted new privacy protections including logging all human access to the data and ensuring its deletion after 30 days in almost all cases (see this post for further details). The data will help us defend against complex and novel attacks (including new jailbreaks and attacks that operate across many requests) as well as help us identify and reduce false positives.
- bagels 4mo agoHow were they not already auditing access to customer data?
- codebje 4mo agoThey were not keeping it beyond the timeframe necessary for the model to process it, so there wasn't access there to audit.
- bethekidyouwant 4mo agoEven worse when you git push something Microsoft gets all your code!
- layer8 4mo agoOnly if you push it to GitHub.
- tcp_handshaker 4mo agoThat is why, for the last five years I have been checking in with them, code with some of the most atrocious quality. So far...its working....
- aurelius_44 4mo agoThe system works!
- vntok 4mo agoThank you for your service.
- dannyw 4mo agoYes, that is your intended purpose of “git push”, it’s to save. And only if you use GitHub. A better analogy here is probably “every time you use VS Code, the files you edit get sent to Microsoft”. Some legitimate concerns: • You have trade secrets. Previously; you can use services like Bedrock, etc, with signed contracts and significant reputations. Your contract is between AWS and you, and stays within your AWS security boundary. • Security breaches. Remember when Anthropic accidentally published the source tree of Claude code? Or Meta’s recent AI recovery bot that didn’t check if the supplied recovery email was actually the email of the Instagram account? The best way to reduce your exposure is to minimise storage. • Weaponised T&S. For example what if Anthropic decided to build a classifier for “usage in unsupported regions” that’s super overbearing (as we see with Fable) and vacuums up all context/input/output if there’s Mandarin? Contractually they could now retain it forever, not just 30 days, for ‘trust and safety purposes’ and perhaps have AI scan for any new or interesting ML techniques at scale, for Anthropic’s own use? They say just can’t train Claude models on the data.
- bethekidyouwant 4mo agoAll analogies are bad.
- GroksBarnacles 4mo agoAll models are wrong, but some are useful
- hexasquid 4mo agoUsing language to represent reality is lossy
- darkwater 4mo agoThe only one doing a very bad analogy in the thread it was you. You got a response with a counter analogy just to play on your same field and then a deep answer with real scenarios. You should respond to those, if you want to continue the discussion.
- deleted 4mo ago[deleted]
- OtomotO 4mo agoUhm, no? I have NO single project on Github. One of my clients has their project on GitHub. Every other client I have ever worked with or for ran and runs their own gitforge.
- tcp_handshaker 4mo agoHalf of my customers will drop them right away, and the other half, after I explain to them what this means.
- vntok 4mo agoYou must have very unrepresentative customers. What will they use?
- OtomotO 4mo agoNo AI at all, like 5/6 of my customers
- usef- 4mo agoIt's only for this model, not the one you're already using. And they're not training on the data. It's supposedly to detect abuse etc (such as someone retrying repeatedly with different variations to get around their protections)
- CorpOverreach 4mo agoStill unacceptable.
- gmerc 4mo agoYet
- Rekindle8090 4mo ago[dead]
- eth0up 4mo agoI cannot help wondering if the 'we won't train on your data' applies across the fence over there in pentagon land, where the classified contracts be. Yeah, of course they are not connected. Or.. Present user-llm activity is a goldmine of intel the agencies literally spent lives and billions on getting hardly close to, yet they elect to just let this one slip by.. Maybe. Really, I don't dispute it. But why? It's what, or precisely what, they always dreamed of.
- arcanemachiner 4mo agoWe've already gone through ECHELON, USAPATRIOT, TIA, PRISM, etc.. Either learn from the pattern and and plan accordingly, or be one of the credulous rubes caught off guard in the next wave of leaks.
- daveshistory 4mo agoI don't know why you'd read literally the last 25 years of leaks from mass surveillance programs and think for one moment that they've just, gosh, overlooked the opportunities.
- rapnie 4mo ago> We won’t use this data to train new Claude models, or for any non-safety-related purpose, and we’ve instituted new privacy protections including logging all human access to the data and ensuring its deletion after 30 days in almost all cases This reads to me as they can use any model that is not a "Claude model", and as for human access to that other model there can be different less restrictive privacy protections. In other words, that anything goes.
- eth0up 4mo agoYes. Words don't mean much these days. Taking corporate doublespeak at face value seems very couragious to me.
- bmitc 4mo agoDoes anyone know about the jailbreaks and attacks they are referring to? These are done through model queries?
- MichaelZuo 4mo agoWhy would you trust anything they say at face value? When they literally just showed you they are being deceptive by sneaking in the weasel word “almost”?
- alexjurkiewicz 4mo agoFirstly, none of this post is the contract people are signing. So it's merely a summary. Secondly, like all contracts I'm sure there will be exceptions for holding data longer than 30 days with reasonable cause, eg a legal hold.
- MichaelZuo 4mo agoThis reply does not make sense. I did not claim it was the literal contract people would sign?
- bmitc 4mo agoI'm asking for information to understand. What about that says I trust what they say as face value?
- deminature 4mo agoOne of the major attack vectors is distillation, where millions of questions are auto-generated and coordinated to produce training data for new LLMs. Anthropic alleges Minimax, Deepseek and Kimi were trained this way. Deepseek 4 compares favorably to Opus, so they're probably trying to prevent Deepseek 5 from being a bootleg Mythos. https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://www.anthropic.com/news/detecting-and-preventing-dist...
- pseudosavant 4mo ago
- kitchi 4mo agoThey seemed to have changed the wording since you posted the comment, now specifying exactly 30 days with seemingly no exceptions. These terms seem to be updated at-will, so I'll take that with a grain of salt however.
- Hamuko 4mo ago[dead]
- ryanisnan 4mo agoThat's strange. Even in my hobby-toy app, I have a TOS that I bump whenever the terms meaningfully change, and in my app, it forces a re-acceptance of the new terms before using the app again.
- SilverElfin 4mo agoWhere are you seeing that updated version?
- cornholio 4mo agoI'm not sure they can actually respect that 30 days absolute commitment. Let's say some internal tool flags a suspect conversation, it bubbles up and a human operator reads it and it looks like evidence of a crime. Then, that employee is legally bound in many jurisdictions to prevent the destruction of that piece of evidence. It's one thing to commit to a "everything is deleted when you press delete" automatic policy. It's quite another to say "we'll keep some stuff for up to 30 days, look inside it for any malfeasance, then pinky promise we'll delete it".
- SilverElfin 4mo agoIt’s even worse than that. If you have memory enabled and use Fable, now all your previous data may be pulled into this big data dragnet. How can Anthropic possibly think this is okay?
- daveshistory 4mo agoWell, it's okay for them.
- abustamam 4mo agoBecause they think people are okay with it, or at the very least, don't care, or don't care to know. Which, judging by how much people are using Fable, appears to be true.
- ithkuil 4mo agoAn interesting way to rate limit access while also getting some data to analyze. They will lift this restriction later when they have more capacity
- Forgeties79 4mo agoRemember when people were trying to pretend anthropic “were the good guys”?
- coldtea 4mo ago>How can Anthropic possibly think this is okay? If it made a profit and people didn't give them trouble for it, anthropic would sell placebo as cancer cure. What they think "is okay" is what they can get away with.
- deleted 4mo ago[deleted]
- nullbio 4mo ago"Even if they usually won't" is generous. I think they usually will, that's the point.
- daveshistory 4mo agoAfter 30 days and before the heat-death of the universe?
- mastermage 4mo agoI mean deleting the Universe also deletes the Data so that counts.
- daveshistory 4mo agoThat's a fair point.
- reinitctxoffset 4mo ago[dead]
- cakeface 4mo agoThe “all human access” is doing work also. Most access will likely be from AI agents.
- thefounder 4mo agoWhatever retention policy they have it will be honoured the same way they comply with DMCA laws(I.e if we’ve got it it’s ours to train/use)
- indoordin0saur 4mo agoAfter the AI companies just blatanty lying that they weren't hoovering up people's IP and art for training I assume they collect any and all data they can get their hands on for training. When it comes to the big AI players feeding their future models I 100% just assume that they suck up any data we send them. Am I cynical?
- mannanj 4mo agoand you can't opt out of data retention for non-training purposes. so I think theres a bit of a psyop occurring here.
- sebazzz 4mo ago> When it comes to the big AI players feeding their future models I 100% just assume that they suck up any data we send them. Am I cynical? There is a reason enterprise contracts and plans exist. And I think even on that account we're going to find out at some point that LLMs are training on that extremely useful data.
- devld 4mo agoI think it's very likely. This is the reason why I stay on GitHub Copilot business for the time being as a solo developer. I assume that Microsoft has less incentive than Antrophic to break the business agreement and use data for training or re-sell it to Antrophic. If I was using the heavily discounted subscription plan from Antrophic, I would 100% assume everything is fed to the machine. I'd rather pay whatever the API costs, than give it an exact recipe to build my product.
- mannanj 4mo agohowever dont all these AI companies retain your non-training data indefinitely? Did I miss something where they suddenly gave you the option to opt-out of retaining your non-training data? I thought that was a big money grab of theirs.