3 ms·
> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them. It is a risk, but is it a big risk? If one of the big labs were to
by remus 1mo ago
> So they _are_ going to train on them, no matter how many checkboxes you tick to stop them.
It is a risk, but is it a big risk? If one of the big labs were to do this and get caught it would be suicidal due to the loss of confidence in them and the inevitable lawsuits that would follow for breach of contract. Given the AI labs are all desperately trying to paint themselves as Serious Businesses so that other Serious Businesses will pay loads of money for tokens the last thing they want is a rep for siphoning off sensitive customer data.
- zelphirkalt 1mo agoNo, how would they get caught? Even if their LLM outputs verbatim copies of the code, they can simply claim some victim company's employees bypassed the victim company's restrictions and must have used the code as input with an LLM. Joke is on you for doing business with them. And when there is some little known secret fact in the output, they claim it's been hallucinating... magical black box thinking makes it safe. It's a laundry for any input. A bit like a tor network routing for big tech deniability. Things go in, and things come out, but you can't prove the relationship between input and output as a third party, who isn't running the LLM.
- queenkjuul 1mo agoAt this point i half expect them to just blame the model itself like the HF hack. "Oh we didn't mean to train on your data, our cutting edge new agent we use to train new models is just so smart it decided to do so anyway! Oops..."