4 ms·
iirc, they said they wouldn't do this? so this could just refer to secret scanning
by paddw 3y ago
iirc, they said they wouldn't do this? so this could just refer to secret scanning
- mtlmtlmtlmtl 3y agoClearly they lied, if their policy says otherwise.
- blackoil 3y agoPolicy says nothing about ChatGPT or LLM/AI models.
- deleted 3y ago[deleted]
- mikeryan 3y agoTheir policy, if you scroll up from this link, is to scan only “aggregate metadata” and only if you opt in. GitHub aggregates metadata and parses content patterns for the purposes of delivering generalized insights within the product. It uses data from public repositories, and also uses metadata and aggregate data from private repositories when a repository's owner has chosen to share the data with GitHub by enabling the dependency graph. If you enable the dependency graph for a private repository, then GitHub will perform read-only analysis of that specific private repository. If you enable data use for a private repository, we will continue to treat your private data, source code, or trade secrets as confidential and private consistent with our Terms of Service. The information we learn only comes from aggregated data. For more information, see "Managing data use settings for your private repository."
- voakbasda 3y agoWhat is AI training but “parsing content” for “delivering generalized insights”? They intentionally use slippery language that can defend their practices.
- hedora 3y ago> The information we learn only comes from aggregated data It seems pretty clear to me that this means they're allowed to use private repos to train copilot, etc. I wonder if any researchers have tried putting fingerprinted source code into a private repo, and then (after it is retrained) getting copilot to suggest stuff that could only have come from the injected supposedly-private source code. That would make a nice paper. I hope someone does it.
- simonw 3y agoI genuinely don't see how "The information we learn only comes from aggregated data" relates to training LLMs, which need raw data, not aggregated data, as their input. Maybe we have different definitions of the term "aggregated"? This suggests to me that GitHub need to extend that text to explain what they mean by "aggregated".
- atq2119 3y agoThe LLM itself is a form of aggregated data.
- simonw 3y agoSure, but the raw training data isn't. I think GitHub need to clarify this themselves.
- atq2119 3y agoSure. All aggregated data is ultimately derived from raw, unaggregated data. One can make the argument that training an LLM is "just" an unusually complicated form of aggregation. Whether that would hold up is another question. But yeah, I agree with the conclusion that they need to clarify this.
- hedora 3y agoTraining the LLM is a form of “learning”, and putting all the data in an input training set is a form of aggregation. The clause seems to mean “we can do whatever we want with your data as long as we violate many people’s privacy at scale at the same time”.
- deleted 3y ago[deleted]