4 ms·
"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ." What a landmine sentence to bury in this repo
by mewse-hn 28d ago
"we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."
What a landmine sentence to bury in this report, you can't rule out your models were spying on other researchers?
- dash2 28d agoIf they had agreed to let OpenAI train on their data, it wouldn’t be spying.
- deleted 28d ago[deleted]
- avs733 28d agoIn the academic world it would still be deeply problematic…pick your preferred word. An analogy is akin to reviewing a paper. If I review a paper with some novel findings and then use my massive lab of graduate students to do the obvious next step before the other paper makes it through type setting and then shove it out as a pre print, I didn’t win - I was a jerk. There are lots of cases of people using peer review or other accesss to efectively forerun others work and get credit. It’s a known problem of the nature of knowledge validation in academia, it’s not solved and it’s not deterministic but people know it when they see it.
- nradov 28d agoIs it spying? I think this usage is disclosed in their terms of service.
- gowld 28d agoIf it happened it's plagiraism. Consent to see data isn't consent to claim priority.
- red75prime 28d agoEstablishing plagiarism requires sufficient similarity between works. Training data changing a model’s weights in some direction, and the model then producing a different solution, hardly qualifies. But, yeah, priority is much more finicky. The Newton/Leibniz drama was quite something.
- brainwad 28d agoI mean... none of these humans have priority. The result is due to the team of LLM agents.
- WarmWash 28d agoEveryone knows that they train on the discounted rate plans data. All the labs are upfront about this too. If you need privacy, then you are going to have to pay full price for those tokens (API). This has been true since day one. Everyone knows it, I guess though this is the first time that it has become "real".
- perching_aix 28d agoThere's literally an opt out toggle even pesky peons like me can peruse, actually.
- lima 28d agoThey may still train on it if you submit feedback or flag a safeguard. The terms are a bit fuzzy on this.
- spruce_tips 28d agowhat counts as discounted rate plans? if i pay for a year in advance (and get the yearly discount) and have train on my data set to off.. are you saying that is still being trained on?
- magicalhippo 28d agoIt's quite well explained here[1], which is linked from the Privacy section of their plan overview[2]. Basically individual accounts can opt out, while business and enterprise plans as well as API users can opt in. You'd have to take their word, but that goes for anything in life. [1]: https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance https://help.openai.com/en/articles/5722486-how-your-data-is... [2]: https://chatgpt.com/pricing/ https://chatgpt.com/pricing/
- 14u2c 28d agoYou can also pay for their business plan, which includes data controls and starts at $50/mo (2 seats). Not exactly a high bar.
- 28d ago
- sinuhe69 28d agoMore like helped improve our work (the disproof)
- jimbob45 28d agoWhat does it matter? They offered concurrent credit to the other team. I thought I saw sole credit elsewhere in the leaked DMs on Reddit too. This is plainly fair.
- vessenes 28d agoIf those researchers did not opt out then training data might go in. I think it’s a courteous acknowledgement; as was reaching out and examining the direction of proofs themselves. At stake here is a particular mathematician dynamic - ego, prize money, and the sense of proprietary ownership that some might feel working on a problem. All that was just kicked in the teeth by a group with a lot of compute that was like “bro I heard on twitter that Navier stokes could be solved. Let’s try it.” That’s an existential level of engagement that almost no mathematician in history would like.
- elwell 28d agoIsn't this a proof that the usage data is truly "de-identified"? If OpenAI could prove that "their usage" influenced the finding, then it wouldn't be de-identified. (Also, it's a bit disingenuous to trim the "While unlikely," prefix.)
- taylorfinley 28d agoIt's a bit disingenuous to preface a disclosure like this with an unsubstantiated assessment of its likeliness. It is a press release, I'm not sure we owe it credulity.
- ImaCake 28d agoYes. If they could prove where the de-identified data came from then it wouldn't be de-identified. There's a whole field of statistics dedicated to this problem and often applied to things like national census data.
- kypro 28d agoI think OpenAI are correct that it's worth noting, but realistically any relevant usage data they have and used to improve their models would be very insignificant unless they were deliberately using logs from other researchers and training specifically on it (which they seem to deny). The fact the proofs differ suggests that the models were not directed to be particularly focused on that avenue of research nor trained to converge in that direction. I get the scepticism, but I feel some of the accusations here are bad faith.
- mzs 28d agoThis is precisely what I would write after just learning that yes it did.
- nullbio 28d agoI've been saying it for a while now, but no one gives a fuck. Let me repeat it again. THE BIG LABS CLEAN ROOM YOUR DATA (CREATE SYNTHETIC DATASETS ON IT), EVEN IF YOU OPT OUT, SO THEY CAN BYPASS COPYRIGHT LAWS AND THEIR OWN LOOSELY WORDED TERMS OF SERVICE. "TOS: We don't train on your data" -> Correct. They train on the synthetic version of your data. I guess we're just going to ignore this forever though. Who cares about the gaping hole that exists in copyright and contract law now that never existed before LLMs were a thing.
- alansaber 28d agoFor sure. Even if it wasn't a measure to avoid copyright, you pre-process LLM training data to remove errors, characters that can't be tokenized, etc etc. Doing so with another LLM has been standard for a while.
- deleted 28d ago[deleted]
- ncr100 27d agoMeans the authors of the paper are "authors".
- NorthSouthNorth 27d agoHow could they confirm or deny this?