3 ms·
When people worry about OpenAI stealing their chats and reproducing them elsewhere, I usually view the situation as unlikely - since chats are "trained" upon an
by instagraham 25d ago
When people worry about OpenAI stealing their chats and reproducing them elsewhere, I usually view the situation as unlikely - since chats are "trained" upon and not necessarily reproduced verbatim, you can assume that unless your chats depict a foundationally new and effective style of communication or ideation, there would be little need or use thereof of training on your chats.
For eg: "Hey ChatGPT my name is X and I am 6 and a half feet tall. Am I anaemic?"
This is a query, and while it might suggest to an AI model that tall people may worry about iron deficiencies, it's not really necessary to include in training. The user may be tall or short, but the idea that one may randomly ask about anaemia is not exclusive to this dataset. At best, this chat is an example of linguistics, not anything else, and the models figured out how to write and answer such questions years ago. It is ignored in training.
But when your work involves solid complex and unique mathematical proofs, the data is suddenly worth training upon. If I understand it correctly, the LLM may view your approach as a brand new path to take to solve an otherwise intractable problem. Its reinforcement training emphasises that it should do this in order to improve. And since it leads to results - large internal teams likely flag the model that reached this stage, the model is rewarded and given compute and attention - it is a desireable outcome both for the model and for OpenAI.
OFC, OpenAI becoming an advertising company will suddenly have incentive to treat all data as valuable. But while they are a "we need to make headlines" company, it's more rational that they view these examples of data as more valuable than others.
I don't doubt that they trained on his chats. This seems like the ideal usecase for "mass surveillance but using training" as a sort of filter.
But even so, one wonders how the model differentiates. If the researcher entered proofs into ChatGPT every day that mentioned "strawberries", while no other math paper on the topic did so, does that mean their chats would be audited?
- coliveira 25d agoAt this point these models have been trained to recognize every important math and science result based on context. They can easily flag conversations concerning the top 100 open problems in mathematics and use them for their advancement.
- camel-cdr 25d agoThere also is an insentive to silently give prominent people (e.g. Linus) or reasearchers like this custom tuned system prompts or even more powerful models.
- defmacr0 25d agoAlso, if we just take "high-quality" input data, which these chats would certainly be classified as, then the models are more than large enough to memorize everything verbatim. Spitballing some numbers, research literature suggests that LLMs are optimally trained with around 20 training tokens per parameter (fairly confident on this figure), that a DNN parameter encodes around 4 bits of data (less confident here) and I found sources in the 1-4 bits of information per token range (least confident here). So, fairly conservatively I would estimate that a model has the capacity to fully memorize around 5% of its training data, presumably high-quality data is a lot less than that.
- 405error 24d agoIt would not be difficult to write a pipeline to remove 99% of low quality posts, especially about specific subjects. It would be very easy to identify accounts as researchers based on their chat logs.
- xpct 24d agoIn a way, I think training on historic chats is akin to caching computation results. The compute cost has already been paid, and we make future retrievals cheaper by encoding it directly in the model. Assuming the results included some external validation such as user's preference, compilation, lean, etc., I'm not sure whether this would lead to model collapse.