3 ms·
I share many of your concerns and frustrations, although I suspect what you're asking for their consider a moat along the lines of a trade secret, rivaled only
by mcint 3y ago
I share many of your concerns and frustrations, although I suspect what you're asking for their consider a moat along the lines of a trade secret, rivaled only by the collection of performance improvement techniques they've amassed in 1000s-10,000s of training runs, 100s of engineers, and (hundreds of?) millions spent on compute. People are hired and praised in the community for their skill in cleaning data.
A non-answer for you, but for curious others, [State of GPT] 10 days ago provides a through introduction to the process used to train, Karpathy speaking at a Microsoft event providing a deep summary review of concepts, training phases, and techniques proving useful in the world of Generative Pretrained Transformers.
[State of GPT]: https://www.youtube.com/watch?v=bZQun8Y4L2A https://www.youtube.com/watch?v=bZQun8Y4L2A
- theptip 3y agoYou seem to be talking about architecture, Simon is discussing training datasets. Of course, the contents of those datasets is trade secret too, but Simon is not looking for the contents, or even the cleaning strategies, just the lineage; is OpenAI using my private data? I don’t care how, just whether they are or not.
- mcint 3y agoExcellent clarification. I suspect this is at a natural Schelling point: don't say much; because more answers would only lead to more questions. The trade secret aspect includes that lineage. It's an edge. Even in his post, only the first and second short sections hint in passing what's linage about training would be wanted or why. I'm not clear what people worry is at stake. I will hunt these comments on this question as well. What are the incremental concerns of AI literate people.
- theptip 3y agoTo expand upon his first paragraph: > People are worried that anything they say to ChatGPT could be memorized by it and spat out to other users. People are concerned that anything they store in a private repository on GitHub might be used as training data for future versions of Copilot Individuals care because if they disclose highly personal secrets or mundane private information (eg health status, PII like address, tastes/preferences) then those could easily be disclosed later. Companies care for more obvious/less speculative reasons; both the trade secret version of the individual concerns above, but also strategically, in that many companies don’t want to aid their competitors by training OpenAI/Copilot how to write code for their domain. (Obviously what you really want is to fine-tune GPT-4 on your code and be able to trust that they aren’t going to use that for training future models.) My gut feel is that they aren’t saying anything because they are doing grey area stuff pushing the boundaries of “fair use” and plan to ask forgiveness later, which will be easier to do when they have demonstrated massive levels of utility & everyone is benefitting from their model.
- godelski 3y agoI think the point more is about advertising masking as research. Yes, OpenAI did a lot of research and there's no question about that and the quality of it and their results. But the question is if *Open*AI is doing internal research (as is common to any big company) or academic/open research. Researchers have different goals and so want to probe these models and understand them. I think the confusion comes from proprietary work looking like academic research. It is nice to peek behind the curtain, but it is unclear what the utility is. I think a lot of the recent pushback is the social immune system going into effect. We just have to decide if this is an auto-immune disorder or not. The question comes to the advertisement-to-utility ratio. Have we crossed that threshold? What is the threshold? I think the immune response is happening because we don't know and our definition of that is as good as trying to define porn. I think there's a lot of confusion because we're not accurately codifying what the issues are. Or rather we all see different issues but are acting as if others have the same concerns; so we are talking on different pages.