3 ms·
Synthetic data generation. You can have a really powerful, expensive model create evals so you can tune a faster, cheaper system with similar performance.
by serjester 2y ago
Synthetic data generation. You can have a really powerful, expensive model create evals so you can tune a faster, cheaper system with similar performance.
- jsheard 2y agoYou could do that, but OpenAI specifically doesn't want you to: https://openai.com/policies/row-terms-of-use/ https://openai.com/policies/row-terms-of-use/ What you cannot do. You may not use our Services for any illegal, harmful, or abusive activity. For example, you may not: Use Output to develop models that compete with OpenAI. Presumably you run the risk of getting banned if they realize what you're doing.
- serjester 2y agoSynthetic data is just as useful for building app layers evals. Probably significantly cheaper ways to get the data if you’re training your own model.
- echelon 2y agoScrew their TOS. OpenAI trained on the world's data. Data they didn't license. Anyone should be able to "rip them off" and copy their capabilities on the cheap.
- jsheard 2y agoThe irony isn't lost on me, but irony isn't going to stop them kicking you off their platform if they feel like it.
- littlestymaar 2y agoI suspect they have no way to enforce that without risking false positive hurting their rich customers (and their business).
- levocardia 2y agoI wonder if some of the high pricing is specifically an attempt to ward off this sort of "slow distillation" of a powerful model
- andyferris 2y ago> You may not use our Services for any illegal, harmful, or abusive activity. For example, you may not: Use Output to develop models that compete with OpenAI. This reads as if they consider developing models that compete with OpenAI as illegal, harmful or abusive. Which is crazy. (The other dot points in their list in the linked terms seem better).
- kelseyfrog 2y agoI compete with AI, not my models.
- SJC_Hacker 2y agoIf it was possible: 1) Why wasn't OpenAI doing it themselves? 2) This means we've reached technological singularity if AI models can improve themselves (as in getting a smarter model, not just compressing existing ones like Deepseek)
- sgillen 2y agoCompressing existing models is exactly what we are talking about here.
- SJC_Hacker 2y agoThen why wasn't OpenAI doing that themselves ?
- nickthegreek 2y agoWho says they arent?
- throwup238 2y agoIt’s not a singularity because the synthetic data generated by the previous frontier model isn’t usually fed directly into the training for the next frontier model - a “discriminator” is applied to select only the highest quality responses. That discriminator could be a field expert or mechanical turk or another model trained to select for higher quality responses (i.e. trained on a dataset of books rather than internet content). As far as I know, OpenAI has been doing this, using both experts and Kenyan workers as well as their own discriminator models. Unfiltered synthetic data is generally used more for distilled models and fine tunes for a specific use case.