3 ms·
The current problem isn’t a lack of data. Meaningful data can be augmented with GPTs, just iterate over sparsely represented subjects and enrich them with conte
by inductive_magic 4y ago
The current problem isn’t a lack of data. Meaningful data can be augmented with GPTs, just iterate over sparsely represented subjects and enrich them with context-aware samples …generated with GPT.
We are actively leveraging GPT4s brainstorming-ability to generate datasets used for finetuning downstream. The fine tuned models downstream can then be used to augment new training data for more complex root language models, just formalize your quality assurance with common sense DSLs. We’ve reached an upward spiral with quite the exponential feel to it, where it will lead, nobody knows.
- throwaway1851 4y agoThat’s only going to get you so far. Sparsely represented subjects have have an actual reality you’re trying to model. So while yes, you can generate synthetic data to interpolate between the points that have already been sampled, whether those interpolated points have anything to do with the underlying reality is a matter of chance. If you really try to drill down the existing tools into a specialized problem area, it becomes very clear that they lack a sufficiently informed model of the subject to be useful. That’s why I’ve found that ChatGPT is great at the beginning of a project and increasingly irrelevant as you approach the core technical challenges of the project.