56 ms·
Yeah totally agree. We've found that a ton of OpenAI usage in practice is a variant of either classification or information extraction. This makes sense -- goin
by kcorbitt 3y ago
Yeah totally agree. We've found that a ton of OpenAI usage in practice is a variant of either classification or information extraction. This makes sense -- going from a human-native form of information (free text) to a computer-native form of information (structured data) is a key component of many pipelines!
Of course, GPT-4 is insanely expensive to use at scale, and still isn't a perfect classifier. So the next step is to take the outputs you get from GPT-4 and use them to fine-tune a smaller model that's really fast and good at your specific problem. In my experience, even without using any human annotations or online learning, a model fine-tuned just on GPT-4 outputs can actually outperform GPT-4 as a classifier! This seems really counterintuitive at first, but my guess is what's happening is that the training process is a kind of regularization, so the weird mistakes GPT-4 occasionally makes are overwhelmed in the training data by all the times when GPT-4 gets it right.
As a disclaimer, we're building open source tooling to ease the transition from prompt to cheaper fine-tuned model at my company OpenPipe.
- sharemywin 3y agoI wonder what the TOS is for a smaller model. I think it's not ok for training another LLM but what point does a Model become an LLM?
- gkbrk 3y agoApproximately nobody cares about the TOS of large language models sold for money that were trained by copying the content of everyone else with 0 compensation.