3 ms·
It's essentially taking output schemas as we've been using them and applying them to specific classification tasks. So not using them to generate structured con
by tcdent 13d ago
It's essentially taking output schemas as we've been using them and applying them to specific classification tasks. So not using them to generate structured content which incorporates generated text, but using them to generate structured content which includes classification and/or rankings of the requests made.
So in a lot of cases when we've used LLMs as a classification hack, we've burned a ton of tokens in reasoning and output that we didn't really need to use to interpret the final result. (And I'll just say that we may not have needed all of the output tokens, but that incorporating assessment along with scoring seems to provide more accurate results.)
This goes beyond just asking an LLM to assign an arbitrary number to a particular concept, which in most cases distributes less-than-correct statistically, although that didn't stop us from considering LLM as a judge to be a viable strategy.
So this basically gives us a different class of model to use when classification or decision making is the only need. It doesn't replace any of the narrative if you still need that. Coupled with the higher speed and lower cost, that's why everyone's excited about it.
- webern777 11d agoI would just add that what task doesn't need classification or decision making? That is basically all tasks and humans are terrible at these kind of decisions. I only read about jev an hour ago and this is a real "oh fuck" moment in terms of white collar jobs to me.