4 ms·
How would you recommend choosing the "best response" programmatically?
by caesil 3y ago
How would you recommend choosing the "best response" programmatically?
- throwup238 3y agoAsk each model to score and rank its own answer and each other. It's AI turtles all the way down.
- riku_iki 3y agoSomeone should build this as a service
- muzani 3y agoOnce they're faster and cheaper, it'll probably end up a standard pattern taught in school.
- nilsherzig 3y agoI think a lot of local LLM benchmarks are evaluated by gpt4 haha
- sorokod 3y agoIf you iteratively score, request improvement, and submit do results converge to a score value? If not, what do you think that means?
- brokensegue 3y agoThey aren't good at scoring their own work in my experience
- vineyardmike 3y agoOr just choose the best response manually as a human.
- patrickhogan1 3y agoThis is what I do today. I input the same prompt across all 3 and gauge the output of the first response. Whichever assistant best “understands” what I want to accomplish, I choose that assistant to continue the follow up prompts with. There is a bias where my lack of prompting technique may be the cause of the assistant not providing the best response. But, im grading on a fair curve since they all have the same input and I see this as the core value proposition of the assistant.
- patrickhogan1 3y agoYou wouldn't. As the originator of the prompt, the human user is the best judge of whether the prompt accurately captures their intent.