3 ms·
Do you think it's possible an LLM might be able to evaluate if another LLM is better than itself?
by grensley 4y ago
Do you think it's possible an LLM might be able to evaluate if another LLM is better than itself?
- Centigonal 4y agoit can definitely evaluate whether an LLM is better aligned than itself (for some narrow definition of "aligned") - this is the motivating principle behind Anthropic's Constitutional AI idea. "Better" is pretty vague, but "more capable" would be difficult, because I don't think ChatGPT has a good idea of its own capabilities.
- geysersam 4y agoIt's possible that a chatbot can be one part of the evaluation process for new chatbots. I'd bet they already are. But they can't be the only part. New models must be evaluated on many different kinds of tasks. And if current models fails at a task, how can they evaluate new models on that task accurately?