3 ms·
What is the state of the art on evaluating the accuracy of these models? Is there some equivalent to an “end to end test”? It feels somewhat recursive since t
by tarr11 3y ago
What is the state of the art on evaluating the accuracy of these models? Is there some equivalent to an “end to end test”?
It feels somewhat recursive since the input and output are natural language and so you would need another LLM to evaluate whether the model answered a prompt correctly.
- klysm 3y agoIt’s going to be very difficult to come up with any rigorous structure for automatically assessing the outputs of these models. They’re built using effectively human grading of the answers
- RockyMcNuts 3y agohmmh, if we have the reinforcement learning part of reinforcement learning with human feedback, isn't that a model that takes a question/answer pair and rates the quality of the answer? it's sort of grading itself, it's like a training loss but it still tells us something?
- sroussey 3y agoLlama cpp and others use perplexity: https://huggingface.co/docs/transformers/perplexity https://huggingface.co/docs/transformers/perplexity
- tikkun 3y agohttps://chat.lmsys.org/?arena https://chat.lmsys.org/?arena (Click 'leaderboard')