3 ms·
The library supports the model-graded factuality prompt used by OpenAI in their own evals. So, you can do automatic grading if you wish (using GPT 4 by default,
by typpo 3y ago
The library supports the model-graded factuality prompt used by OpenAI in their own evals. So, you can do automatic grading if you wish (using GPT 4 by default, or your preferred LLM).
Example here: https://promptfoo.dev/docs/guides/factuality-eval https://promptfoo.dev/docs/guides/factuality-eval
- westurner 3y agoOpenAI/evals > Building an eval: https://github.com/openai/evals/blob/main/docs/build-eval.md https://github.com/openai/evals/blob/main/docs/build-eval.md "Robustness of Model-Graded Evaluations and Automated Interpretability" (2023) https://www.lesswrong.com/posts/ZbjyCuqpwCMMND4fv/robustness-of-model-graded-evaluations-and-automated https://www.lesswrong.com/posts/ZbjyCuqpwCMMND4fv/robustness... : > The results inspire future work and should caution against unqualified trust in evaluations and automated interpretability. From https://news.ycombinator.com/item?id=37451534 https://news.ycombinator.com/item?id=37451534 : add'l benchmarks: TheoremQA, Legalbench