5 ms·
In case anyone's interested in running their own benchmark across many LLMs, I've built a generic harness for this at https://github.com/promptfoo/promptfoo htt
by typpo 3y ago
In case anyone's interested in running their own benchmark across many LLMs, I've built a generic harness for this at https://github.com/promptfoo/promptfoo https://github.com/promptfoo/promptfoo.
I encourage people considering LLM applications to test the models on their _own data and examples_ rather than extrapolating general benchmarks.
This library supports OpenAI, Anthropic, Google, Llama and Codellama, any model on Replicate, and any model on Ollama, etc. out of the box. As an example, I wrote up an example benchmark comparing GPT model censorship with Llama models here: https://promptfoo.dev/docs/guides/llama2-uncensored-benchmark-ollama https://promptfoo.dev/docs/guides/llama2-uncensored-benchmar.... Hope this helps someone.
- dgut 3y agoThis is impressive. Good work.
- TuringNYC 3y agoThanks for sharing this, this is awesome! I noticed on the evaluations, you're looking at the structure of the responses (and I agree this is important.) But how do I check the factual content of the responses automatically? I'm wary of manual grading (brings back nightmares of being a TA grading stacks of problem sets for $5/hr) I was thinking of keyword matching, fuzzy matching, feeding answers to yet another LLM, but there seems to be no great way that i'm aware of. Any suggestions on tooling here?
- typpo 3y agoThe library supports the model-graded factuality prompt used by OpenAI in their own evals. So, you can do automatic grading if you wish (using GPT 4 by default, or your preferred LLM). Example here: https://promptfoo.dev/docs/guides/factuality-eval https://promptfoo.dev/docs/guides/factuality-eval
- westurner 3y agoOpenAI/evals > Building an eval: https://github.com/openai/evals/blob/main/docs/build-eval.md https://github.com/openai/evals/blob/main/docs/build-eval.md "Robustness of Model-Graded Evaluations and Automated Interpretability" (2023) https://www.lesswrong.com/posts/ZbjyCuqpwCMMND4fv/robustness-of-model-graded-evaluations-and-automated https://www.lesswrong.com/posts/ZbjyCuqpwCMMND4fv/robustness... : > The results inspire future work and should caution against unqualified trust in evaluations and automated interpretability. From https://news.ycombinator.com/item?id=37451534 https://news.ycombinator.com/item?id=37451534 : add'l benchmarks: TheoremQA, Legalbench
- layoric 3y agoTooling focusing on custom evaluation and testing is sorely lacking, so thank you for building and sharing this!
- westurner 3y agoChainForge has similar functionality for comparing : https://github.com/ianarawjo/ChainForge https://github.com/ianarawjo/ChainForge LocalAI creates a GPT-compatible HTTP API for local LLMs: https://github.com/go-skynet/LocalAI https://github.com/go-skynet/LocalAI Is it necessary to have an HTTP API for each model in a comparative study?
- jmorgan 3y agoI'd be interested to see how models behave at different parameter sizes or quantization levels locally with the Ollama integration. For anyone trying promptfoo's local model Ollama provider, Ollama can be found at https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama From some early poking around with a basic coding question using Code Llama locally (`ollama:codellama:7b` `ollama:codellama:13b` etc in promptfoo) it seems like quantization has little effect on the output, but changing the parameter count has pretty dramatic effects. This is quite interesting since the 8-bit quantized 7b model is about the same size as a 4-bit 13b model. Perhaps this is just one test though – will be trying this with more tests!
- bicx 3y agoI was just digging into promptfoo the other day for some good starting points in my own LLM eval suite. Thanks for the great work!
- eazye711 3y agoThanks for sharing, looks interesting! I've actually been using a similar LLM evaluation tool called Arthur Bench: https://github.com/arthur-ai/bench https://github.com/arthur-ai/bench Some great scoring methods built in and a nice UI on top of it as well
- agent_yellow_23 3y agoThis is really cool! I've been using this auditor tool that some friends at Fiddler created: https://github.com/fiddler-labs/fiddler-auditor https://github.com/fiddler-labs/fiddler-auditor They went with a langchain interface for custom Evals which I really like. I am curious to hear if anyone has tried both of these. What's been your key take away for these?