3 ms·
I've found that context is everything to getting consistently good output, and augmenting your prompts with known truths with SERPAPi, embeddings and a vector d
by celestialcheese 4y ago
I've found that context is everything to getting consistently good output, and augmenting your prompts with known truths with SERPAPi, embeddings and a vector db, brought really flakey results into the >90% accuracy threshold.
As an aside - does anyone have good tools or methods for testing and evaluating prompt quality over time? Like performance monitoring in the web space, but for prompt quality. The techniques to use LLMs as evaluation tools of themselves always has seemed flakey when I've tried it, I'd like to use a more grounded baseline.
For example, if you have a prompt that says "What is the weather today in {city}?", you can run it against a list of cities and expected outputs (using a lookup to some known truthful API). That way, when you make changes to the prompt, you can compare performance to a baseline.
- 30minAdayHN 4y agoWe personally came across the evaluation problem while building MakerDojo[1]. My current workflow is to run manually on 50 different test cases and compare the results against the previous versions of the prompt. This is extremely time consuming. And to be honest, I no longer test for every little change. Some more contect - as a way to support MakerDojo[1], we are building TryPromptly[2] - a tool to do better prompt management. In that tool, we are building the ability to create a test suite, run the test suite and compare the results. At least knowing for which test case the results varied and reviewing them would go a long way. Here is the format we are thinking: https://docs.google.com/spreadsheets/d/1kLBIb7W0jrY-IkNPqJsNeh8CdGcJw7E-1-fulKZJrv0/edit?usp=sharing https://docs.google.com/spreadsheets/d/1kLBIb7W0jrY-IkNPqJsN... In addition, we are about to launch after test suite is to have live A/B tests in the production based on user feedback. Users can upvote or downvote their satisfaction with the results and that will inform you which version of the prompt yielded better results. If you have other ideas on how to test them better, it would be super helpful to us. [1] MakerDojo - https://makerdojo.io https://makerdojo.io [2] TryPromptly - https://trypromptly.com https://trypromptly.com