3 ms·
Are there established best practices for "engineering" prompts systematically, rather than through trial-and-error? Editing prompts is like playing whack-a-mol
by typpo 3y ago
Are there established best practices for "engineering" prompts systematically, rather than through trial-and-error?
Editing prompts is like playing whack-a-mole: once you clear an edge case, a new problem pops up elsewhere. I'd really like to be able to say, "this new prompt performs 20% better across all our test cases".
Because I haven't found a better way, I am building https://github.com/typpo/promptfoo https://github.com/typpo/promptfoo, a CLI that outputs a matrix view for quickly comparing outputs across multiple prompts, variables, and models. Good luck to everyone else out there tuning prompts :)
- nico 3y agoAmazing, so useful, thank you
- sitkack 3y agoSeems like you would want to apply some NLP to the prompts themselves Take the gradient of the prompt wrt adjectives, verbs, nouns, etc. I forget the technique, but they add garbage words to the prompt to effectively increase the temperature.
- thomasfromcdnjs 3y agoGreat work, the space needs some more tooling in this direction.
- tlarkworthy 3y agoI use observablehq notebooks so I have programming reactively attached. https://observablehq.com/@tomlarkworthy/colossal-cave-chatgpt-challange https://observablehq.com/@tomlarkworthy/colossal-cave-chatgp...
- ukuina 3y agoThank you so much for this, especially for allowing custom LLM calls to allow testing of local models.