4 ms·
This is a great summary of why productionizing LLMs is hard. I'm working on a couple LLM products, including one that's in production for >10 million users. T
by typpo 3y ago
This is a great summary of why productionizing LLMs is hard. I'm working on a couple LLM products, including one that's in production for >10 million users.
The lack of formal tooling for prompt engineering drives me bonkers, and it compounds the problems outlined in the article around correctness and chaining.
Then there are the hot takes on Twitter from people claiming prompt engineering will soon be obsolete, or people selling blind prompts without any quality metrics. It's surprisingly hard to get LLMs to do _exactly_ what you want.
I'm building an open-source framework for systematically measuring prompt quality [0], inspired by best practices for traditional engineering systems.
0. https://github.com/typpo/promptfoo https://github.com/typpo/promptfoo
- darkteflon 3y agoThis looks excellent, thank you - really nails the UI. Going to use this this week.
- jmccarthy 3y agoVery nice, thank you! Will give it a try.
- anotherpaulg 3y agoThis looks really useful. Any thoughts on managing costs? I've been developing against gpt-4, and it runs up charges quickly. I've been thinking I will need to be careful about adding live api calls in any sort of testing situations. Wondering if your tool has any features to help avoid/minimize wasted api usage?
- typpo 3y agoThe tool maintains an LRU cache on disk by default - which means repeat identical requests will be fetched from cache instead of the live API.
- chaxor 3y agoIf you have a task that requires something suggested by "__exact__", then a full LLM is probably not the answer anyway. Try distilling step by step, especially if the goal to to generate a DSL or some restricted language. It can be helpful to have a different set of tokens available to the model for decoding, such that the only possible outcome is something like 'ATTCGGTCCCGGG' given some question to predict a DNA sequence.
- typpo 3y agoFor sure. I'm dealing with fuzzier stuff, more in the sense of "don't refer to yourself as a chatbot", "this input should trigger X tool", and things of that nature.