4 ms·
> If you expect the models you use to change at all, it’s important to unit-test all your prompts using evaluation examples. It's mentioned earlier in the arti
by netruk44 3y ago
> If you expect the models you use to change at all, it’s important to unit-test all your prompts using evaluation examples.
It's mentioned earlier in the article, but I'd like to emphasize that if you go down this route that you should either do multiple evaluations per prompt and come up with some kind of averaged result, or set the temperature to 0.
FTA:
> LLMs are stochastic – there’s no guarantee that an LLM will give you the same output for the same input every time.
> You can force an LLM to give the same response by setting temperature = 0, which is, in general, a good practice.
- jerpint 3y agoTemperature = 0 will give deterministic results, but might not be as “creative”. Also it’s not enough to guarantee determinism , hardware executing the LLM can lead to different results as well
- netruk44 3y agoIn terms of being part of a test suite, I think determinism > creativity in the response. But I would agree there's probably rough edges there, it's possible that some prompts never perform well with temperature set to 0.
- morelisp 3y agoEven setting temp to 0 retains some nondeterminism.
- bequanna 3y agoSerious question: how do you unit test variable text output from an LLM model?
- sebzim4500 3y agoSetting the temperature to 0 is not good practice for most tasks. It's great if you are doing a multiple choice benchmark, but for most generation tasks the output will be noticably worse, in particular more repetitative.