3 ms·
> There is almost zero value in evaluating a prompt by only running it once. To the user But these tools are marketed as if you do only need to run them once
by ADeerAppeared 2y ago
> There is almost zero value in evaluating a prompt by only running it once.
To the user
But these tools are marketed as if you do only need to run them once to get a good result; The companies behind them would really want you to stop hammering the button that deletes their money.
As an aside:
> For evaluating prompts and running in production; your hallucination rate is inversely proportional to the number of times you sample.
This isn't really true, and requires you to fuzz the prompt itself for best effect. Making the "spam the LLM with requests" problem much worse.