4 ms·
This demonstrates a need for more robust engineering practices. Typically you would solve this by testing a validation set against a new model. Some tools like
by deepsquirrelnet 2y ago
This demonstrates a need for more robust engineering practices. Typically you would solve this by testing a validation set against a new model. Some tools like DSPy or Agenta are helping to encourage this. But creating good evaluators for generative responses isn’t easy, and in my experience people tend to punt.
- rmbyrro 2y agoPrompting is different than programming. Auto-testing your prompts is good practice and helps during a migration. But it doesn't change the fact that the prompts won't work the same in a different model. It's like saying having Python tests eliminates bugs when rewriting a codebase into JavaScript.