3 ms·
Experienced similar between 5.4-mini vs 5.6-luna in our own pipelines but after spending some time on prompt optimization and testing out various reasoning effo
by gbnwl 2mo ago
Experienced similar between 5.4-mini vs 5.6-luna in our own pipelines but after spending some time on prompt optimization and testing out various reasoning effort levels 5.6-luna was well worth it. Did you just replace model selection while keeping everything else in place or spend some time on evaling with newer prompts etc?
- baalimago 2mo agoNo we kept prompts as is, just swapped model. The prompt is already quite optimized for the task. How would updating it possibly make a more intelligent model spend less tokens than a less intelligent model? Care to elaborate?
- Tankenstein 2mo agoMost of the time when upgrading models we have needed to change prompts to get the same performance (let alone better performance). Usually, your prompt is overfit to the specific model doing the specific task. For example often your previous prompt is overspecifying and creating contradictions that a dumber model would just gloss over whereas a smarter model will try even harder to follow.
- steveklabnik 2mo agoHere is an example of a guide from OpenAI on how you should prompt 5.6 differently than their previous models. https://developers.openai.com/api/docs/guides/latest-model#prompting-best-practices https://developers.openai.com/api/docs/guides/latest-model#p...
- gbnwl 2mo agoI think the fundamental difference between our assumptions is you believe prompts to be optimized for tasks rather than model-task pairs. The only elaboration I can give you is empirical observations and model providers own guidance (as someone has already linked here). I'm pretty sure you probably have specific parts of your prompts that came about due to specific failure modes observed in your evals of running the task against first model. These vary across models in my experience, and it's always worth redoing this calibration process.
- weird-eye-issue 2mo agoYikes, you can't really expect prompts to just be model agnostic