8 ms·
What? Prompts are natural language, so I don't see how they can possibly become obsolete.
by devit 3y ago
What? Prompts are natural language, so I don't see how they can possibly become obsolete.
- electroly 3y agoThe performance of the model can be improved with tweaks to the prompt, but the tweaks end up being model-specific. This is why "prompt engineering" exists for productionized use cases instead of people just spitting words semi-randomly into a textbox. Your old prompts probably won't completely fail but they'll behave differently under a different model.
- SkyPuncher 3y agoThey're loosely natural language. There's a lot of tweaking and non-natural language that goes into them to get the exact results you expect.
- tyree731 3y agoThey can. Results change as you change your models, and results aren't always strictly better or worse, which is why testing gold-standard results with any prompt and model changes is so important for applications utilizing LLMs.
- ineedasername 3y agoI have to prompt engineer a lot more with 3.5 than 4. The way I asked questions and convey what I want tends to be much more structured with 3.5, in a less natural way than I can do with 4. Hopefully 4 would be even better at answering a structured prompt like that, but also maybe not: For quick questions with a short answer 3.5 will sometimes give a simpler answer than 4, but 3.5 is correct. 4 isn't necessarily wrong, but it sort of reads into the question a bit more, the the answer is less succinct, more caveats and nuances explained, etc. In examples like this even though both give a correct answer, the one from 4 may be undesirable. You don't want to have to read through an extra paraph to pick out the answer to your question. There's more frictions. Of course the above scenario is easily solved: Change your prompt to include "Be Brief", but that's exactly the argument-- the old prompt is at least in part obsolete and much changes to achieve functional equivalency in 4. And the you need to check for unanticipated changes to the answer that "be brief" would cause: maybe it would now be too brief! Maybe not, but you have to have some method of checking.
- cosmojg 3y agoRLHF and fine-tuning! While these methods make prompting more accessible and approachable to people unfamiliar with LLMs and otherwise expecting an omniscient chatbot, they make the underlying dynamics a lot more unstable. Personally, I prefer the untuned base models. In fact, I depend upon a set of high-quality prompts (none of which are questions or instructions) which perform similarly across different base models of different sizes (e.g., GPT-2-1.5B, code-davinci-002, LLaMA-65B, etc.) but frequently break between different instruction-tuned models and different versions of the _same_ instruction-tuned model (I think Google's Flan-T5-XXL has been the only standout exception in my tests, consistently outperforming its corresponding base model, and although it's not saying much, I admit that GPT-4 does do a lot better than GPT-3.5-turbo in remaining consistent across updates).
- TeMPOraL 3y agoPrompts are natural language, but you're using them with the model in a way similar to getting a split-second gut feel reaction from a human - that reaction may very well vary between people.
- sp332 3y agoPrompts read like natural language, but you can’t always write them the way you’d write for a human. Here are some concrete examples of semantically similar prompts being interpreted quite differently by an LLM. https://twitter.com/mitchellh/status/1645562198935347205 https://twitter.com/mitchellh/status/1645562198935347205 And these chaotic, butterfly-effect areas are going to be different for different models, which is what prompted (lol) the original question.