3 ms·
The techniques in this article are good practice for general model tuning and testing with a correct answer. So for tasks like extraction, labelling, classific
by binarymax 3y ago
The techniques in this article are good practice for general model tuning and testing with a correct answer. So for tasks like extraction, labelling, classification, this is a great guide.
The challenge comes when the response is a subjective answer. Tasks like summarization, open question answering generation, search query/question/result generation, are the hard things to test. Those typically will need another manual step in the process to grade the success of each result, and then you need to worry about bias/subjectivity of your expert graders. So then you might need multiple graders and consensus metrics. In short it makes the process very very slow, expensive, and tedious.
- IsaacL 3y agoI pretty much agree. The "scientific" approach the author pushes for in the article -- running experiments with multiple similar prompts on problems where you desire a short specific answer, and then running a statistical analysis -- doesn't really make much sense for problems where you want a long, detailed answer. For things like creative writing, programming, summaries of historical events, producing basic analyses of countries/businesses/etc, I've found the incremental, trial-and-error approach to be best. For these problems, you have to expect that GPT will not reliably give you a perfect answer, and you will need to check and possibly edit its output. It can do a very good job at quickly generating multiple revisions, though. My favourite example was having GPT write some fictional stories from the point of view of different animals. The stories were very creative but sounded a bit repetitive. By giving it specific follow-up prompts ("revise the above to include a more diverse array of light and dark events; include concrete descriptions of sights, sounds, tastes, smells, textures and other tangible things" -- my actual prompts were a lot longer) the quality of the results went way up. This did not require a "scientific" approach but instead knowledge of what characterized good creative writing. Trying out variants of these prompts would not have been useful. Instead, it was clear that: - asking an initial prompt for background knowledge to set context - writing quite long prompts (for creative writing I saw better results with 2-3 paragraph prompts) - revising intelligently Consistently led to better results. On that note, this was the best resource I found for more complex prompting -- it details several techniques that you can "overlap" within one prompt: https://learnprompting.org/docs/intro https://learnprompting.org/docs/intro
- jimbokun 3y agoJust like it is with grading a student’s English class essay, for example.