3 ms·
>We prompt three top-performing LLMs: GPT3.5, GPT4, and Claude V1.3 to generate a story of similar length to each New Yorker story, based on the one-sentence pl
by caesil 3y ago
>We prompt three top-performing LLMs: GPT3.5, GPT4, and Claude V1.3 to generate a story of similar length to each New Yorker story, based on the one-sentence plot summary
This is kinda like resizing an image down to 100x100px and then asking AI to upscale it and comparing against the original. Of course the attempts to reverse lossy compression won't be as good, even if the system is capable of similar quality work when correctly prompted.
- vannevar 3y agoNot really. The point of the LLM exercise is to measure creativity. Upscaling is not a creative process, it's basically the opposite of creativity. The only reason to keep the prompt in the same vein as the original story is for the judges, to at least keep the premise in the same ballpark so they're comparing comparable stories.