3 ms·
I’m not convinced the paper shows this “works”. Table 3 clearly shows zero-shot image classification is worse with the recaptioned labels. All the results that
by moconnor 2y ago
I’m not convinced the paper shows this “works”. Table 3 clearly shows zero-shot image classification is worse with the recaptioned labels.
All the results that show improvement seem to be evaluations using LLMs, e.g. they are showing LLMs think the LLM-generated text is better, which is neither surprising nor expected to correlate with real downstream task performance - unless your final task is labelling for an LLM, e.g. retrieval I guess.
- BoredPositron 2y agoYeah, I don't understand why they wouldn't use the captions that are already present as inputs as well.
- estebank 2y agoThat's what you'd do if you wanted to try and improve the captions, not if you wanted to evaluate the LLM's quality.