3 ms·
maybe someone more informed can help me understand why they didn't compared to Llava (https://llava-vl.github.io/ https://llava-vl.github.io/)?
by tracyhenry 3y ago
maybe someone more informed can help me understand why they didn't compared to Llava (https://llava-vl.github.io/ https://llava-vl.github.io/)?
- dartos 3y agoMaybe they just didn’t know about llava while conducting their research. It can take days to train a model sometimes.
- buildbot 3y agoWeeks to months at larger scales even.
- kolja005 3y agoThe purpose of this research is to compare large vision-language models where the vision component is pre-trained using different techniques, namely on image classification versus unsupervised contrastive pre-training (see OpenAI's CLIP). PaLI-3 also isn't an instruction-tuned model, so comparing it to Llava would be a little apples-to-oranges.