30 ms·
I strictly meant using the embeddings for training a reward model, not the generator. By good for generation I just meant the reward model might find more visua
by E-Reverance 1mo ago
I strictly meant using the embeddings for training a reward model, not the generator. By good for generation I just meant the reward model might find more visual cues for aesthetic preference and avoid some of the spurious semantic correlation CLIP has
- schopra909 1mo agoThat might work! Off the dome, it’s not clear to me whether spatial/depth priors are better/worse than an LLM for this type of task. Only reason I can think why the LLM might still work better here is that it’s trained to solve a bunch of different image/video related questions, so it’s perceptual modules may be more robust adaptive for this aesthetic grading task versus something like LingBot
- E-Reverance 1mo agoApologies for making the reply chain so long but I think a video like this somewhat proves how a lot of aesthetic preferences can be *ultra* sensitive to small visual details : https://youtu.be/twcMra_67-w?t=88 https://youtu.be/twcMra_67-w?t=88 The video is timestamped to open at the comparison frame. I don't think an LLM can tell the quality difference without direct reference for comparison