5 ms·
I agree that it's not possible to have them do 13th century architectural style perfectly right now. But I believe it will be soon. The image/video models are i
by samplank2 2y ago
I agree that it's not possible to have them do 13th century architectural style perfectly right now. But I believe it will be soon. The image/video models are improving, but so are the reasoning models, and they can check for and fix anachronisms.
- lukev 2y agoI hope you're right. Are you aware of any image-gen models that apply chain-of-thought style reasoning (either agentic or via reinforcment learning to shape outputs?) For example, consider this imagery from today's challenge: https://firebasestorage.googleapis.com/v0/b/fastab-f08e9.appspot.com/o/timetravel%2F93fb81ef-002c-459d-9726-b1493e3631e0%2Fvideos%2Fvideo_2494c3e1-4511-4889-8913-c4be7c395e74.mp4?alt=media https://firebasestorage.googleapis.com/v0/b/fastab-f08e9.app... These are some incredible monoliths: if they were real, I feel like I would have heard about them? And if they did... that's so cool. But because it's AI generated, I have a very low confidence level that this ever existed at all. Which is sad.
- samplank2 2y agoNo, not aware of image models that do chain-of-thought reasoning. But there are vision models that do it, so you can have them review the generated images and iterate on the prompts.
- tralarpa 2y ago[Spoiler] I guess it's this: https://madainproject.com/northern_stelae_park https://madainproject.com/northern_stelae_park Which is funny, because the monoliths in the AI video look more eroded than the real ones today. This looked like a nice idea at first glance. At second glance, it's really bad because you have to assume that everything you see in these videos can be wrong or misleading.
- nl 2y agoReasoning models aren't needed for this. The loss function for the image models needs to take year into account. This is entirely possible, as the incredible accuracy[1] of non-generative picture location models (a very similar problem) shows. [1] https://paperswithcode.com/sota/image-based-localization-on-cvusa-1 https://paperswithcode.com/sota/image-based-localization-on-...
- littlestymaar 2y agoWhy not using img2vid starting from an historically accurate picture or painting?
- samplank2 2y agoThis does use img2vid but with AI generated images. Using real pictures or paintings could definitely be fun too.
- peishang 2y agoYou might look into era specific LoRas if they exist, and if not consider training a few to help better capture architectural detail from that specific time frame.
- samplank2 2y agogood idea! It would be fun to have a ton of LoRas for different places x eras