4 ms·
RE: "Genie 2 is capable of remembering parts of the world that are no longer in view and then rendering them accurately when they become observable again." -- T
by psb217 2y ago
RE: "Genie 2 is capable of remembering parts of the world that are no longer in view and then rendering them accurately when they become observable again." -- This claim is almost certainly wildly misleading. This claim is technically true if there's any scenario where their agent, eg, briefly looked down at the ground and then back up at the sky and at least one of the clouds in the sky was the same as before looking down. However, I expect most people will interpret the claim far more broadly than the model can support. It's classic weasel wording.
- pfortuny 2y ago"remember parts of the world..." not even "some"... That is a tell-tale.
- isotypic 2y agoLooking at how no samples other than the 3 samples in the "Long horizon memory" section have any camera movement which puts something offscreen and then back onscreen, it certainly seems that they are stretching the capabilities as far as they can in writing.
- drusepth 2y agoYeah, my best guess is they're probably including the previous N frames as context into generating the next model. This works to preserve continuity over a short amount of time (as you say, briefly looking at the ground and then back up), but only a short period of time. For these kinds of models to be "playable" by humans (and, I'd argue, most fledgling AI agents), the world state needs to be encoded in the context, not just a visual representation of what the player most recently saw.