4 ms·
LoGeR – 3D reconstruction from extremely long videos (DeepMind, UC Berkeley)
- msuniverse2026 7mo agoTruly don't understand what is happening in the heads of these researchers. Can't they see how the main use of this is going to be mass surveillance?
- KeplerBoy 7mo agoThese seems to be much more robotics / autonomous vehicle focused? I don't quite see the mass surveillance angle you get from this you don't already get from cheap ubiquitous cameras, basic computer vision and networking (aka flock) .
- haritha-j 7mo agoI think you've made the erroneous assumption that the researchers care. I work in 3D reconstruction and I've not really seen too many people care about the actual use case, and indeed have had some friends join defence.
- imtringued 7mo agoI'm not sure what you mean. The input video feed already constitutes "surveillance". You'd need cameras everywhere and if you have a camera, you can also just use regular models like China already does.
- endymion-light 7mo agoI mean, i think if you want to perform mass surveilance, you can do it far cheaper and more efficiently via facial recognition, mobile phone surveillance and a variety of different other methods. If you want reconstruction and training of robotic movement, this is far more appropriate. I believe we're going to see robots being able to "dream" in terms of analysing historical video information on spaces and improving movement and navigation. So not mass surveilance, but probably there's a future of mass subjugation using robot enforcement.
- KaiserPro 7mo agoThis bit isn't that surveillance-y Relocalisation is the bit thats surveillance-y. But its also crucial for accurate visual only navigation.
- Tklaaaalo 7mo agoThe main use case is aligning virtual world with physical world for robotics and co.
- IshKebab 7mo agoVery cool. Doesn't seem like they've actually released the code: > This is a reimplementation of LoGeR; complete code and models will be released upon approval. I don't understand why it's a reimplementation either? I would guess it's "research" code anyway so not really usable unless you are an expert.
- ninjagoo 7mo agoFrom the paper website [1], the code is here [2] model weights are here [3] [1] https://loger-project.github.io/ https://loger-project.github.io/ [2] https://github.com/Junyi42/LoGeR https://github.com/Junyi42/LoGeR [3] https://huggingface.co/Junyi42/LoGeR https://huggingface.co/Junyi42/LoGeR
- Dead_Lemon 7mo agoWhat is the actual objective of this, is it solving an issue or creating a solution to a problem, that is still to be determined? It seems like a lot of energy to replicate a lidar mapping system. It's not like you can expect accurate dimensions from this approximate guess work, excluding the expected hallucinations adding to inaccuracy.
- flipbrad 7mo agoN00b question from me, perhaps, but how easy is it to mount and run Lidar on aerial drones?
- petargyurov 7mo agoIt's easy but it's not cheap. Well, price is relative but capturing video is certainly cheaper. Also, I am not sure how heavy LIDAR units are, but remember that the heavier the payload the more the flight time is reduced. Some drones can only have a single payload, so if you also want to capture (high-res) video/imgs you need to fly again. It all depends on the use-case.
- Daub 7mo agoThe most available lidar is found on your iPhone, but the results are orders of magnitude less detailed than that derived from photogrammetry. How ever an advantage is that lidar is not confused by reflections.
- taneq 7mo agoHuh? LIDAR absolutely is confused by reflections. Not always the reflections you can see (because often it’s using IR wavelengths) but nonetheless, reflections.
- voidUpdate 7mo agoVideo cameras are much cheaper and easier to use than LIDAR, like anyone can just pull out their phone, take a video and send it to this algorithm to get a reasonable point cloud of the environment. Sure, if you want an exact model of an environment and you have the time and money, LIDAR would give better results, but this is about doing more with less
- tmilard 7mo agoVery interesting paper. I can see street-view using it to perfect the 3D analysing of the photo-video they catch with there google-car. What a wonderfull time we are living in ! Specificaly in the Video to 3D reconstruction. Every month, a new brick is put in place.Super
- wumms 7mo agoStreet View cars added Velodyne LiDAR around 2017 [0][1], but it's optional. I found no data on 'LiDAR vs image only'-percentage. [0] https://arstechnica.com/gadgets/2017/09/googles-street-view-cars-are-now-giant-mobile-3d-scanners https://arstechnica.com/gadgets/2017/09/googles-street-view-... [1] https://en.wikipedia.org/wiki/Google_Street_View https://en.wikipedia.org/wiki/Google_Street_View
- KeplerBoy 7mo agoIt's safe to assume street view cars capture way more data than the stuff that ends up on the street view product.
- overfeed 7mo ago> I can see street-view using it to perfect the 3D analysing of the photo-video they catch with there google-car. Waymo recently announced[1] a World Model that does exactly this: using footage from a single-camera dashcam, it can predict/simulate multiple inputs that would have been sensed by a Waymo vehicle on the same travel path (i.e. multiple camera angles, Lidar cloud, etc). On top of this, the model can be prompted to customize the scenario (adding an elephants or a tornado were the example given) 1. https://waymo.com/blog/2026/02/the-waymo-world-model-a-new-frontier-for-autonomous-driving-simulation/ https://waymo.com/blog/2026/02/the-waymo-world-model-a-new-f...
- _fw 7mo agoThis is like something straight out of Cyberpunk 2077 - the braindances investigation scenes.
- realberkeaslan 7mo agoIt reminds me of that as well.
- Karliss 7mo agoMore like the opposite. Point cloud data captured with varying means has existed for a long time with raw data visualized more or less just like this. And SciFi movies/games use the effect of raw visualization as something futuristic/computer tech looking. Just like wireframe on black background, although that one is getting partially downgraded to more retro scifi status since drawing 3d wireframe isn't hard anymore. It started when any 3d computer graphics even basic wireframe was futuristic and not every movie could afford it, with some of them faking it with analog means. Any good scifi author takes inspiration from real world technology and extrapolate based on it, often before widespread recognition of technology by general population. Once something reaches the state of consumer product beyond just researchers and trained professionals, the visuals tend to get more polished and you loose some of the raw, purely functional, engineering style.
- priowise 7mo ago[dead]
- quadrature 7mo agoIn a traditional SLAM pipeline you do periodically fix drift by detecting when you've visited an area that you've mapped before this lets you align your sub maps so they are globally consistent. In the areas you have visited previously you have two estimates of your position one from your frame-to-frame estimates and another from the map you built of the area the first time. You can then solve an optimization problem to bring those two estimates closer together. In order to find out if you've already visited an area you store a description of the locations in a DB and search through them. The paper says they use a compressed representation of the "maps" and use test time training to optimize the global consistency between their sub maps.
- priowise 7mo agoMakes sense, thanks for the explanation. The compressed map representation + test-time training part sounds especially interesting. Does the approach hold up well when the environment changes over time (lighting, objects moved, etc.), or does it assume mostly static scenes?
- raphaelmolly8 7mo ago[dead]