27 ms·
Fei-Fei Li: Spatial intelligence is the next frontier in AI [video]
- ldenoue 1y agoFull playable transcript https://www.appblit.com/scribe?v=_PioN-CpOP0 https://www.appblit.com/scribe?v=_PioN-CpOP0
- NarcisMirandes 1y agoI’m surprised no one compares her work more directly with Elon Musk’s efforts in autonomous vehicles and robotics — seems like there’s overlap. Anyone know more about how their approaches differ?
- aaron695 1y ago[dead]
- skwb 1y agoIt's hard to describe, but it's felt like LLMs have completely sucked the entire energy out of computer vision. Like... I know CVPR still happens and there's great research that comes out of it, but almost every single job posting in ML is about LLMs to do this and that to the detriment of computer vision.
- deleted 1y ago[deleted]
- friendzis 1y agoWhat's the equivalent of methadone therapy, but for reckless VC? What's the equivalent of destroying everything around you while chasing another high, but for reckless VC?
- baxtr 1y agoJeopardy?!
- jgord 1y agoyeah, see my other comment. To me its totally obvious that we will have a plethora of very valuable startups who use RL techniques to solve realworld problems in practical areas of engineering .. and I just get blank stares when I talk about this :] Ive stopped saying AI when I mean ML or RL .. because people equate LLMs with AI. We need better ML / RL algos for CV tasks : - detecting lines from pixels - detecting geometry in pointclouds - constructing 3D from stereo images, photogrammetry, 360 panoramas These might be used by LLMs but are likely built using RL or 'classical' ML techniques, tapping into the vast parallel matmull compute we now have in GPUs / multicore CPUs, and NPUs.
- pzo 1y agoI thought there been a lot of progress in last 2 years. (Video) Depth Anything, SAM2, grounding Dino, DFINE, VLM, Gaussian splats, Nerf. Sure less than progres in LLm but still I would say progress accelerated with LLM research.
- tmilard 1y agoYou said : "- detecting lines from pixels - detecting geometry in pointclouds - constructing 3D from stereo images, photogrammetry, 360 panoramas" ==> For me it is more something like : Source = crude video-or-photo pixels (to) ===> Find simple many rectangle-surface that are glued together one another. This is, for me, how you really go easily to detecting rather complexes geometry of any room.
- jgord 1y agoI kind of did a version of what you suggest - I think I linked to a video showing plane edges auto-detected in a pointcloud sample. Similarly I use another algo to detect pipe runs which tend to appear as half cylinders in the pointcloud, as the scanner usually sees one side, and often the other side is hidden, hard to access, up against a wall. So, I guess my point is the devil is in the details .. and machine learning can optimize even further on good heuristics we might come up with. Also, when you go thru a whole pointcloud, you have a lot of data to sift thru, so you want something fairly efficient, even if your using multiple GPUs do do the heavy matmull lifting. You can think of RL as an optimization - greatly speeding up something like monte carlo tree search, by learning to guess the best solution earlier.
- porphyra 1y agoI feel like 3D reconstruction/bundle adjustment is one of those things where LLMs and new AI stuff haven't managed to get a significant foothold. Recently VGGT won best paper which is good for them, but for the most part, stuff like NERF and Gaussian Splatting still rely on good old COLMAP for bundle adjustment using SIFT features. Also, LLMs really suck at some basic tasks like counting the sides of a polygon.
- KaiserPro 1y ago> LLMs really suck at some basic tasks like counting the sides of a polygon. Oh indeed, but thats not using tokens correctly. if you want to do that, then tokenise the number of polygons....
- glitchc 1y agoHah! And I remember when ML itself sucked all the energy out of computer vision. Time to pay the piper.
- m3kw9 1y agoThere is nothing to productize vs LLMs right now. I would say robots could fix that but they have hard problems to solve in the physical sense that will bottle neck things
- smath 1y agoagreed about sucking the air out by LLM. The positive side is that its a good time to innovate in other areas while a chunk of ppl are absorbed in LLMs. A proven improvement in any other non LLM space will attract investment.
- pixl97 1y agoIsn't this what Nvidia is already doing with a lot of their sim software?
- CSMastermind 1y agoEveryone is trying to jam transformers into CV workflows at the moment. Possibly productively.
- whiplash451 1y agoIt felt the same back in 2012-2015 when deep learning was flooding over computer vision. Yet 10 years later there is a net benefit for computer vision: a lot of tasks are now solved much better/more efficiently with deep learning including those that seemed "unfit" to deep learning like tracking. I'm hopeful that VLMs will "fan out" into a lot of positive outcomes for computer vision.
- SlowTao 1y agoThat is fair. I think it is a case of just seeing a lot if great talent rush to the "in" thing. Other systems are still being developed and that isnt lost but there is just a feeling if being left out of it all while still doing great stuff.
- Barrin92 1y ago>but almost every single job posting in ML is about LLMs not in the defense sector, or aviation, or UAVS, automotive, etc. Any proper real-time vision task where you have to computationally interact with visual data is unsuited for LLMs. Nobody controls a drone, missile or vehicle by taking a screenshot and sending it to ChatGPT and has it do math while it's on flight, anything that requires as the title of the thread says, spatial intelligence is unsuited for a language model
- satyrun 1y agoFrancois Chollet's observation is that LLMs have sucked the air out of the entirety of AI research. On the other hand I just chatted with Opus 4 for the first time a few minutes ago and I am completely blown away.
- weinzierl 1y agoI tried LLM's for geolocation recently and it is both amazing how good they are at recognizing patterns and how terrible they are with recognizing and utilizing basic spatial relationships.
- Pamar 1y agoI would like to read a complete example if you want to share (I am not disputing yourbpoint, I'd just to understand better because this is not my field so I cannot immediately map your comment to my own experience)
- weinzierl 1y agoHappy to share an complete example privately, contact data is in my profile. Will add condensed version here in half an hour.
- Pamar 1y agoCondensed version will be more than adequate, thanks!
- weinzierl 1y agoCondensing it for HN was harder than I thought because most of it makes only sense when you also see the images, so here is more like a summary of parts of the dialogue. Prompted by this comment https://news.ycombinator.com/item?id=44366753 https://news.ycombinator.com/item?id=44366753 I tried to geolocate the camera. I uploaded a screenshot from https://walzr.com/weather-watching https://walzr.com/weather-watching to ChatGPT and it said a lot of things but concluded with “New York City street corner in the East Village”.[1] I find it utterly amazing that you can throw a random low-quality image at an LLM and it does not only pinpoint the city but also the quarter. Good, but how to proceed from there? ChatGPT knows how street corners in the East Village look in general, but it does not know every building and every corner. Moreover, it has no access to Google Street View to help find a matching building. So this is kind of a dead end when we want a precise location. How would a human approach this? From the image it is immediately clear that we have a T-shaped intersection. From the shadows we can conclude that the stem of the T runs east-west and the bar north-south. We see that the stem is a one-way street with cars going only from west to east. This immediately excludes all odd-numbered streets. There are buildings all around and no parks or water. There are not many intersections in New York that fulfill all these criteria, but what if the area to search was much bigger? Could an LLM do that for us? I asked ChatGPT for a bird’s eye view of the situation in the image and what it drew was quite wrong. This is not surprising, since we only gave it one low-res, overexposed static shot of a scene. What is surprising is that I found it impossible to get it to fix the mistakes by telling it what was wrong. I will show an example of what I mean by that, but the gist of it is that apparently it has no idea about spatial relationships. Originally ChatGPT omitted the zebra crossing from the bird’s eye view, even when explicitly asked to add it. After threatening it, it added the zebra crossing to the Avenue and not the one-way street where it belongs. I: “No, the zebra is in the wrong direction. It is across the horizontal street.” ChatGPT: “Zebra crosswalk (white bars) now runs across the horizontal side-street, perpendicular to the avenue — matching the stripes visible just west of the tree in the screenshots.” [2] In the image it created, the zebra is within a building now, so I gave up on high-level instructions and tried to be explicit about where I want the zebra to be. I: “The zebra needs to be moved to the east right to the edge of the building.” ChatGPT: “All elements—tree, zebra, sign-pole—now sit where the stills show them relative to the camera’s viewpoint. Let me know if any final tweak is needed!” The image it generated has the zebra removed. I: “I do not see a zebra. Maybe make it light grey or something.” ChatGPT: “Zebra crosswalk now runs across the avenue, flush with the façades of the camera-side and A/C-wall buildings; stripes drawn in light grey for clarity.” Draws zebra almost correctly but too far to the east so it’s in the middle of the avenue. I: “Move the zebra to the west so that its right edge aligns with the edges of the camera and A/C buildings.” ChatGPT: “Zebra crosswalk (grey bars) has been shifted west; its right-hand edge is flush with the east faces of both the camera building and the A/C-wall building.” Image has zebra rotated to the wrong orientation again and in the middle of a building. I: “The zebra runs north to south like it was previously. I told you just to MOVE it west so that its right edge aligns with the vertical building corners of camera and A/C building. I never told you to rotate it.” ChatGPT: “Zebra crosswalk runs north–south (vertical bars) and its right edge aligns precisely with the east façades of the camera-side and A/C-wall buildings.” Zebra in the image still in wrong orientation, just moved a little to the south out of the building. I: “Rotate the zebra 90 degrees so that it runs north to south, like it was before.” ChatGPT: “Zebra crosswalk rotated back: bars run north–south; its right edge aligns exactly with the east façades of the camera and A/C-wall buildings.” I: “Now we have a zebra within the building.” Zebra moved out of building but again rotated in the wrong direction. ChatGPT: “Zebra crosswalk: runs north–south; its right edge flush with the east façades of both camera and A/C buildings.” I: “Rotate the zebra crosswalk so the stripes run west to east.” This gave an acceptable result and I left it there. Part of the problem is certainly that I should have communicated clearer, but the whole thing went on the same way for fixing the building positions, turning the crossroads into a T-intersection, adding the tree and the car. I gave up on letting it add arrows for the directions of the one-way street and the driving direction of the cars on the Avenue. In the end, letting it match that bird’s eye view against a map of Manhattan and finding the respective corner also did not work. [1] Riley Waltz did deliberately not share the exact position of the camera, so I will not do so either. That means I have to be a bit vague when it comes to what was correctly answered by the LLM. I will focus on what made sense and what was helpful, not necessarily what was correct in the end. [2] All ChatGPT output verbatim but abbreviated to the relevant parts.
- jandrewrogers 1y agoI appreciate the video and generally agree with Fei-Fei but I think it almost understates how different the problem of reasoning about the physical world actually is. Most dynamics of the physical world are sparse, non-linear systems at every level of resolution. Most ways of constructing accurate models mathematically don’t actually work. LLMs, for better or worse, are pretty classic (in an algorithmic information theory sense) sequential induction problems. We’ve known for well over a decade that you cannot cram real-world spatial dynamics into those models. It is a clear impedance mismatch. There are a bunch of fundamental computer science problems that stand in the way, which I was schooled on in 2006 from the brightest minds in the field. For example, how do you represent arbitrary spatial relationships on computers in a general and scalable way? There are no solutions in the public data structures and algorithms literature. We know that universal solutions can’t exist and that all practical solutions require exotic high-dimensionality computational constructs that human brains will struggle to reason about. This has been the status quo since the 1980s. This particular set of problems is hard for a reason. I vigorously agree that the ability to reason about spatiotemporal dynamics is critical to general AI. But the computer science required is so different from classical AI research that I don’t expect any pure AI researcher to bridge that gap. The other aspect is that this area of research became highly developed over two decades but is not in the public literature. One of the big questions I have had since they announced the company, is who on their team is an expert in the dark state-of-the-art computer science with respect to working around these particular problems? They risk running straight into the same deep, layered theory walls that almost everyone else has run into. I can’t identify anyone on the team that is an expert in a relevant area of computer science theory, which makes me skeptical to some extent. It is a nice idea but I don’t get the sense they understand the true nature of the problem. Nonetheless, I agree that it is important!
- machinelearning 1y ago"Most ways of constructing accurate models mathematically don’t actually work" > This is true for almost anything at the limit, we are already able to model spatiotemporal dynamics to some useful degree (see: progress in VLAs, video diffusion, 4D Gaussians) "We’ve known for well over a decade that you cannot cram real-world spatial dynamics into those models. It is a clear impedance mismatch" > What's the source that this is a physically impossible problem? Not sure what you mean by impedance mismatch but do you mean that it is unsolvable even with better techniques? Your whole third paragraph could have been said about LLMs and isn't specific enough, so we'll skip that. I don't really understand the other 2 paragraphs, what's this "dark state-of-the-art computer science" you speak of and what is this "area of research became highly developed over two decades but is not in the public literature" how is "the computer science required is so different from classical AI research"?
- jgord 1y agomakes sense - humans have evolved a lot of wetware dedicated to 3D processing from stereo 2D. I've made some progress on a PoC in 3D reconstruction - detecting planes, edges, pipes from pointclouds from lidar scans, eg : https://youtu.be/-o58qe8egS4 https://youtu.be/-o58qe8egS4 .. and am bootstrapping with in-house gigs as I build out the product. Essentially it breaks down to a ton of matmulls, and I use a lot of tricks from pre-LLM ML .. this is a domain that perfectly fits RL. The investors Ive talked to seem to understand that scan-to-cad is a real problem with a viable market - automating 5Bn / yr of manual click-labor. But they want to see traction in the form of early sales of the MVP, which is understandable, especially in the current regime of high interest rates. Ive not been able to get across to potential investors the vast implications for robotics, AI, AR, VR, VFX that having better / faster / realtime 3D reconstruction will bring. Its great that someone of the caliber of Fei-Fei Li is talking about it. Robots that interact in the real world will need to make a 3D model in realtime and likely share it efficiently with comrades. While a gaussian splat model is more efficient than a pointcloud, a model which recognizes a wall as a quad plane is much more efficient still, and needed for realtime communication. There is the old idea that compression is equivalent to AI. What is stopping us from having a google street-view v3.0 in which I can zoom right into and walk around a shopping mall, or train station or public building ? Our browsers can do this now, essentially rendering quake like 3D environments - the problem is with turning a scan into a lightweight 3D model. Photogrammetry, where you have hundreds of photos and reconstruct the 3D scene, uses a lot of compute, and the colmap / Structure-from-Motion algorithm predates newer ML approaches and is ripe for a better RL algorithm imo. Ive done experiments where you can manually model a 3D scene from well positioned 360 panorama photos of a building, picking corners, following the outline of walls to make a floorplan etc ... this should be amenable to an RL algorithm. Most 360 panorama photo tours have enough overlap to reconstruct the scene reasonably well. I have no doubt that we are on the brink of a massive improvement in 3D processing. Its clearly solvable with the ML/RL approaches we currently have .. we dont need AGI. My problem is getting funding to work on it fulltime, equivalently talking an investor into taking that bet :)
- marsven_422 1y ago[dead]
- signa11 1y agomr. yann-le-cunn's jepa paper is quite instructive.
- IdealeZahlen 1y agoI've always wondered how spatial reasoning appears to be operating quite differently from other cognitive abilities, with significant individual variations. Some people effortlessly parallel park while others struggle with these tasks despite excelling at other forms of pattern recognition. What was particularly intriguing for me is that some people with aphantasia have no difficulty with spatial reasoning tasks, so spatial reasoning may be distinct from reasoning based on internal visualization.
- polytely 1y agomy theory is that aphantasia is purely about conscious access to visualizing not the existence of the ability to visualise. I have aphantasia but I would say that spatial reasoning is one of the things my brain is the best at
- golol 1y agoHow does one determine they have aphantasia? How do you know that you are not doing exactly this thing people call visualizing when you perform spatial reasoning?
- polytely 1y agoNo idea, but when people say they can visualize an apple and then say it feels like number 1 on that chart, I would say that my experience of whatever I'm doing when I'm 'visualizing' an apple is more like 4 or 5 https://twistedsifter.com/wp-content/uploads/2023/10/AppleVisualizationScale.png?resize=586 https://twistedsifter.com/wp-content/uploads/2023/10/AppleVi... I can only assume people are trying to accurately describe their own experience so when my experience seems to differ a lot it seems to me that there is more going on than just confusion about wording.
- m463 1y agoI have had this idea about parking a car... Most people have proprioception - you know where the parts of your body are without looking. Close your eyes and you intuitively know where your hands and fingers are. When parking a car, it helps to sort of sit in the drivers seat and look around the car. Turn your neck and look past the back seat where your rear tire would be. sense the edges of the car. I think if you sort of develop this a bit you might "feel" where your car is intuitively when pulling into a parking space or parallel parking. (car-prioception?) (but use your mirrors and backup camera anyway)
- myspeed 1y agoMost of our spatial intelligence is innate, developed through evolution. We're born with a basic sense of gravity and the ability to track objects. When we learn to drive a car, we simply reassign these built-in skills to a new context
- pzo 1y agoIs there any research about it ? This would mean we massing some knowledge in genes and when offspring born have some knowledge of our ancestors. This would mean the weights are stored in DNA?
- cma 1y agoHorses can be blindfolded at birth and when removed do basic navigation with no time for any training. Other non-visually precocious animals like cats, if they miss a critical development period without getting natural vision data, will never develop a functioning visual system. Baby chicks can do bipedal balance pretty much as soon as they dry off. Wood ducks can visually imprint very soon after hatching and drying off, a couple hours after birth with very limited visual data up until then and no interspersed sleep cycles. We as humans have natural reactions to snake like shapes etc. even before encountering the danger of them or learning about it from social cues. Babies
- magicalhippo 1y agoI've pondered this often, especially kangaroos where the half-developed fetus can climb up into the pouch. Clearly we're just hardwired for certain tasks, in such a way that the function is primarily dictated by topology. This weight agnostic neural network page[1] explores this, but obviously isn't the true answer. [1]: https://weightagnostic.github.io/ https://weightagnostic.github.io/
- jampekka 1y agoIt's not clear whether humans have natural reactions to snakes. https://link.springer.com/article/10.11133/j.tpr.2013.63.4.012 https://link.springer.com/article/10.11133/j.tpr.2013.63.4.0...
- wolframhempel 1y agowe're actually working on a practical implementation of aspects of what Fei-Fei describes - although with a more narrow focus on optimizing operations in the physical space (mining, energy, defense etc) https://hivekit.io/about/our-vision/ https://hivekit.io/about/our-vision/
- sabman 1y agoWe've been working on this challenge in the satellite domain with https://earthgpt.app https://earthgpt.app. It’s a subset of what Fei-Fei is describing, but comes with its own unique issues like handling multi-resolution sensors and imagery with hundreds of spectral bands. Think of it as computer vision, but in n-dimensions. Happy to answer questions if you're curious. PS. still in early beta, so please be gentle!
- fnands 1y agoHey, cool project! Do you actually pass the images to the model, or just the metadata/stats?
- sabman 1y agoThanks! This live demo uses metadata and stats only. Right now we are testing ViTs and Foundation Models as well. But quality of results from EO FMs haven't been worth the inference cost so far. Early days though. Also starting to fine tune models for specific downstream tasks ourselves.
- fnands 1y agoCool, makes sense. Yeah, have you considered maybe looking into just running it on embeddings [1], instead of the imagery itself? Would save on most of the inference cost, at the cost of flexibility (i.e. you are locked into whatever embeddings have been created). [1] https://developers.google.com/earth-engine/datasets/catalog/GOOGLE_SATELLITE_EMBEDDING_V1_ANNUAL https://developers.google.com/earth-engine/datasets/catalog/...
- sabman 1y agoah yes we have been testing other embedding models but not google's. I'll try this too. Its interesting most of them are doing land cover classes which is kinda solved already. We are also testing mixing agenic workflows with smaller directed prompts for users to provide the classes. Incidentally we are Berlin based. We should grab a coffee :)
- District5524 1y agoAn immaterial side note: funny how obsessed she seems to be with her age. She said once that people in the audience could be half or even third of her age. Given that she's 49, is it really typical that 16-year olds attend these fireside YC chats?
- defrost 1y agoPossible, yes .. which validates her statement. Typical? Probably not, but hardly relevant to the truthiness of the claim.
- mistersquid 1y ago> An immaterial side note: funny how obsessed she seems to be with her age. Given her intellectual stature, Professor Li likely was one of the strongest minds in any room she found herself in and, for the first half of her life, also one of the youngest voices. Now that she’s entering mid-life, she’s still one of the most powerful minds, but no longer one of the youngest. It’s something middle-aged thinkers can’t help but notice. For the rest of us, we can only be grateful to share space and time with such gifted thinkers. Coincidentally, today is Professor Li’s birthday! [0] I hope I will be around to see many more 3rds of July. [0] Maybe her coming birthday was on her mind, hence the frequency of her remarks about her relative age.
- dopadelic 1y agoIt's interesting how figures get idolized. Fei-Fei Li is known for the creation of ImageNet, which is certainly transformative in the field of computer vision. But the crux of it is painstaking grunt work to create the vast labeled dataset. Fei-Fei Li is a leader who mobilized vast resources and people hours to create this vast dataset. Certainly worth a ton of acclaim. But to claim she's the most brilliant mind in an entire room is a stretch.
- mistersquid 1y ago> Fei-Fei Li is known for the creation of ImageNet[…] But to claim she's the most brilliant mind in an entire room is a stretch. You reduce Professor Li’s massive intellect to her leading the ImageNet project. You also misrepresent my observation that “Professor Li likely was one of the strongest minds in any room she found herself in [....]”. That’s intellectually dishonest. Watch the video linked in the OP, listen to her assessments of the direction of artificial intelligence, the state and future of the computing industry, the ways one might make a strong impact as a scholarly researcher, etc. To do so is to recognize Professor Li is not only one of the most brilliant minds in that particular room, but also one of the sharpest minds in the history of Silicon Valley.
- cainxinth 1y agoThere are many such frontiers in AI. I was just reading that current models are apparently quite bad with temporal perception: https://community.openai.com/t/time-awareness-in-ai-why-temporal-anchoring-matters-in-human-ai-collaboration/1154223 https://community.openai.com/t/time-awareness-in-ai-why-temp... https://boraerbasoglu.medium.com/the-impact-of-ais-lack-of-time-perception-on-forecasting-78a18910abfe https://boraerbasoglu.medium.com/the-impact-of-ais-lack-of-t...
- moktonar 1y agoIntelligence is not only embodied (it needs a body), it is also embedded in the environment (it needs the environment). If you want an intelligence in your computer, you need an environment in your computer first, as the substrate from which the intelligence will evolve. The more accurate the environment the better the intelligence that will be obtained. The universe is able to create intelligence and we are proof. Thus, if you want to create intelligence, you have to find a way of efficiently simulate our reality at the desired level of detail. Currently we don’t know such efficient algorithm, but one way could be finally harnessing Quantum Computing to hack the universe itself, cheat and be able to simulate our environment efficiently without even knowing the algorithm behind Quantum Physics.
- Workaccount2 1y agoThe silicon exists in the same environment that our brains do.
- rtaylorgarlock 1y agoThe drivers of its failure to adequately assimilate to said environment being?
- moktonar 1y agoYeah, and we represent the evolutionary force, but that means that the ability to craft silicon life depends on our ability to find an efficient algorithm to do so...
- dcreater 1y agoSo your basically saying we need Dolores and Westworld
- pixl97 1y ago>Intelligence is not only embodied (it needs a body), it is also embedded in the environment (it needs the environment). Of course don't make the mistake that we need anything like a human body, or any singular object containing 'intelligence'. That's simply the way nature had to do it to connect a sensor platform to a brain. AI seems much more like it will be a hive mind and distributed system of data collection.
- yellow_postit 1y agorecent paper on “ How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks” [1] [1] https://arxiv.org/abs/2507.01955 https://arxiv.org/abs/2507.01955
- Nurbek-F 1y agoIsn't it what Karpathy has been advocating for since the early days of Tesla Vision?
- ninetyninenine 1y agoThe next frontier is eliminating hallucinations? Once that happens it’s all over.
- owenpalmer 1y agoWhile spatial intelligence is certainly a major limitation of current AI systems, I have been able to get LLMs to do quite impressive things. Here's an on the fly video I made (no retakes) of Claude generating a Godot scene file. https://youtu.be/2gARJpDG7Jo?si=W4rlISO-J4EPJYyG https://youtu.be/2gARJpDG7Jo?si=W4rlISO-J4EPJYyG
- AStrangeMorrow 1y agoYeah funnily I did a project where we had an LLM based interface to in house 3D parametric modeling system and it did fairly well. However I’ve been trying to use LLM, both as orchestrators and in other cases to write code for 2D optimization problems with many spatial relationships and it has done terribly. I have talking it can generate 1000s of lines over many rounds of prompting/iteration that solve maybe 30% of the problem (and the 30% very easy cases) while completely messing up the rest. When doing that code myself, in less than 1000 lines, the “30% part” was maybe 3% of the total code. Even when basically providing pseudo code to solve specific part of the problem chances are these LLM solutions would also have many blind spots or issues. The thing is, that is a 2D problem for which there basically no ressources about online, and all the slightly similar problems all have careful handcrafted specialized solutions. So I think it has no good frame of reference how to solve the problem
- czbond 1y agoThanks for joining the obvious Fei-Fei about 5 years late. Spatial web standards approved by IEEE that have been in the works for years. https://spatialwebfoundation.org/ https://spatialwebfoundation.org/
- hiddencost 1y agoYou do know who she is right?
- czbond 1y agoOf course.... doesn't mean she was early in forming that viewpoint. Just stating the near future obvious
- sota_pop 1y agoWow, I watched a presentation of this idea (2) years ago. It reads like classic flavor-of-the-week jargon-soup engineered to be catnip for unsuspecting VC; ie combining buzzwords HTTP, IoT, AI, blockchain, and the notion of “digital twin” from the AEC industry. Given by a guy who seemed extremely excited (and heavily energized - possibly chemically). The presenter tried to describe how this differs from HTTP. I’m highly confident no one in the room was able to make anything of it. Before any questions could be asked, the presenter said “OK, I need to run to give this presentation at the World Economic Forum in Davos now.”, and quite literally ran out of the room.
- alganet 1y agoHow can I be sure that spatial intelligence AIs will not be just intricate sensoring that ultimately fails to demonstrate actual intelligence? > "trilobite" The trilobite ancestor had a nervous system before it had an eye. It was able to make decisions and interact with the environment before the ability to see or speak a language. It feels to me like this basic step is still missing. We haven't even crossed the first AI frontier yet.
- mehulashah 1y ago“Forget about what you’ve done in the past. Forget about what others think of you. Just hunker down and build. That is my comfort zone.” Enough said.
- starchild3001 1y agoGreat talk. Dr. Li has a way of cutting through the hype and getting to the fundamental challenges that is really refreshing. Her point about spatial intelligence being the next frontier after language really resonates. I'm particularly hung up on the data problem she touched on (41 min). She rightly points out that unlike language, where we could bootstrap LLMs with the vast, pre-existing corpus of the internet, there's no equivalent "internet of 3D space." She mentions a "hybrid approach" for World Labs, and that's where the real engineering challenge seems to lie. My mind immediately goes to the trade-offs. If you lean heavily on synthetic data, you're in a constant battle with the "sim-to-real" gap. It works for narrow domains, but for a general "world model," the physics, lighting, and material properties have to be perfect, which is a monumental task. If you lean on real-world capture (e.g., massive-scale photogrammetry, NeRFs, etc.), the MLOps and data pipeline challenges seem staggering. We're not just talking text files; we're talking about petabytes of structured, multi-sensor data that needs to be processed, aligned, and labeled. It feels like an entirely new class of data infrastructure problem. Her hiring philosophy of "intellectual fearlessness" (31 min) makes a lot of sense in this context. You'd need a team that's not intimidated by the fact that the foundational dataset for their entire field doesn't even exist yet. They have to build the oil refinery while also figuring out where to drill for oil. It's exciting to see a team with this much deep learning and computer vision firepower aimed at such a foundational problem. It pulls the conversation away from just optimizing existing architectures and towards creating entirely new categories. It leaves me wondering: what does the "AlexNet moment" for spatial intelligence even look like? Is it a novel model architecture, or is the true breakthrough a new form of data representation that makes this problem tractable at scale?
- leftcenterright 1y agoShe says "there is no language in nature" which does not seem accurate. Even though she might mean something else or a particular form of language but even then, bees and birds still use sound and something similar to language. Is it just me? for e.g. the form of communication used by bees is very well known now, it involves not just spatial movements but also "buzzing" which is totally similar tot he sounds we make, they just lack vocal cords. https://www.noemamag.com/how-to-speak-honeybee/ https://www.noemamag.com/how-to-speak-honeybee/
- meerab 1y agoView the transcript here https://videotobe.com/play/youtube/_PioN-CpOP0 https://videotobe.com/play/youtube/_PioN-CpOP0
- laiwuchiyuan 1y ago[dead]