13 ms·
FastVLM: Efficient vision encoding for vision language models
- porphyra 1y agoIt seems that the future of robotics is VLA models. Even Tesla FSD is an end-to-end VLA model. Efficient vision encoding will be a huge part of making robots safe and responsive.
- deleted 1y ago[deleted]
- BryanLegend 1y agoSeems like the main thing holding these new minds back is being able to see well. Breakthroughs like this will fix that.
- efnx 1y agoThat and the ability to hold on to knowledge.
- static_void 1y ago... or say they don't know.
- kamranjon 1y agoApple out here playing 5d chess, installing neural cores in their hardware and writing crazy efficient vision models to run on em. Cool stuff.
- vFunct 1y agoCan it fill a wine glass to the rim?
- mkl 1y agoIt's for interpreting images, not generating them.
- turnsout 1y agoApple has gotten a slow start in the LLM world, but they have the only long term strategy that makes sense. They’re going to dominate the 2030s.
- jfarina 1y agoWhat strategy is that?
- ryanmcgarvey 1y agoI presume they mean that distribution is king and they make all the devices.
- boroboro4 1y agoWhat exactly the strategy is?
- generalizations 1y agoThey can run locally on-device: a win for cost, latency and privacy (privacy is pragmatic: it means you can use all the user's data as context without qualms). There's a reason Microsoft tried so hard to push for the neural processors a year or two ago. Avoiding the cost of the datacenter while offering good-enough inference (emphasis on good) is a massive win.
- turnsout 1y agoYes, thank you; this is the strategy I was referring to. It will take some time for the models and chips to get there, but on-device inference will have massive advantages for privacy, speed and cost. Plus it will drive demand for hardware—at first, iPhones, but soon AirPods and glasses.
- xnx 1y agoGoogle already has some of the best on device models (Gemma) and chips (Tensor).
- insane_dreamer 1y agoAs the father of a young child whose optic nerves are highly deteriorated (compression) and is expected to lose his sight (when exactly is unknown; based on original projections he should be blind by now, but an experimental treatment run in a trial at the NIH (KEEP FUNDING SCIENCE) has stabilized his sight), I'm overjoyed with the advances being made in VLMs. I can now envision a future where even if he loses his sight he'll be able to interact with the world around him, go to college, have a fulfilling career (he loves science and engineering, and is talented for his young age), etc.
- lynx97 1y agoI grew up in the 80s as a 100% blind child. Technology was by far not as advanced as today. Computers were just coming up when I was around 12. I learnt to type on a oldschool typewriter, and I also learnt to write braille with a pretty heavy full-metal embossing device. OCR was still quite bad. When I switched to what you call high scooll, I used a laptop with integrated Braille display to follow classes. Used good old DOS as OS and Word 5.5 as my "notepad". Except for PC Lingua for Latin, I basically had no tools specialized for learning. A electronic notepad and my brain was all I had to follow school. And I still made it. I have a great job I love, my own appartment, a sweet girlfriend and I am basically completely independent. To a point where I had to forcefully send away my mother since her continued attempts to "help" me were basically detrimental to my own development. I can not emphasis how important it is how you deal with it as a parent. Since parents are indeed the biggest hinderence to development, we have a saying around here amongst disabled people: "additional disability due to parental overprotection" (Zusatzbehinderung Eltern). Please take a moment to understand what this means, without feeling personally attacked. Its important. Your child can leave home around 18, just like every other kid. I did. Don't slow that process down artificially. The more this is prolonged, the harder it gets for the individual to actually obtain independence. I am telling you this because I read between the lines that you believe current technology is a reason for you to be hopeful. Sure, it should be. But never forget, your child can do much more then you as a sighted person will ever be able to understand. Don't let them drown in your own misery. Let them discover what they can do. You will be surprised what they come up with. And dont fall for Gear Acquision Syndrome. Sure, tools are nice, and they do get better, which is also nice. I LOVE vision models, to stay on topic somehow. However, I still leave my house with only a cane and my phone in my pocket. I do occasionally ask Siri "Where am I" to get an address if I happen to have forgotten where I am exactly, currently. But at the end of the day, my cane is what shows me the way. Most tech is hype, plain old hearing and your sense of touch gets you much farther then you might think. Wish you all the best for your own journey, and the development of your child.
- liamwire 1y agoIt feels like this is the required level of speed-up needed re. time-to-first-token to make continuous vision useful for on-device applications like an assistant that can see and take action on your screen, ala the original Apple Intelligence demos. It’s very impressive seeing the app in the repo and I’m excited to build it tonight and play around.
- nine_k 1y agoWith that, a really helpful aid for blind people can be made, running just on their phone, fed from a camera in their eyeglasses. Somebody who could not move around without an assistant could become autonomous in daily life.
- jdiff 1y agoIt might be useful for telling Cream of Chicken from Cream of Mushroom, but for locomotion I can't see this adding anything over existing strategies people use to get around sans sight. "There's a tree. There's a tree. There's a tree. There's a number of pedestrians. There's a tree. There's a sign." does not strike me as useful feedback for getting around.
- nine_k 1y agoConsider a city. It's full of signs and inscriptions, traffic lights, and other key interaction elements. Consider a store. It has shelves with stuff, again with inscriptions, price tags, etc. "Pavement. Row of stores to the left. Joe's Grocery Store. Doors. Door handle. A shelf with bakery. A shelf with canned goods. A shelf with bottles. Coke bottle. Large Pepsi bottle. Apple juice bottle. Passageway. Checkout. Payment terminal. Door. Door handle. Pavement. ..."
- jdiff 1y agoNone of that gives me any useful spatial sense of where. "Payment terminal." Okay. Where is it? Left? Left where? How much left? How far? The only truly useful bits I see in your stream of text is, again, "Cream of Mushroom" vs "Cream of Chicken." I am actively holding something, so I know where it is, but need to differentiate it from printed detail.
- adamsiem 1y agoAnyone using vision to parse screenshots? QVQ was too slow. Will give this a shot.
- abrichr 1y agoYou might be interested in https://github.com/OpenAdaptAI/OpenAdapt https://github.com/OpenAdaptAI/OpenAdapt
- logankeenan 1y agoI used molmo to parse screenshots in order to detect locations of UI elements. See the repo below. I think Omni parser from Microsoft would also work well. https://github.com/logankeenan/george https://github.com/logankeenan/george https://github.com/microsoft/OmniParser https://github.com/microsoft/OmniParser
- nprateem 1y ago[flagged]
- Aeroi 1y agoI built/building a realtime voice+vision app called Sen, its currently live in beta and streams frames over webrtc. It's fast and smart, but Im super curious to see how these models do as we get closer to the metal. I can see these running on-device in the future with super fast ttfb.
- keyle 1y agoDo you have a write up of the tech stack and setup? Or willing to give the gist here? I'd like to make a private Qwen or similar for my kids to prompt with a button and voice control. It doesn't need vision... Although eventually that'd be very cool. Siri just sucks. We might not be there yet...
- Aeroi 1y agoyeah i made a post on here, but the algo sent it to the gulag abyss. https://news.ycombinator.com/item?id=43926673 https://news.ycombinator.com/item?id=43926673
- keyle 1y agoThat's a good product site but it doesn't help me in anyway...
- Aeroi 1y agoI also ran across an interesting robot toy demo today that had voice built in. it was whimsical and seemed like it was aimed towards primary education and kids. Someone here might know the name.
- stavros 1y agoYou can use Ollama or LM Studio, both in API mode, to return the responses. I believe they offer audio support, but I'm not entirely sure. However, if you're looking for instruction following (like an agent), I've tried to implement my own agent and have lost faith. Even GPT-4.1 will regularly gaslight me that no, it definitely ran the tool call to add the event to my calendar, when it just didn't. I can't get any more adherence out of it.
- deleted 1y ago[deleted]
- nikolayasdf123 1y ago2GB for 0.5B smallest model. it does not make sense for each app to download this. apple must have plans to pre-load these models on os level and expose SDK for all apps to call these models locally. exciting times! opened issue for them to confirm this: https://github.com/apple/ml-fastvlm/issues/7 https://github.com/apple/ml-fastvlm/issues/7
- cube2222 1y agoThat’s what they suggested about LLMs at last year’s WWDC iirc. The core models are provided by the OS, while apps bring LORAs to fine-tune them / bring custom heads for them.
- babl-yc 1y agoYou could probably get away with f16 or even quantize to int8 and have a much smaller model, but your point stands. Users won't be thrilled to download a 500MB model for an app either.
- nikolayasdf123 1y agohaha latest Uber build for iOS 18 is 500MB... without LLM models <face-palm/>
- ukuina 1y agoWhat are they doing in there? Is it mostly visual assets?
- bastawhiz 1y agoIf I was going to guess, I'd get there's a ton of third party code for things like payment method SDKs. Every local payment method around the world is going to have its own package that you need to import, and you can't just load in new executable code on the fly after the app is installed.
- 1y ago
- nikolayasdf123 1y agogoogle and cloud LLM providers must be biting their teeth now! haha
- nikolayasdf123 1y agodistributing this heavy compute and moving it close to device where 1. source of data happens; 2. decision and output about the result of analysis is done; is way to go. super low latency, no network traffic, privacy, less overhead in cloud. this is amazing
- lynx97 1y agoI wonder, can I convert/run this with llama.cpp? It being LLaVA based seems promising.
- vessenes 1y agoUm wow. The on-device realtime videos are worth a watch, and compelling. Looking forward to this being deployed and widely adopted. Getting much faster time to first token opens up a ton of features and usability benefits.
- buyucu 1y agowhere is my gguf?
- kristel100 1y ago[dead]
- simianparrot 1y agoI have a feeling feeding tesseract the image every 1 second would be significantly faster and take far less space and processing power? Haven't tested it yet but given how fast tesseract is on large images, it wouldn't surprise me.
- regularfry 1y agoIf all you want is OCR, possibly.
- coredog64 1y agoIf all you want is OCR of typewritten text. Tesseract is awful for handwriting.
- d3k 1y agoVery nice! I wish they were more keen to contribute to AI/ML community an publish weights and model definition on HuggingFace. Funny enough I have just seen today a similar demo that is using a freely available VLM: https://github.com/ngxson/smolvlm-realtime-webcam https://github.com/ngxson/smolvlm-realtime-webcam
- tough 1y agoSmolVLM is from huggingface team cool to see people doing stuff with smaller models https://huggingface.co/blog/smolvlm https://huggingface.co/blog/smolvlm https://arxiv.org/abs/2504.05299 https://arxiv.org/abs/2504.05299
- labadal 1y agoI'm absolutely thrilled that there is an effort to make models smaller and run with less resources instead of blindly throwing more resources at the problem and expecting it to get solved.