7 ms·
Video Surveillance with YOLO+llava
- _giorgio_ 2y agoAll I see, usually, is some AI YOLO algorithm applied to an offline video. This is the first time that I've seen a "complete" setup. Any info to learn more on applying YOLO and similar models to real time streams (whatever the format)?
- hug 2y agoThis repository seems to be exactly what you are asking for. It's YOLO analysis of video frames passed in through Real Time Streaming Protocol.
- _giorgio_ 2y agoYes, probably it's only one reasonably sized, let's say that with a lot of patience you can study it! I'll search for some online resources too. I thought that this topic yolo object recognition would have much more following, instead there are really only a few projects. https://github.com/search?q=yolo+rtsp&type=repositories&s=forks&o=desc https://github.com/search?q=yolo+rtsp&type=repositories&s=fo...
- llm_trw 2y agoJust stream it one frame at a time to the model and eat the latency: https://www.youtube.com/watch?v=IHbJcOex6dk https://www.youtube.com/watch?v=IHbJcOex6dk if you need more hand holding. There's a reason why there's a whole family of models from tiny to huge.
- yeldarb 2y agoIf you do it naively your video frames will buffer waiting to be consumed causing a memory leak and eventual crash (or quick crash if you’re running on a device with constrained resources). You really need to have a thread consuming the frames and feeding them to a worker that can run on its own clock.
- llm_trw 2y agoThat's not how loop devices work on Linux.
- _giorgio_ 2y agoSorry for the newbie question Under windows, say that I have an RTSP stream (or something similar) Would you use a single python script with which one of this multithreading solutions? 1 import concurrent.futures 2 import multiprocessing 3 import threading
- _giorgio_ 2y agoThanks for the link, but what happens when you have a video stream, be it a usb webcam, or an RTSP stream, and the hardware can't keep up? I'm on windows. Ideally I'd like the frames to be dropped, so the inference is done on the last received frame? Is this a standard behaviour?
- yeldarb 2y agoWe’ve got an open source pipeline as part of inference[1] that handles the nuances (multithreading, batching, syncing, reconnecting) of running multiple real time streams (pass in an array of RTSP urls) for CV models like YOLO: https://blog.roboflow.com/vision-models-multiple-streams/ https://blog.roboflow.com/vision-models-multiple-streams/ [1] https://github.com/roboflow/inference https://github.com/roboflow/inference
- ferar 2y agoCan you specify ideal hardware (camera, computer) to deploy the solution? Thanks
- llm_trw 2y agoDefault yolo models are stuck at 640x640, so literally any camera that is at least capable of that resolution. Llava I believe is about the same. You'd need ubuntu and something that can run a llava model in vaguely real time, so a 4090/4080.
- moandcompany 2y agoYou'll want to find an IP Camera that supports the RTSP protocol, which is most of them. If your budget supports commercial style or commercial grade cameras, looking at Dahua or Hikvision manufactured cameras would be a good starting point to get an idea of specs, features, and cost.
- meow_catrix 2y agoMaybe don’t buy surveillance hardware from those brands
- sinuhe69 2y agoNot OP, but the reason may be: US - FCC Ban The US Federal Communications Commission (FCC) banned Dahua and Hikvision from new equipment authorizations in November 2022. Most products that use electricity require FCC equipment authorizations; otherwise, they are illegal to import, sell, market, or use, even for private individuals. Jul 5, 2024
- hcfman 2y agoShame, they are the best cameras available.
- formerly_proven 2y ago
- rocauc 2y agoA suggestion: I'd swap llava for Florence-2 for your open set text description. Florence-2 seems uniformly more descriptive in its outputs.
- jerpint 2y agoI found grounding-dino better than Florence and faster
- netdur 2y agoI found YOLOS to be faster and better, bot real time but 22k objects under half second
- Eisenstein 2y agoThey are using Ollama which is based on llama.cpp; florence is not supported on that backend.
- yu3zhou4 2y agoCongrats! What hardware you use to run the inference 24/7? I built a simpler version for running on low end hardware [0] for recognizing if there’s a person on my parcel, so I know someone have trespassed and I can launch siren, lights etc. https://github.com/jmaczan/yolov3-tiny-openvino https://github.com/jmaczan/yolov3-tiny-openvino
- pmontra 2y agoThis runs with a Geforce GTX 1060. By a quick search it's 120 W. Maybe it's only the peak power consumption but it's still a lot. Do commercial products, if there are any, consume that much power?
- hcfman 2y agoI have something similar. It's not tracking though. Drawing around 10W on a pi, around 7W on a Jetson.
- 4ggr0 2y agonot sure if i'm misunderstanding - you've got a similar GPU to a 1060 hooked up to a pi?
- lelag 2y agoOP is probably using an AI accelerator like this: https://coral.ai/products/accelerator https://coral.ai/products/accelerator which works great on a PI and uses very little power. It will do the Yolo part, but you can't really expect it to do the multimodal LLM part, although you could try to run Florence directly on the PI too.
- mobilemidget 2y agoThis works better in my experience, https://hailo.ai https://hailo.ai https://www.raspberrypi.com/news/raspberry-pi-ai-kit-available-now-at-70/#comment-1597156 https://www.raspberrypi.com/news/raspberry-pi-ai-kit-availab...
- synergy20 2y agocoral has pcie module which is 1/4 to 1/3 of the price
- hcfman 2y agoNot a pi. A Jetson. Still an arm SBC though.
- nikolayasdf123 2y agohow about llama3.2 vision? should it get better performance?
- vaylian 2y agoHello from the privacy crowd! Please use this responsibly. Tech can be a lot of fun and I encourage you to play around with things and I appreciate it when you push the boundaries of what is technically feasible. But please be mindful that surveillance tech can also be used to oppress people and infringe on their freedoms. Use tech for good!
- deleted 2y ago[deleted]
- doctorhandshake 2y ago>> It calculates the center of every detection box, pinpoint on screen and gives 16px tolerance on all directions. Script tries to find closest object as fallback and creates a new object in memory in last resort. You can observe persistent objects in /elements folder I’ve never implemented this kind of object persistence algo - is this a good approach? Seems naive but maybe that’s just because it’s simple.
- xrd 2y agoI'm confused about why you need yolo and llava. Can't you simply use yolo without a multimodal LLM? What does that add? You can use yolo to detect and grab screen coordinates on its own, right?
- michaelt 2y agoAlmost certainly using yolo to segment the cars, then llava for the more detailed "silver sedan" description
- andblac 2y agoSkimming through the source it seems to run 'car' and 'person' objects through llava with the following prompt: - "person": "get gender and age of this person in 5 words or less", - "car": "get body type and color of this car in 5 words or less". So YOLO gives the bounding box and rough category, while llava describes the object in more details.
- anshumankmr 2y agoCould try with Florence by Microsoft instead of Yolo and Llava, though the results are not going to be as great. Florence will do the inference on CPU. This is just for fun.
- 01100011 2y agoIf you're interested in DIY security+AI, check out Frigate NVR(https://frigate.video/ https://frigate.video/), Scrypted(https://www.scrypted.app/ https://www.scrypted.app/) and Viseron(https://viseron.netlify.app/ https://viseron.netlify.app/).
- gh02t 2y agoI've been using Frigate for a long time and it's a really cool project that has been quite reliable. The configuration can be a little bit of a headache to learn, but it gets better with every release. Viserion is new to me though, that looks really cool.
- dfc 2y agoI just recently got frigate up and running. How do the other two compare?
- 01100011 2y agoBeats me, I'm just getting into this now. I started with a Reolink NVR, but it's a piece of crap, so I'm looking for a better alternative. It looks like either Frigate or Viseron will do what I want. I started setting up Frigate, but realized I should downgrade my Reolink Duo 3 to a Duo 2 before I go too far. The Duo 3 really doesn't offer much better image quality but forces you to use h265 and consumes a lot more bandwidth. Once I stabilize my camera setup I'll get back to setting up both Frigate and Viseron and see what performs better. I like that the pro upgrade of Frigate allows you to customize the model and may make use of that.
- taikon 2y agoI've been running frigate for a while now and I find it's object detection has a higher than preferred false-positive rate. For instance, it kept thinking the tree in my back yard is a person. I find it hilarious that it often assigns a higher likelihood the tree is a person than me! I've needed to put a mask over the tree as a last resort.
- 2y ago
- matrik 2y agoMobileNetV3 and EfficientDet are othwr possible alternatives to YOLO. I was able to get higher than 1.5 FPS on Raspberry Pi Zero 2W which draws 1W on average. With efficient queuing approach, one can eliminate all bottlenecks.