8 ms·
Tensorflow on edge, or – Building a “smart” security camera with a Raspberry Pi
- brutus1213 6y agoNice writeup but the Raspberry Pi isn't running tensorflow. It is mentioned in the article that the author is sending images to an edge machine. The big question I had was about hardware video encoding/decoding ... doesn't really cover that. I've found sending single image frames over zeromq to be fairly limiting if you care about high frame rate/low latency processing. Key issue I have run into is while many chips support hardware video encoding/decoding, the APIs to interface with this aren't there or not in open source. Anyone who has ideas on this, I'd welcome your comment. As an aside, another option is to run Intel's Movidus USB stick (aka Neural compute stick) and then you get a smart camera on the raspberry pi itself. That raises other issues though.
- snowzach 6y agoShameless plug, check out DOODS: https://github.com/snowzach/doods https://github.com/snowzach/doods It's a simple REST/gRPC API for doing object detection with Tensorflow or Tensorflow Lite. It will run on a Raspberry Pi. It actually did support the EdgeTPU hardware accelerator to make the Pi pretty quick for certain models. They broke something so I need to fix EdgeTPU support but it's still usable on the Pi withe the mobilenet models or inception if you're not in a hurry.
- agibsonccc 6y agoFew questions: 1. Did you build this for your own use cases? Interesting side project? 2. How do you feel about the need for base64 being a requirement on the endpoints? Isn't GRPC the wrong medium for this? Also, what do you see as the main limitations right now? The models?
- snowzach 6y ago1. I built it to integrate with Home Assistant and security systems. I was trying to use Tensorflow on a Raspberry Pi and the dependencies were a nightmare. Tensorflow in general is a nightmare to compile and run IMO. I got to thinking, what if I could make all the deps inside of a docker container. What if I could run it remotely. It was born out of that. 2. As for base64, I'm not sure of a better way to support sending raw image data over JSON (in REST mode) In some ways I think GRPC is a better medium than JSON (it supports either) as GRPC supports sending the RAW bytes. What leads you to believe GRPC isn't the right transport? Plus you can do it in a stream format if you want to do a lot of video. The only limitations I can think of are that Tensorflow supports a myriad of CPU optimizations so providing a single container image that has all the right options is basically impossible. I created one that has what I think are some of the better options (AVX, SSE4.X) and then an image that basically should run on any 64 bit intel compatible CPU. To get optimized options you need to build the docker container yourself which can take the better part of a day on slower CPUs. With that said, I also provide ARM32 and ARM64 containers that actually run semi-okay on Raspberry Pis and and other ARM SBCs. I can run the inception model on a Pi4 on a 1080p image in about 5 seconds which is pretty good IMO.
- barake 6y agoCurious what you found limiting about zeromq? Just not enough throughput for high FPS? So far I’ve found it to be the sanest multicast solution since clients pull.
- brutus1213 6y agoIssue isn't ZeroMQ. The simple/inefficient way to do it is to capture frames one at a time, and send them via ZeroMQ. Video is pretty bandwidth intensive .. the only reason things like YouTube work as smooth as they do is that they use codecs such H264/265 (which are proprietary unfortuantely) and stream compress frames over the network. Now doing the codec in software burns a lot of CPU as this is very math intensive .. most processors support hardware video codecs for this purpose. There are just no open source tools/libraries that make this good/simple enough that I have found.
- ww520 6y agoAren't IP cameras capable of H264 already encoded the video before output it to the network? H264 video stream has very good compression ratio and shouldn't consume too much bandwidth.
- brutus1213 6y agoIn the project I used zeromq, we were not using external IP cameras. My experience with IP cameras is still full of some seconds of latency .. I have no idea why.
- thebruce87m 6y agoThis is where GStreamer normally steps in. A lot of hardware manufacturers provide a gstreamer plugin for their module. I’ve had experience with NVIDIA and atmel SoCs and that seemed to be the default path. Good luck with the gstreamer pipeline learning curve however!
- NikolaeVarius 6y agoCan confirm not fun
- agibsonccc 6y agoCould you elaborate on some of the problems you had overall?
- brutus1213 6y agoI've unsuccessfully dabbled in gstreamer in the past. I was doing a project this weekend, and the comments on this thread motivated to give it another shot .. after a couple of hours (2-4ish?), I was able to get video off the Pi to my desktop (on the same LAN) but the performance was pretty bad. I didn't optimize much yet but let me summarize the key issues I experienced with gstreamer these last few hours: 1) Very little documentation; poorly explained pipelines. I tried to read what docs I could find but things quickly devolved into trying out random gstreamer pipelines posted in comments. People don't explain why they use one particular element over another. So it felt like whack-a-mole. 2) Installing gstreamer on the Pi was a breeze. I wanted to pull video off the connected camera and sent to VLC on my desktop. Sounded like something that would work out-of-the-box? Nope. Kept seeing lots of stackoverflow comments of people stabbing in the dark, getting errors (or have the thing just sit there and not work) with very little feedback on what was wrong. 3) I have very little indication of what is hardware and what is software accelerated in my pipeline. I have no idea where latency is coming into my pipeline. Overall .. my modern expectation for software frameworks is "batteries included" .. it is totally reasonable for sophisticated software tools to be complex .. but gstreamer is just not designed that way. While I got it to work, I see massive latency (likely because my pipeline is inefficient) and degraded quality (no idea why).
- gambiting 6y agoYeah I was hoping it was a raspberry pi maybe using one of these neural nets USB sticks, instead it was just using RPi as a dumb terminal for sending video. You could probably do the same with an old android phone set to stream video over lan.
- quietbritishjim 6y ago> Nice writeup but the Raspberry Pi isn't running tensorflow. It is mentioned in the article that the author is sending images to an edge machine. Yeah I was a bit surprised by this, and although the article is very clear about it I think it's generated a bit of confusion in the comments here. My understanding of edge computing is that it means the processing of data is done at the point the data is captured, so to me that would mean right there on the raspberry pi. But the author considers their whole LAN to be the "edge", so basically anything that doesn't involve sending the data over the internet: > ... doing the heavy lifting on a machine physically close to the edge node – in this case, running the Tensorflow Object detection. By doing so, we avoid roundtrips over the internet, as well as having to pay for Cloud compute on e.g., AWS or GCP. I think their strategy of capturing the data on a very low-power device and then processing on a server on your network is a very reasonable one, I just wouldn't have used that term.
- theblackcat2004 6y agoSince you already using openCV, you can write a neat motion detection and only start sending frames for detection when motions are detected
- simlevesque 6y agoGood thinking.
- otter-in-a-suit 6y agoJust saw this thread (I'm the author) - great idea, thank you!
- thebruce87m 6y agoIs it really edge computing if the pi isn’t running tensorflow? I know the definition is kind of woolly. I wonder what the performance would be on a $100 jetson nano.
- brutus1213 6y agoJetson nano can run Cuda code. It is pretty decent. The article does employ a reasonable definition of edge computing IMO (scientist who works in this area). The rpi is the client, and the processing happens on a beefy edge node. But yeah .. there is not one clear, accepted definition here.
- dnautics 6y agoit's not going to the cloud, so I'd say that counts.
- hn_check 6y agoThe Jetson Nano is fantastic, and I use Tensorflow on it with two home security cameras. It is a wonderful device, and deserves far more attention than it gets. I'm going to upgrade in the next week to a Jetson Xavier NX. Not because I need to, but because I like playing around and it's a silly powerful device. I also run a NextDNS CLI client on it, various automation stuff, etc.
- deleted 6y ago[deleted]
- simlevesque 6y agoTensorflow on the pi itself is hard but I get great results for a similar system with just a rpi4 and opencv.
- alkonaut 6y agoWould object detection like this work out of the box for deer like he demonstrates for humans? I need this for deer.
- agibsonccc 6y agoHi, could you describe your use case a bit? Just an alarm trigger for deer in the backyard?
- alkonaut 6y agoYes, I'd point a camera at my precious vegetables and if a deer walks into the video feed, something that scares it off is triggered so it runs off before eating the whole garden.
- joshu 6y agoI've just started down this path using a plain RTSP-serving camera and a low end box using the Coral EdgeTPU to process the frames. It looks like there are a variety of solutions available. https://github.com/blakeblackshear/frigate https://github.com/blakeblackshear/frigate https://docs.ambianic.ai/users/configure/ https://docs.ambianic.ai/users/configure/ etc
- fareesh 6y agoI am trying to achieve something similar but at a higher scale. I have about 48 different cameras where I want to count people and get their approximate location in the frame. I want to run an object detection model on all of those video streams simultaneously. My AWS instance maxes out after 7 simultaneous streams so I figured I don't really need real-time monitoring. One frame every couple of seconds, even every minute could potentially suffice, since I am dealing with larger time-frames. Since I don't want to run too many instances at the same time, what are some viable strategies to achieve this? My plan is to have 5-6 instances of the ML model loaded up and waiting to accept a frame. When one of them is ready, it will instruct one of the RTSP streams to send it a frame, which it will process and store / send the result to an application server. I feel like I may not even be able to consume so many RTSP streams at once (I've never tried so I don't know), so I may have to have some other method of priming the handshake etc. before the model asks for a frame to process. Is there a better / non-hacky way of achieving this (i.e. managing the workload on a single GPU instance) ? I don't have any control of the camera hardware at all.
- agibsonccc 6y agoHi Fareesh, I'd love to hear more about your use case. Email's in my profile.
- shiftpgdn 6y ago48 RTSP streams is a lot of bandwidth to consume at once. Why not use an edge PC or Jetson system to do it in small blocks? A new Jetson Xavier NX can do 8-12 streams depending on FPS and model.
- xchip 6y agothe Raspberry Pi isn't running tensorflow
- staycoolboy 6y agoI've done this with a Jetson Xavier, 4 CCTV cameras and a PoE hub. You really want to use DeepStream and C/C++ for inference, not Python and TensorFlow. I'm streaming ~20 fps (17 to 30) 720P directly from my home IP4 address, and when a person is in-frame long enough and caught by the tracker, a stream goes to an AWS endpoint for storage. I've experimented with both SSDMobileNet and Yolo3, which are both pretty error prone but they do a much better job filtering out moving tree limbs and passing clouds, unlike Arlo. You need way more processing power than an RPi to do this at 30fps, and C/C++, not Python. (There are literally dozens of projects for the RPi and TFlow online but they all get like 0.1 fps or less by using Flask and browser reload of a PNG... great for POC but not for real video) I wrote very little of the code, honestly: only the capture pipe required a new C element. I started with NVidia DeepStream which is phenomenally well-written, and their built-in accelerated RTSP element, and added a custom GStreamer element that outputs a downsampled MPEG capture to the cloud when the upstream detector tracks an object. NVidia also wrote the tracker, you just need to provide an object detector like SSDMobileNet or YOLO. NVidia gets it. The main 4 camera-pipe mux splits into the AI engine and into a tee to the RTSP server on one side and my capture element on the other side. It was amazingly simple, and If I turn the CCD cameras down to 720P with h265 and a low bitrate, I don't need to turn on the noisy Xavier fan. The onboard Arm core does the detected downsampling (one camera only, a limitation right now) and pushes the video with a rest endpoint on a node server in the AWS cloud. I'm very pleased with it, I haven't tested scaling but if I turned off the GPU governors I could easily go to 8 cameras. I went with PoE because WiFi can't handle the demand.
- exhilaration 6y agoDo you have any links you could share to build something like this?
- reichardt 6y agoTensorFlow Lite with SSDLite-MobileNet gets you around 4 fps on a Raspberry Pi 4 (23 fps with a Coral USB Accelerator): https://github.com/EdjeElectronics/TensorFlow-Lite-Object-Detection-on-Android-and-Raspberry-Pi/blob/master/Raspberry_Pi_Guide.md https://github.com/EdjeElectronics/TensorFlow-Lite-Object-De...
- szczepano 6y agoesp32-cam microcontroller costs around $6 and you have face detection build in. It have bluetooth and wifi and most of it's drivers code if not everything is on github. Only problem is you need to program it using arduino or other microcontroller hardware.
- ed25519FUUU 6y agoFace detection is nice but body/person detection is much more useful in these setups.
- szczepano 6y agoYou mean esp-who ? It uses mobilenetv2 so it's quite possible to train it to detect person instead of face. Didn't tried myself, just started playing with it.
- noja 6y agoOff topic almost: does anyone know of any (long life) battery powered wifi cameras (with IR) for a project like this? Off the shelf, with a battery life of months and nice looking (like Arlo) but not cloud?
- milofeynman 6y agoFor people running Blue Iris you can do something similar to this with blue iris, deepstack[0], and an exe[1] someone wrote that sends the images to deepstack. Video guide: https://youtu.be/fwoonl5JKgo https://youtu.be/fwoonl5JKgo (Links in the comments of video as well) [0] https://deepstack.cc/ https://deepstack.cc/ [1] https://ipcamtalk.com/threads/tool-tutorial-free-ai-person-detection-for-blue-iris.37330/ https://ipcamtalk.com/threads/tool-tutorial-free-ai-person-d... https://github.com/gentlepumpkin/bi-aidetection https://github.com/gentlepumpkin/bi-aidetection
- deleted 6y ago[deleted]
- canada_dry 6y agoI'm hoping advances like YoloV5 [i] will allow a rpi4 to more ably do this without piping the video to another processor. [i] https://github.com/ultralytics/yolov5 https://github.com/ultralytics/yolov5
- dheera 6y agoCan we all please stop using the term "edge" computing? It's nothing but a hype term and in reality it's really what we already had for the decades before the internet.
- hikarudo 6y agoI disagree. The term "edge computing" actually adds precision to a description of a distributed system. Nowadays, with a lot of machine learning inference happening on the cloud, when seeing the term "edge inference" you immediately know you don't have to send heavy bandwidth-clogging video streams to the cloud. Inference on the edge is a clear trend in computer vision applications, now that we each year there are better low-power neural network accelerators.
- dheera 6y ago> Nowadays, with a lot of machine learning inference happening on the cloud Right, and if it's not on the cloud, it runs locally, as everything did before "cloud" became popular. We don't need to call it "edge" just to raise VC money or put out some PR. We can just say it runs locally, on-device, etc. If (big if) and when Adobe realizes that their Creative Cloud was a bad idea, are they going to call the next product "Adobe Edge Edition! Wow you can actually run PhotoShop on your own desktop!"?
- gspr 6y ago> Right, and if it's not on the cloud, it runs locally, as everything did before "cloud" became popular. We don't need to call it "edge" just to raise VC money or put out some PR. We can just say it runs locally, on-device, etc. To me, "edge" means more than just "not cloud". It's appropriately used when making the point that computations happen where the data is gathered and the output is required (which seems actually not to be the case in TFA, but still). It's when computations are not offloaded elsewhere at all, not just "not to the cloud".
- dheera 6y ago
- acidburnNSA 6y agoThis is extraordinarily neat. Home Assistant does have a tensorflow integration [1] that allows you to run other home assistant automations (including various alerts, alarms, and scare sequences) based on person detection with basically any camera (since it's kind of a hub-and-spoke model to all other possible IoT devices). [1] https://www.home-assistant.io/integrations/tensorflow https://www.home-assistant.io/integrations/tensorflow I struggled recently to get it running on my actual GPU since I run Home Assistant on a home server. I ended up making a custom component using pytorch instead on Pop OS 20.04 and it works gloriously. CPU usage way down and GPU has something to do now. My super awesome self-hosted alarm system is now extra-super awesome. Of course burglars are going to all just start wearing AI adversarial t-shirts.
- anp 6y agoThis reminds me of a hypothetical project I would take up if I still had a dog and a small yard: building a poop cleanup map from CV processing of camera footage. Stepping stone toward a Poopba, obviously.
- b34r 6y agoWhat about package delivery people lol
- 9nGQluzmnq3M 6y agoCoral's Edge TPU products are built specifically for this kind of thing: https://coral.ai/ https://coral.ai/ Hands-on video (4 min): https://www.youtube.com/watch?v=-RpNI4ZrfIM https://www.youtube.com/watch?v=-RpNI4ZrfIM
- monkeydust 6y agoInteresting have RPi's lying around might get the USB Accelerator
- josteink 6y ago> We’ll use a Raspberry Pi 4 with the camera module to detect video. ... Now, here’s an issue for you: My old RasPi runs a 32bit version of Raspbian. So why not just use the 64-bit Ubuntu RPi image instead then? https://ubuntu.com/download/raspberry-pi https://ubuntu.com/download/raspberry-pi
- hathym 6y agoIt's google edge by the way.
- 8fingerlouie 6y agoI did something similar, but because i had no requirement to playback audio "real time", i opted for a simpler solution. I run a simple video capture from a Raspberry Pi Zero W running motion, meaning all motion events are captured, including leaves blowing in the wind. The captured files are stored on a NFS share per camera. On the server i then monitor the parent directory for every camera for new files, and run my object detection there, which in turn generates push notifications with a screengrab if certain objects are detected. It also stores a bounding box annotated version of the file. Not really needed except for figuring out why you got an alert without any clear reason. doing it this way however allows me to save a bit on each camera, and use dedicated hardware for object detection on the server. I currently use an Intel Neural Compute Stick 2 (https://software.intel.com/content/www/us/en/develop/hardware/neural-compute-stick.html https://software.intel.com/content/www/us/en/develop/hardwar...), and while it is far from dedicated GPU performance, it is equally far from dedicated GPU power consumption.