26 ms·
Yoloface-500k: ultra-light real-time face detection model, 500kb
- bawana 6y agoWe are all essentially evil because 1) tomorrow will find some behavior that is accepted today as bad and we are all doing it 2) creating an AI that can manipulate people is too tempting for the predator in us to avoid (this is a REAL turing test - making an AI that can make tools out of humans and everything else.) The only way to stay ahead of the 'evil' corrupting influence of new tech is to prevent it from widespread use controlled by a single entity. So, yolo is ok as long as you cannot deploy it in a cloud at scale. So, just as nuclear weapons (the massive concentration of energy release at tremendously fast rates) are bad, so is a super AI/AGI (the massive computational ability at nanosecond scale). No evil was ever perpetrated by institutions of learning - only business entities and governments who scaled up those discoveries caused evil. And now for the flame bait- So, by this argument we should elect Luddites to govern us , especially ones that are not imaginative or creative.
- yoloyoloyolo 6y agoFreedom is slavery.
- johnghanks 6y agowhat
- ddrt 6y agoI shudder for the future.
- stu2b50 6y agoYolo does bounding box detection on faces. Classifying those faces is something else's job.
- jakear 6y agoXolo does feature vector extraction on face regions. Extracting those regions from larger images is something else’s job. Zolo does feature vector similarity analysis on facial feature vectors. Extracting those vectors is something else’s job. Oh look, now every authoritarian government has free access to never before seen levels of data harvesting. But nobody has to feel any guilt because they only contributed a third of the machine. Hooray!
- zo1 6y agoThe future is coming, and the tech that can enable authoritarian behavior will come no matter what we do as they're just tools. It's been here and around us in ever-increasing forms for decades and yet we haven't necessarily devolved into a big brother state. What we should be worrying is actual usages of grand scale citizen-control and monitoring projects that are enabled by technology. Think China, not UK or US.
- longtom 6y ago> and yet we haven't necessarily devolved into a big brother state. Well this is damn close. It just needs an executive apparatus which is where drones surely come in handy. https://en.wikipedia.org/wiki/Global_surveillance_disclosures_(2013%E2%80%93present) https://en.wikipedia.org/wiki/Global_surveillance_disclosure...
- 8fingerlouie 6y agoAnd yet, a $0.50 facemask will completely destroy the surveillance.
- eunos 6y agoNah, gait based subject recognition might come in the future.
- 6y ago
- rocauc 6y agoPurpose-built, small, and fast models appears to be the inevitable evolution for computer vision. Where can the "Easy Set, "Medium Set, and "Hard Set" evaluations referenced in the "Wider Face Val" be found?
- qiuqiu-dog 6y ago## Wider Face Val Model|Easy Set|Medium Set|Hard Set ------|--------|----------|-------- libfacedetection v1(caffe)|0.65 |0.5 |0.233 libfacedetection v2(caffe)|0.714 |0.585 |0.306 Retinaface-Mobilenet-0.25 (Mxnet) |0.745|0.553|0.232 version-slim-320|0.77 |0.671 |0.395 version-RFB-320|0.787 |0.698 |0.438 yoloface-500k-320|0.728|0.682|0.431|
- rocauc 6y agoThanks, I see the table. Are the source datasets available for creating additional benchmarks?
- haditab 6y agoI believe this is exactly why pjreddie quit computer vision research. It must kill him to see such projects based off of his work.
- jonex 6y agoCould you elaborate? What is the problem with the linked project? Training a slightly faster, smaller and less accurate version of an existing model?
- chvid 6y agohttps://pjreddie.com/darknet/yolo/ https://pjreddie.com/darknet/yolo/ Is that pjreddie used horses, dogs, and bicycles as training data? Not realising that his technology could also be used on human faces?
- viraptor 6y ago> Not realising that his technology could also be used on human faces? I'm not sure how you got to this idea, but it's just not plausible.
- mlyle 6y agoHe said: "I stopped doing CV research because I saw the impact my work was having. I loved the work but the military applications and privacy concerns eventually became impossible to ignore." Initial good results on CV doesn't mean that you realize all the ways it'll start to be used and the implications thereof.
- viraptor 6y agoThis is about implications - that's not the same as not realising that face is an object.
- atoav 6y agoIMO this is ethically more simple than many like to believe. There are existing power structures (planned ones and emerged ones) in this world and the technology we create can either be used to reinforce them or to question them. Sometimes it is both and things cancel each other out and move on a sideways trajectory – but in the case of CV, it is quite clear who will benefit: those in power, those who need to quantify, control and punish the human element, but don't have the manpower (=legitimacy?) or funds (=priority?) to do so manually. I get that working in CV is interesting and cool stuff, but the collective suffering it might help creating and keeping is something one should seriously think about as well.
- hirundo 6y agoImagine an app like Pokémon VR but instead of virtual pocket monsters it targets members of <outgroup> whose faces are detected by cadre smart phones, then tracked, flash mobbed and dealt with, for the crime of making members of <ingroup> feel unsafe. Watching leadership supine and carefully uncritical of burning, looting mobs offers little confidence that they will stand in the way of this. After all only <outgroup epithet>s have anything to fear. Please tell me why this is an unlikely scenario.
- jcahill 6y agoThe submission is an implementation of a core task in computer vision. Your response is an appeal to vividness unrelated to the submission except in its recruitment of computer vision to sell the FUD. If researchers getting good at something is sufficient priming to cause you to direct your imagination toward hyperbolically negative outcomes, the problem on your hands is a constitutional resistance to further progress in the research area. In that case, challenging readers to produce arguments on the finer details of the narrative you've painted in support of the technopessimism is bad faith rhetoric.
- Kiro 6y agoI don't understand why you decided to respond to an almost-dead comment when the top-voted comment thread shows the same technopessimism.
- jcahill 6y agoI'm similarly unclear on what sort of response you expect. There weren't many comments when I replied. My criticism was tailored to a fairly specific phenomenon: asymmetrically imaginative doomsaying that appeals to a vivid vignette / sketch of an adjacent possible future featuring some hyperbolically elaborated extension of trending tech, like Flash Mob Gone Wrong[1] and Slaughterbots[2]. HN flagging and points are irrelevant to me. ____________________ [1]: https://youtube.com/watch?v=RyMdOT8YJgY https://youtube.com/watch?v=RyMdOT8YJgY [2]: https://en.wikipedia.org/wiki/Slaughterbots https://en.wikipedia.org/wiki/Slaughterbots
- layoutIfNeeded 6y agoWew, 500kb is ultra-light nowadays. I wonder how much space would the original Viola-Jones face detector take.
- srg0 6y agoCheck out https://github.com/opencv/opencv/tree/master/data/haarcascades https://github.com/opencv/opencv/tree/master/data/haarcascad... Plain-text XML for the frontal face detector is 912 KB. 132 KB gzipped. It should be smaller in binary.
- rjeli 6y agoWow, 100MFlop’s. That could run real time on a $5 dsp.
- chvid 6y agoI think it is very cool. Trying to think of some applications for this. For example one could create a mechanism that watched people entering and exiting a shop providing the shop owner more quantitative data that he could use to optimize his sales. Or you could have it watch a soccer game. Generating all sorts of data on how the game went. All on relative cheap piece of hardware.
- andrewnc 6y agoYou could make this a commodity https://andrewnc.github.io/projects/projects.html#heart-rate https://andrewnc.github.io/projects/projects.html#heart-rate
- mycall 6y agoEntering/exiting buses for automatic passenger counters is more important than ever now. Being able to broadcast GTFS-Occupancy in real-time when only 50% (or less) of the bus can be filled with passengers, is a real issue transit is facing today.
- Tempest1981 6y agoI'm looking for something to run on a Raspberry Pi, to detect humans on a security camera. The built-in camera software has false triggering, esp. on windy days. When looking at these projects, how do I figure out what hardware they're aimed at? This one mentions NVidia/CUDA. Is there any sort of hardware abstraction layer that YOLO or R-CNNs can operate on? Can I use any of this code (or models) for my R-Pi?
- dmm 6y agoChezck out the Frigate project. It uses the Coral tpu accelerator which could be used with the rpi. https://github.com/blakeblackshear/frigate https://github.com/blakeblackshear/frigate
- w_t_payne 6y agoYou might want to investigate the Movidius Neural Compute stick (which you can use with a RPi), or the Nvidia Jetson Nano, which has a lot more oomph.
- 8fingerlouie 6y agoI use the Movidius NCS and it's pretty much the optimal solution for light OpenCV work. It draws very little power, but as the Pi itself is powered by USB, don't expect to run much else off of USB on it. I use it in the exact same scenario, a Raspberry Pi (Zero W) with a camera with motion detection and notifications on movement, my implementation may be specific though. Each of my Raspberry Pi cameras runs motion (https://motion-project.github.io/index.html https://motion-project.github.io/index.html), and recorded files are stored on a NFS share. Each camera has it's own directory within this share (or rather each camera has it's own share within a parent directory), and the server then runs a python script that monitors for changed/added files, and runs object detection on the newly created/changed files. If a person is detected in the file, it then proceeds to create a "screenshot" of the frame with the most/largest bounding box, and sends a notification through Pushover.net including the screenshot with bounding box. There implementation is not quite as simple as described here, i.e. i use a "notification service" listening on MQTT for sending pushover notifications, but the gist of it is described above. Edit: I should probably clarify that my cameras are based on Raspberry Pi Zero W. They have enough power to run motion at 720p - at around 30fps. Not great, but good enough for most applications. I've since migrated most to Unifi Protect instead. A little higher hardware cost, a lot better quality :)
- m0zg 6y agoThere's also Blazeface: https://arxiv.org/abs/1907.05047 https://arxiv.org/abs/1907.05047, which the authors do not seem to mention.
- DSingularity 6y agoI love how every new YOLO project inevitably leads to the discussion of the ethics. At the very least more people will be wondering if they should also be taking ethics into consideration wrt their lines of work. Pjreddie is a giant for this. It is a real contribution.
- rgrieselhuber 6y agoThis also puts social distancing into perspective.
- sp332 6y ago"Bflops"? I'm guessing this is a measure of the total processing power needed, in billions of floating-point ops, and not a measure of operations per second?
- ta1234567890 6y agoIs it possible to run this "in reverse" so it generates faces instead of detecting them? If so, how?
- phonebucket 6y agoThere are much better ways to generate faces than using this, e.g. https://github.com/tkarras/progressive_growing_of_gans https://github.com/tkarras/progressive_growing_of_gans
- ta1234567890 6y agoThank you for the link. I'm still curious about the possibility of taking a detection algorithm/network and just running it in reverse. Is that feasible? Are there people doing it?