12 ms·
YOLOv5: State-of-the-art object detection at 140 FPS
- rocauc 6y agoEfficientDet was open sourced March 18 [1], YOLOv4 came out April 23 [2], and now YOLOv5 is out only 48 days later. In our initial look, YOLOv5 is 180% faster, 88% smaller, similarly accurate, and easier to use (native to PyTorch rather thank Darknet) than YOLOv4. [1] https://venturebeat.com/2020/03/18/google-ai-open-sources-efficientdet-for-state-of-the-art-object-detection/ https://venturebeat.com/2020/03/18/google-ai-open-sources-ef... [2] https://arxiv.org/abs/2004.10934 https://arxiv.org/abs/2004.10934
- KMnO4 6y agoThose numbers are quite impressive. YOLOv4 -> YOLOv5 Inference time: 20ms -> 7ms (on P100) Frames per second: 50 -> 140 Size: 244mb -> 27 mb
- WanderPanda 6y agof(x)=c, zero size, infinite fps. You should also take some accuracy metric into account ;)
- yodon 6y ago> "Similarly accurate"
- normanluhrmann 6y agoMS COCO looks to be improved even overall. OT: the stats above should be part of the PyTorch marketing material, indeed impressive
- eeZah7Ux 6y ago> open sourced This is not a verb.
- sudosysgen 6y agoArguably it became one.
- dang 6y agoIt's behaving like one: https://hn.algolia.com/?dateRange=all&page=0&prefix=true&query=open%20sources&sort=byDate&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que... https://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=open%20sourced&sort=byDate&type=story&storyText=none https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
- gok 6y agoEr so this "Ultralytics" consulting firm just borrowed the name YOLO for this model and didn't actually publish their results yet?
- yeldarb 6y agoYeah, they made the most popular PyTorch implementation of YOLOv3 as well so they're not entering out of the blue, though. https://github.com/ultralytics/yolov3 https://github.com/ultralytics/yolov3 The author of YOLOv3 quit working on Computer Vision due to ethical concerns. YOLOv4, which built on his work in v3, was released by different authors last month. I'd expect more YOLOvX's from different authors in the future. https://twitter.com/pjreddie/status/1230524770350817280 https://twitter.com/pjreddie/status/1230524770350817280
- deleted 6y ago[deleted]
- sonofaragorn 6y agoI'm a bit fascinated by this Ultralytics. It has super nice website but according to LinkedIn, I think it's just one guy who does consultancy. The intriguing part is that he has also done research in particle physics (as Ulatrlytics) that has been published in Nature [1]. I had never seen anything like that. [1] https://www.nature.com/articles/srep13945 https://www.nature.com/articles/srep13945
- rememberlenny 6y agoTwo interesting links from the article: 1. How to train YOLOv5: https://blog.roboflow.ai/how-to-train-yolov5-on-a-custom-dataset/ https://blog.roboflow.ai/how-to-train-yolov5-on-a-custom-dat... 2. Comparing various YOLO versions https://yolov5.com/ https://yolov5.com/
- bcatanzaro 6y agoWhy benchmark using 32-bit FP on a V100? That means it’s not using tensor cores, which is a shame since they were built for this purpose. There’s no reason not to benchmark using FP16 here.
- joshvm 6y agoNot sure about the benchmark, but the code includes the option for mixed precision training via Apex/AMP.
- bcatanzaro 6y agoIf you click around enough you’ll see they benchmarked in 32-bit FP. Glad they have a mixed precision training option but I really think it’s a mistake in 2020 to do work related to efficient inference using 32-but FP. The problem is that your conclusions aren’t independent of this choice. A different network might be far better in terms of accuracy/speed tradeoffs when evaluated at a lower precision. But there is no reason to use 32-but precision for inference, so this is just a big mistake.
- hnarayanan 6y agoWhat does it take to now use this name?
- ulam2 6y agoJust guts, i guess. You would also need to show some real performance improvement.
- bonoboTP 6y agoIf it's not trademarked, perhaps not much? I think it's pretty misleading, but the fight for attention is on! Using an established brand in your title will get more clicks.
- newen 6y agoYeah, it's pretty unethical. Looks like they just stole the name without any care. There doesn't seem to be any relationship between these guys and the original YOLO group.
- 0xcoffee 6y agoIs it possible to run these models in the browser, something similar to tensorflow.js?
- m00dy 6y agoI would try convert to ONNX model and then try to infer with tensorflowjs.
- osipov 6y agoJust recently IBM announced with a loud PR move that the company is getting out of the face recognition business. Guess what? Wall Street doesn't want to keep subsidizing IBM's subpar face recognition technology when open source and Google solutions are pushing the state of the art.
- ahelwer 6y agoNot something to brag about. Facial recognition has very few applications outside of total surveillance. We should not respect those who lend it their time and effort.
- monocasa 6y agoI thought the real focus on the bad actors at this point was on gait detection. Works in civil unrest situations where everyone covers their face. Not the the difference matters that much.
- anewdirection 6y agoIts not exclusive. Bad actors are working on whatever they are paid to build, by other bad actors with less technical acumen and more money. Edit: I should add, that most of the actual progress is being made by smart people who think its an interesting problem and are unaware or uncaring of the clear outcome of such tech.
- monocasa 6y agoIn this case I meant bad actors as in who is funding the research with the idea of increasing surveillance for the purpose of squashing dissent.
- jcims 6y ago>Facial recognition has very few applications outside of total surveillance. That's not really for you to decide, is it? You're absolutely free to have that opinion of course. >We should not respect those who lend it their time and effort. Also your choice of course. Facial recognition is essentially a light integration of powerful underlying technologies. Should 'we' ostracize those working on machine learning, computer vision, network and distributed computing, etc?
- sillysaurusx 6y agoWe made a site that lets you collaboratively tag a bunch of images, called tagpls.com. For example, users decided to re-tag imagenet for fun: https://twitter.com/theshawwn/status/1262535747975868418 https://twitter.com/theshawwn/status/1262535747975868418 And the tags ended up being hilarious: https://pbs.twimg.com/media/EYXRzDAUwAMjXIG?format=jpg&name=large https://pbs.twimg.com/media/EYXRzDAUwAMjXIG?format=jpg&name=... (I'm particularly fond of https://i.imgur.com/ZMz2yUc.png https://i.imgur.com/ZMz2yUc.png) The data is freely available via API: https://www.tagpls.com/tags/imagenet2012validation.json https://www.tagpls.com/tags/imagenet2012validation.json It exports the data in yolo format (e.g. it has coordinates in yolo's [0..1] range), so it's straightforward to spit it out to disk and start a yolo training run on it. Gwern recently used tagpls to train an anime hand detector model: https://www.reddit.com/r/AnimeResearch/comments/gmcdkw/help_build_an_anime_hand_detector_by_tagging/ https://www.reddit.com/r/AnimeResearch/comments/gmcdkw/help_... People seem willing to tag things for free, mostly for the novelty of it. The NSFW tags ended up being shockingly high quality, especially in certain niches: https://twitter.com/theshawwn/status/1270624312769130498 https://twitter.com/theshawwn/status/1270624312769130498 I don't think we could've paid human labelers to create tags that thorough or accurate. All the tags for all experiments can be grabbed via https://www.tagpls.com/tags.json https://www.tagpls.com/tags.json, so over time we hope the site will become more and more valuable to the ML community. tagpls went from 50 users to 2,096 in the past three weeks. The database size also went from 200KB a few weeks ago to 1MB a week ago and 2MB today. I don't know why it's becoming popular, but it seems to be.
- rocauc 6y agoI remember following this as it came out (and learning windshield wipers should be called "swipey bois") Surprised and happy to hear you're seeing high labeling quality. We'll re-host with credit on https://public.roboflow.ai https://public.roboflow.ai What license is this?
- sillysaurusx 6y agoThanks! We've decided to license the data as CC-0. We'll add that to the footer. We don't host any images directly – we merely serve a list of URLs (e.g. https://battle.shawwn.com/tfdne.txt https://battle.shawwn.com/tfdne.txt). But any data served via the API endpoints is CC-0.
- qchris 6y agoIf anyone's interested in the direct GitHub link to the repository: https://github.com/ultralytics/yolov5 https://github.com/ultralytics/yolov5
- travisporter 6y agoHm on this page it has something written in an eastern language under YOLO, https://github.com/ultralytics https://github.com/ultralytics says Madrid, Spain, but then they say "Ultralytics is a U.S.-based particle physics and AI startup"
- hikarudo 6y ago"You only look once" in Chinese.
- ely-s 6y agoThere seems to be an unfair comparison between the various network architectures. The reported speed and accuracy improvements should be taken with a bit of scepticism for two reasons. * This is the first yolo implemented in Pytorch. Pytorch is the fastest ml framework around, so some of YOLOv5's speed improvements may be attributed to the platform it was implemented on rather than actual scientific advances. Previous yolos were implemented using darknet, and EfficientDet is implemented in TensorFlow. It would be necessary to train them all on the same platform for a fair speed comparison. * EfficientDet was trained on the 90-class COCO challenge (1), while YOLOv5 was trained on 80 classes (2). [1] https://github.com/ultralytics/yolov5/blob/master/data/coco.yaml https://github.com/ultralytics/yolov5/blob/master/data/coco.... [2] https://github.com/google/automl/blob/master/efficientdet/inference.py#L42 https://github.com/google/automl/blob/master/efficientdet/in...
- cgarciae 6y agoSide note: I like Pytorch but eager pytorch is not faster the jax.jit or tf.function code
- rocauc 6y agoGreat points, and hoping Glenn releases a paper to complement performance. We are also planning more rigorous benchmarking nonetheless. re: PyTorch being a confounding factor for speed - we recompiled YOLOv4 to PyTorch to achieve 50 FPS. Darknet would likely top out around 10 FPS on the same hardware. EDIT: Alexey, author of YOLOv4, provided benchmarks of YOLOv4 hitting much higher FPS here: https://github.com/AlexeyAB/darknet/issues/5920#issuecomment-642213028 https://github.com/AlexeyAB/darknet/issues/5920#issuecomment...
- DEDLINE 6y agoDoes anyone know of an open-source equivalent to YOLOv5 in the sound recognition / classification domain? Paid?
- fattire 6y agoLike it would identify what you're hearing? "Trumpet!" "Wind whistling through oak leaves!" "Male child!" etc?
- DEDLINE 6y agoUbicoustics [1] would be the closest example to what I am looking for in a FOSS / Commercial offering. Is anyone working on this? [1] https://github.com/FIGLAB/ubicoustics https://github.com/FIGLAB/ubicoustics
- ebg13 6y agoIt looks like this is YOLOv4 implemented on PyTorch, not actually a new YOLO?
- bArray 6y agoYOLO is a neural network, Darknet is the framework. Without both YOLOv4 and "YOLOv5" on the same framework, it makes it near impossible to make any kind of meaningful comparison.
- franciscop 6y agoI am very interested on loading YOLO into a Raspberry Pi + Coral.ai, anyone knows a good tutorial on how to get started? I tried before and with Darknet it was not easy at all, but now with pytorch there seem to be ways of loading that into Coral. I am familiar with Raspberry Pi dev, but not much with ML or TPUs, so I think it'd be mostly a tutorial on bridging the different technologies. (might need to wait a couple of months since this was just released)
- boscon 6y agoLatency is measured for batch=32 and divided by 32? This means that 1 batch will be processed in 500 milliseconds. I have never seen a more fake comparison.
- jcims 6y agoHas anyone (beyond maybe self-driving software) tried using object tagging as a way to start introducing physics into a scene? E.g. human and bicycle have same motion vector, increases likelihood that human is riding bicycle. Bicycle and human have size and weight ranges that could be used to plot trajectory. Bicycles riding in a straight line and trees both provide some cues as to the gravity vector in the scene. Etc. etc. Seems like the camera motion is probably already solved with optical flow/photogrammetry stuff, but you might be able to use that to help scale the scene and start filtering your tagging based on geometric likelihood. The idea of hierarchical reference frames (outlined a bit by Jeff Hawkins here https://www.youtube.com/watch?v=-EVqrDlAqYo&t=3025 https://www.youtube.com/watch?v=-EVqrDlAqYo&t=3025 ) seems pretty compelling to me for contextualizing scenes to gain comprehension. Particularly if you build a graph from those reference frames and situate models tuned to the type of object at the root of each each frame (vertex). You could use that to help each model learn, too. So if a bike model projects a 'riding' edge towards the 'person' model, there wouldn't likely be much learning. e.g. [Person]-(rides)->[Bike] would have likely been encountered already. However if the [Bike] projects the (rides) edge towards the [Capuchin] sitting in the seat, the [Capuchin] model might learn that capuchins can (ride) and furthermore they can (ride) a [Bike].
- craftinator 6y agoI've been wondering these same thoughts for years. I don't do much work in the neural network subfield, but have done a lot with computer vision, and always found myself wanting more robust physical estimation techniques that didn't require external data.
- joshvm 6y agoRGB-D based semantic segmentation is certainly a thing. I'm sure it's also been done with video sequences as well.
- jcims 6y agoYeah I wish the flagship phone manufacturers would put the hardware back into the phone to take 3d photos...even better if you can get point cloud data to go with it. The applications right now are kind of cheesy but they will get better and if the majority of photos taken pivot to including depth information i think it could really drive better capabilities from our phones. Eyes are very hard to make and coordinate, yet there are almost no cyclops in nature.
- david_draco 6y ago> In February 2020, PJ Reddie noted he would discontinue research in computer vision. It would be fair to state also why he chose to discontinue developing YOLO, as it is relevant.
- heavyset_go 6y agoI like to think that the name is also a reference to the fact that this will inevitably be used in some autonomous driving systems.
- nharada 6y agoI welcome forward progress in the field, but something about this doesn't sit right with me. The authors have an unpublished/unreviewed set of results and they're already co-opting the YOLO name (without the original author) for it and all of this to promote a company? I guess this was inevitable when there's so much money in ML but it definitely feels against the spirit of the academic research community that they're building upon.
- syntaxing 6y agoTotally agreed, kinda seems dirty to call something "v5" when it this is a derivative work of the original.
- oehtXRwMkIs 6y agoI think derivative is a bit generous. This is just a reimplementation of v4 with a different framework.
- salty_biscuits 6y agoWell, very unlikely to get the original author. He doesn't do that kind of thing anymore https://twitter.com/pjreddie/status/1230524770350817280?s=19 https://twitter.com/pjreddie/status/1230524770350817280?s=19
- sicariusnoctis 6y ago> there's so much money in ML What do you mean? I thought the DL hypetrain was dying as companies failed to make returns on their investments.
- bArray 6y agoI'm just going to call this out as bullshit. This isn't YOLOv5. I doubt they even did a proper comparison between their model and YOLOv4. Someone asked it to not be called YOLOv5 and their response was just awful [1]. They also blew off a request to publish a blog/paper detailing the network [2]. I filed a ticket to get to the bottom of this with the creators of YOLOv4: https://github.com/AlexeyAB/darknet/issues/5920 https://github.com/AlexeyAB/darknet/issues/5920 [1] https://github.com/ultralytics/yolov5/issues/2 https://github.com/ultralytics/yolov5/issues/2 [2] https://github.com/ultralytics/yolov5/issues/4 https://github.com/ultralytics/yolov5/issues/4
- rcpt 6y agoI love that the response to them is "you can you up,no can no bb" Learned a new phrase today.
- catalogia 6y agoCan you explain it? I can't figure out what that means.
- FriendlyNormie 6y agoIt means teaching pajeets how to use computers was a mistake.
- arctangent 6y agoApparently it is Chinese internet slang meaning: "If you can do it, then you go and do it. If you can’t do it, then don’t criticise others." via: http://www.chinesetimeschool.com/zh-cn/articles/chinese-internet-slang-u-can-u-up/ http://www.chinesetimeschool.com/zh-cn/articles/chinese-inte...
- kbenson 6y agoJust found these.[1][2] That is pretty awful, if it's from a dev. Edit: Although as yeldarb explains in a comment here[3], it's probably a bit more complicated than that. 1: https://www.urbandictionary.com/define.php?term=you%20can%20you%20up https://www.urbandictionary.com/define.php?term=you%20can%20... 2: https://www.quora.com/Whats-the-meaning-of-you-can-you-up-no-can-no-bibi https://www.quora.com/Whats-the-meaning-of-you-can-you-up-no... 3: https://news.ycombinator.com/item?id=23478983 https://news.ycombinator.com/item?id=23478983
- ma2rten 6y agoIn February 2020, PJ Reddie noted he would discontinue research in computer vision. He actually stopped working on it because of ethical concerns. I'm inspired that he made this principled choice despite being quite successful in this field. https://syncedreview.com/2020/02/24/yolo-creator-says-he-stopped-cv-research-due-to-ethical-concerns/ https://syncedreview.com/2020/02/24/yolo-creator-says-he-sto...
- jjcon 6y agoIn other words he stopped working on a project and needed an excuse to virtue signal.
- kuzee 6y agoJust read this. Nice overview of the history of the "YOLO" family, and summary of what YOLOv5 is/does.
- tapatio 6y agoLess weights, more accuracy. Magic :)
- heisenburgzero 6y agoThis is not the first time something is fishy. Back in the early stages of the repo. They were advertising on the front page that they are achieving similar MAP to the original C++ version. But only to be found out they haven't train it on COCO dataset and test it.
- darknet-rider 6y agoI really like the work done by AlexAB on darknet YOLOv4 and the original author Joseph Radmon with YOLOv3. These guys need a lot more respect than any other version of YOLO.