13 ms·
Building a deeper understanding of images
- mrfusion 12y agoThat's pretty amazing. It seems like we're at a point where we could build really practical robots with this? Robots to do dishes, weed crops, pick fruit? Why isn't this being applied to more tasks?
- sp332 12y agoBecause it was just invented over the past 2 years? And it is being used, but mostly by Google and Facebook.
- drcode 12y agoI would eat a "hat with a wide brim" if Google isn't going to release a robot that can do basic household chores (laundry, dishes, dusting) within the next 3 years. Google has been gobbling up robotics startups, and given how Google also loves gobbling up personal data, having robot "boots on the ground" in every home must be extremely appealing to them.
- GhotiFish 12y ago... Do you want me to note that in my calender?
- HSO 12y agoabsolutely --> https://news.ycombinator.com/item?id=8274902 https://news.ycombinator.com/item?id=8274902
- VikingCoder 12y agoI'll congratulate you on not specifying a time duration for your meal. Indigestible material can cause a blockage in your colon, which can cause severe pain, damage to your colon, and even death. I suggest chopping the hat up into very, very small pieces, and eating them over a very long period, such as a month or longer.
- jschwartzi 12y agoIf you'll excuse me, I have to write some letters to the legal department at the hat factory where I work.
- robotresearcher 12y agoI doubt it. The easy parts of laundry and dishes are already done by simple robots sold in every white goods department. The remaining parts are very demanding indeed. Research labs are not robustly demonstrating these capabilities yet, even with very expensive robots.
- bfung 12y agoI had this thought the other day - not robots washing the laundry, but folding it. Folding is actually not that simple after thinking about it.
- robotresearcher 12y agohttps://www.youtube.com/watch?v=gy5g33S0Gzo https://www.youtube.com/watch?v=gy5g33S0Gzo Work is progressing. That's a $300K robot. (edit: no affiliation. Video shows PR2 robot at Berkeley folding towels competently but very slowly in 2010)
- botman 12y agoAnother robot researcher agrees. Classifying images is totally different from enabling a robot to perform a difficult task. Object recognition is a supervised learning problem, mapping image -> vector of probabilities. Robotic manipulation is a control problem, mapping a long sequence of images -> a long sequence of actions.
- jrussino 12y agoThe idea is right but if by "release" you mean "sell to consumers" I think you're a little over-optimistic about the timeline. Google's driverless car project is a direct descendant of the tech developed for the DARPA Grand/Urban Challenges. Those took place in 2005-2007 [1]. It's been ~9 years since the first Grand Challenge, and while there has been great progress in this area no one will sell me an autonomous car just yet. The first DARPA robotics challenge was held in December of 2013 [2]. If you check out some of the videos from that competition [3] [4] (or some of the PR2 videos another user posted), you'll see that we can make general-purpose humanoids that can do some pretty neat stuff, but there's still a lot of work to be done before I can buy one that will reliably do many different household tasks. The Atlas platform that many of the competitors used in the DRC was developed by Boston Dynamics, who were subsequently bought by Google. So yes, they are definitely in this space. Maybe progress in humanoids will go faster just by virtue of the fact that we've got an extra decade of research to build on, but the journey from "works in a lab/demo, barely/slowly" to a viable commercial product is a long one. [1] http://en.wikipedia.org/wiki/DARPA_Grand_Challenge http://en.wikipedia.org/wiki/DARPA_Grand_Challenge [2] http://en.wikipedia.org/wiki/DARPA_Robotics_Challenge http://en.wikipedia.org/wiki/DARPA_Robotics_Challenge [3] http://www.youtube.com/watch?v=hzmMVHGNXvI http://www.youtube.com/watch?v=hzmMVHGNXvI (sped-up highlights) [4] http://www.youtube.com/watch?v=mwWm3HaDbnQ http://www.youtube.com/watch?v=mwWm3HaDbnQ (full 10-hour recording of competition live-stream)
- cscurmudgeon 12y agoJust a small list of other things you need: 1. Natural language understanding 2. 3D object understanding 3. Planning 4. Object manipulation 5. Navigation 6. Speech recognition
- mrfusion 12y agoReally, just for pulling weeds? Maybe you thought I meant a robot that can do arbitrary household chores. I actually meant specialized robots for specific tasks.
- icegreentea 12y agoYou skip 1 and 6 then. 3 and 5 isn't bad right now. 4 might be tricky (dependent on task at hand - weed pulling can be surprisingly difficult depending on context). 2 would likely be difficult. The results for the actual competition are here. You can take a quick skip at the error rates of the first place approaches. For some tasks they're down to 5-10% which might be acceptable for some tasks, but could also be completely unacceptable for others. shrug The day is probably coming, but it probably won't be soon. Not the least because robots are actually ridiculous to work with / develop. http://www.image-net.org/challenges/LSVRC/2014/results http://www.image-net.org/challenges/LSVRC/2014/results
- tachyonbeam 12y agoYou can realistically skip 1 and 6, control it with your smartphone instead. Or, you can simplify the problem by using more constrained verbal commands. I was told that speech recognition works very well when you have a simple grammar instead of an open-ended language.
- deleted 12y ago[deleted]
- michaelbuckbee 12y agoWe don't really classify it this way, but Google's self driving car is the practical robotic implementation of this. There has been a bunch of press around how the Google self-driving car has to have streets mapped out for it, etc. This research will go directly to issues like: "Is that a 'domestic cat' or 'paper bag' in the street up ahead?"
- jschwartzi 12y agoWhat if it's a domestic cat in a paper bag?
- robotresearcher 12y agoRobots are being applied as fast as (i) they really work, and (ii) they make economic sense. Precise, powerful mechanical actuation is really expensive. High quality motors use expensive rare-earth magnets and require very accurate manufacturing. Right now, only very valuable manipulation tasks can justify using a robot. Hence, robots built most of my car, but I take out my own trash. Non-physical AI stuff will catch on much faster than robots. edit: The exception is clever things like Roomba, where the mechanical parts are dirt cheap, and the robot's behaviour compensates for its lack of precision. Neat.
- Houshalter 12y agoI wonder if using imprecise motors would be acceptable with sufficient AI. So it could account for errors and imprecision in it's movement.
- robotresearcher 12y agoExample project partly motivated by that idea: http://cswww.essex.ac.uk/staff/owen/machine/cronos.html http://cswww.essex.ac.uk/staff/owen/machine/cronos.html (no affiliation)
- krasin 12y agoI only partially agree. Just recently, pretty decent stepper motors have become very cheap (~$6 for NEMA-17), thx to all this 3d printer boom, and lots of things are now finally possible to implement on a budget. But I agree that good servos with high torque are still super-expensive.
- robotresearcher 12y agoSome manipulation tasks require force control, too.
- krasin 12y agoSurprisingly, this is also getting solved. For example, there're demos of self-calibrating cheap Delta robots which use force sensors on the plate: https://groups.google.com/forum/#!topic/deltabot/6fxnM20nYKc https://groups.google.com/forum/#!topic/deltabot/6fxnM20nYKc Sure, it's not yet there, but with the space of feasible is growing fast.
- karpathy 12y agoI think one of the main reasons is that the improvements have been drastically significant and very recent; There hasn't been enough time to convert the research code into open source code and libraries, but you can confidently expect these models to become pervasive over the next few years not just in robotics, but in all perception systems. A few good open source libraries out there that have the building blocks for putting together similar models: - Caffe (C++) http://caffe.berkeleyvision.org/ http://caffe.berkeleyvision.org/ - cudaconvnet2 (C++) https://code.google.com/p/cuda-convnet2/ https://code.google.com/p/cuda-convnet2/ - DeepBeliefSDK (iOS,Android,OS X, Raspberry Pi, JS) https://github.com/jetpacapp/DeepBeliefSDK https://github.com/jetpacapp/DeepBeliefSDK
- mdda 12y agokarpathy is being too modest here. He's also created his own convnet.js [1], which you can play with online [2] (complete with relevant, working demos), etc. [1] https://github.com/karpathy/convnetjs https://github.com/karpathy/convnetjs [2] http://cs.stanford.edu/people/karpathy/convnetjs/ http://cs.stanford.edu/people/karpathy/convnetjs/
- imaginenore 12y agoWe have robots that pick fruits: https://youtube.com/watch?v=RCBQqEGp8Go https://youtube.com/watch?v=RCBQqEGp8Go https://youtube.com/watch?v=fUGVBTxheHo https://youtube.com/watch?v=fUGVBTxheHo
- colanderman 12y agoNow if only Google could develop a way to serve static text content without using JavaScript! (All I get is a B with twirling gears in it...)
- gavinpc 12y agoThis is a longstanding Blogger bug that happens when cookies are blocked. They haven't fixed it because you and I are the only people on the planet who whitelist cookies.
- karpathy 12y agoI am one of the people who helped analyze the results of the mentioned ILSVRC challenge. In particular, I performed an experiment comparing Google's performance to that of a human a week ago and wrote up the results in this blog post: http://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/ http://karpathy.github.io/2014/09/02/what-i-learned-from-com... TLDR is that it's very exciting that the models are starting to perform on par with humans (on ILSVRC classification at least), and doing so on orders of milliseconds. The included page also has a link to our annotation interface where you can try to compete against their model yourself, and see its predictions and mistakes.
- mrfusion 12y agoThanks for being here. Why do you think the results improved so much this year?
- karpathy 12y agoGoogle hasn't released details about their model yet, but the VGG team in close 2nd place had a very simple and beautiful "vanilla" CovnNet model, but with more careful hyperparameter settings and training/testing protocol. In other words, the source of improvement are tweaks on the original Krizhevsky architecture from 2012, not completely new and unexpected ideas. That is not to belittle the contribution - these experiments take forever to run and require very good practices and intuitions for what to try next. Karen released the details of the model only yesterday on arXiv (http://arxiv.org/pdf/1409.1556v1.pdf http://arxiv.org/pdf/1409.1556v1.pdf) More specifically, as can be seen in the paper, it seems that very deep stacks of conv/conv/pool modules with tiny 3x3 filters work well (which is satisfying because it's really simple and beautiful), and from being more thorough with the training and testing protocol (data augmentations, averaging, multiscale approaches in both train/test time, etc). Google will release details of their method on Sept 12 so we'll know more. From their abstract, it seems they have a more significant departure from a basic convnet architecture.
- mentat 12y agoSo what the process for handling misclassification? http://cs.stanford.edu/people/karpathy/ilsvrc/val/ILSVRC2012_val_00042717.JPEG http://cs.stanford.edu/people/karpathy/ilsvrc/val/ILSVRC2012... is definitely a grand piano not an upright which is the "official" answer. Edit: http://cs.stanford.edu/people/karpathy/ilsvrc/val/ILSVRC2012_val_00042797.JPEG http://cs.stanford.edu/people/karpathy/ilsvrc/val/ILSVRC2012... is also misclassified. There isn't a bee house there. Closest thing is a barbecue. Edit 2: http://cs.stanford.edu/people/karpathy/ilsvrc/val/ILSVRC2012_val_00042836.JPEG http://cs.stanford.edu/people/karpathy/ilsvrc/val/ILSVRC2012... appears to be a building based on scaling clues.
- botman 12y ago"typical incarnations of which consist of over 100 layers with a maximum depth of over 20 parameter layers)" Anyone know exactly what that means? I'm guessing that that there are 100 layers total, 20 of which have tunable parameters, and the other 80 of which don't--e.g., max pooling and normalization.
- hyperion2010 12y agoI wonder whether some of the intermediate layers in these models might correspond to something like "living room" or other locations that provide additional information about the objects that might be in the scene. For example, I suspect it was much easier for me to identify the preamp and the wii in one of the pictures because I knew it was a living room/den instead of an office or study.
- pdenya 12y agoIn this case, no they didn't have pre-labeled location/setting available. You can see one of the datasets they used here: http://image-net.org/challenges/LSVRC/2013/#data http://image-net.org/challenges/LSVRC/2013/#data. Generally speaking, Neural Networks are black boxes. The layers interact with each other but not in a defined categorical manner like that. Layer size/depth are parameters you provide when setting up that have tradeoffs in result accuracy, space, time spent, etc like jpeg quality.
- joelthelion 12y agoHow big is the model? Training these kinds of networks is expert work and requires enormous infrastructure; but if they released the model, I'm sure people like us could come up with all sorts of very useful applications.
- gcr 12y agoIf you're interested, Caffe http://caffe.berkeleyvision.org/ http://caffe.berkeleyvision.org/ comes with some pre-trained models for ImageNet, which was close to state-of-the-art a year or two ago.
- Someone1234 12y agoI wish this was available as a translation app. You point your phone at a fruit stand and it names every single item, and you can then ask the vendor for the item by name. It isn't that crazy, in fact that's exactly what they have right now but just in English only.
- Igglyboo 12y agoNot exactly what your'e talking about but you should take a look at Word Lens(which google just bought). Basically you hold your phone up and position it over a piece of text using the camera. It then OCRs the text, translates it, and replaces it in realtime in the camera feed. It's pretty remarkable.
- MichaelAza 12y agoThese classifications are amazing but the fact that the first image in the article is classified as "a dog wearing a wide-brimmed hat" and not as "a chihuahua wearing a sombrero" is telling of how far we are from true understanding of images. Only a human possessed with the relevant cultural stereotypes (chihuahua implies Mexican, ergo, the hat must be a sombrero) could make that conclusion. Even so, I firmly believe that at this rate of improvement, we're not far from that kind of deep understanding.
- wodenokoto 12y agoI'm not sure how your example is more true.