5 ms·
Thanks for announcing this, looks amazing! As somebody that's been toying around with deep learning and machine learning, I've been wondering what the steps are
by ddlutz 10y ago
Thanks for announcing this, looks amazing! As somebody that's been toying around with deep learning and machine learning, I've been wondering what the steps are to move from 'cool example' to viable product. I know somebody else mentioned something general like that in another comment, but I had a concrete example.
For instance, it's extremely easy to set up an MNIST clone and achieve almost world-record performance for single character recognition with a simple CNN. But how do you expand that to a real example, for instance to do license plate OCR? Or receipt OCR? Do you have to do two models, 1 to perform localization (detecting license plates or individual items in a receipt) and then a second model which can perform OCR from the regions detected from the first model? Or are these usually done with a single model that can do it all?
I'm not sure if answering these questions is a goal of your course, or if they're perhaps naive questions to begin with.
- jph00 10y agoThey are excellent questions and that indeed is exactly the goal of this course. I hope you try it out and let us know if you find the answers you need. For this particular question, a model that does localization and then integrated classification is called an "attentional model". It's an area of much active research. If your images aren't too big, or the thing you're looking for isn't too small in the image, you probably won't need to worry about it. And if you do need to worry about it, then it can be done very easily - lesson 7 shows two methods to do localization, and you can just do a 2nd pass on the cropped images manually. For a great step by step explanation, see the winners of the Kaggle Right Whale competition: http://blog.kaggle.com/2016/01/29/noaa-right-whale-recognition-winners-interview-1st-place-deepsense-io/ http://blog.kaggle.com/2016/01/29/noaa-right-whale-recogniti... (There are more sophisticated integrated techniques, such as the one used by Google for viewing street view house numbers. But you should only consider that if you've tried the simple approaches and found they don't work.)
- Eridrus 10y agoThe big struggle with deep models is their thirst for data. MNIST is considered a simple toy example, and it has 50k images spread across 10 classes. ImageNet has 1m images spread across 1k samples. One of the things that has made image recognition in the form of categorisation easier is that using a network pre-trained on ImageNet, and then finetuning it to your task actually works pretty well and requires far fewer images. The struggle with doing something like license plate OCR is that it's unlikely that doing that you can transfer the learning from ImageNet to your target task. So, in reality your struggle is going to be more around the data than the model. If you already had a system deployed that was getting data in and you were getting some feedback of when your model failed, then this problem would be easily solved, but if you're building from scratch this is going to be your biggest problem. And since you don't necessarily know ahead of time how easy or hard your problem is, you don't know how many samples you will need or how much it will cost you. So, if you did actually want to build a license plate reader using deep learning, my suggestion would be to try and artificially create a dataset by generating images that look like license plates and sticking them in photos in the state you expect to see them in (i.e. blurred, at weird angles, etc) and then training a neural net to recognise them. That would give you a sense for how hard the problem is, and how much data you will need to collect. In terms of the model; I would probably just try having 6 outputs with 36 classes per output corresponding to the characters/digits in order. I don't know if it will work well, but it's a good baseline to start with before trying more complicated things like attention models or sequence decoders (https://github.com/farizrahman4u/seq2seq https://github.com/farizrahman4u/seq2seq )