17 ms·
How Google Translate squeezes deep learning onto a phone
- BigC15 11y agoI'm gonna squeeze some deep learning into your mom.
- cossatot 11y agoInternational travel now has a new source of entertainment: On-the-spot generation of humorous mistranslations.
- cjslep 11y agoJust capture the screenshot and you have a meme generator as well!
- chipgap98 11y agoReddit is going to have a field day
- joosters 11y agoThe oddest result I ever got from WordLens was when using it to translate a page of poetry on a plaque. The output was wonderful :) WordLens was awesome for translating fragments of foreign languages - stuff like signs, menus and so on. But its offline translation seemed to be little more than a word->word translation, so there is a huge scope for improvement there. Very difficult when working offline!
- rasz_pl 11y agoChinese restaurants did it first.
- joosters 11y agoWordLens was an awesome app and it's good to see that Google is continuing the development. The new fad for using the 'deep' learning buzzword annoys me though. It seems so meaningless. What makes one kind of neural net 'deep' and are all the other ones suddenly 'shallow' ?
- StavrosK 11y agoWell, if all it cares about is looks...
- raverbashing 11y ago> What makes one kind of neural net 'deep' and are all the other ones suddenly 'shallow' Number of layers It's that simple
- dnautics 11y agoIt is that simple but the more complex story is that when the number of hidden layers exceeds 2, training becomes difficult. Also convnets for example cheat by having the connections between layers be incomplete bipartite graphs (not every node is connected to every other node), usually chosen because of some physical property - for computer vision nearest neighbors - eg.
- lisivka 11y agoUse another deep learning network to supervise training of your DLN. You can also use it to supervise itself. It is simple idea invented about decade ago (at least I heard it about decade ago here, in Ukraine).
- discardorama 11y agoTo expand on this some more: for a long time, thanks to Cybenko's theorem[1], people just used 1 hidden layer in their neural networks (also because computing was sloowww..). So, your typical NN architecture was input_layer --> hidden_layer --> output_layer. Eventually, people realized that you could improve performance by adding more hidden layers. So while theoretically Cybenko was correct, practically stacking a bunch of hidden layers made more sense. These network architectures with stacks of hidden layers were then labelled as "deep" neural networks. [1] https://en.wikipedia.org/wiki/Universal_approximation_theorem https://en.wikipedia.org/wiki/Universal_approximation_theore...
- teraflop 11y ago> What makes one kind of neural net 'deep' and are all the other ones suddenly 'shallow' ? If this is a serious question, then googling "what is a deep neural network" would take you to any number of explanations. But to summarize very briefly, it's not a buzzword; it's a technical term referring to a network with multiple nonlinear layers that are chained together in sequence. Deep networks have been talked about for as long as neural networks have been a research subject, but it's only in the last few years that the mathematical techniques and computational power have been available to do really interesting things with them. The "fad" (as you call it) is not mainly because the word "deep" sounds cool, but because companies like Google have been seeing breakthrough results that are being used in production as we speak. For example: http://papers.nips.cc/paper/4687-large-scale-distributed-deep-networks.pdf http://papers.nips.cc/paper/4687-large-scale-distributed-dee... http://static.googleusercontent.com/media/research.google.com/en//pubs/archive/43793.pdf http://static.googleusercontent.com/media/research.google.co... http://static.googleusercontent.com/media/research.google.com/en//pubs/archive/42538.pdf http://static.googleusercontent.com/media/research.google.co...
- teraflop 11y agoA possibly relevant research paper that they didn't mention: "Distilling the Knowledge in a Neural Network" http://arxiv.org/abs/1503.02531 http://arxiv.org/abs/1503.02531
- api 11y ago"Squeezes" is very relative. These phones are equal to or larger than most desktops 10-15 years ago, back when I was doing AI research with evolutionary computing and genetic algorithms. We did some pretty mean stuff on those machines, and now we have them in our pockets.
- afsina 11y agoThe main issue here is probably not squeezing memory but squeezing performance. Even using regular SIMD is not good enough if your network is medium sized. They apply linear quantization, lookups and special SIMD operations to make it speedy. See here for what they did for offline speech recognition: http://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41176.pdf http://static.googleusercontent.com/media/research.google.co...
- zippzom 11y agoWhat are the advantages of using a neural network over generating classification trees or using other machine learning methods? I'm not too familiar with how neural nets work, but it seems like they require more creator input than other methods, which could be good or bad I suppose.
- boomzilla 11y agoNeural networks, and the plain old trusted logistic regression :) handles raw, continuous data better than the other learning algorithms. For example, if your inputs are images or audio recordings, it's really hard to do classification with decision trees or random forests as you'd need to construct the features manually. What would be a feature: color densities, color histograms, edges, corners, Haar-like, etc.? The promise of multilayer neural network is that given a lot of data, the right network structures, an appropriate learning strategy, and a huge farm of GPUs, the network can automatically learn the right features from raw data in the first layers, and utilizes the features in later layers. The big advantage of this approach is that you abstract away the domain problems (hopefully), and focus on picking the right network design, the right learning strategy, collecting a good data set etc. Neural network training is also easy to parallelize, so Google and the like can leverage their huge infrastructures. Now if the features in the domain problem is more well defined, like credit ratings, and data is sparse, and domain expertise is available, decision trees are perfectly valid options.
- danieldk 11y agoFor example, if your inputs are images or audio recordings, Just wanted to add: and word/character/phrase embeddings.
- josu 11y agoWordLens/Google Translate is the most futuristic thing that my phone is able to do. It's specially useful in countries that don't use the latin alphabet.
- motoboi 11y agoI am 15 years into this computers thing and this blog post made me feel like "those guys are doing black magic". Neural networks and deep learning are truly awesome technologies.
- deleted 11y ago[deleted]
- dr_zoidberg 11y agoThey are, but once you start learning about them, you realize the "black magic" part comes mostly from their mathematical nature and very little from them being "inteligent computers". A neural net is a graph, in which a subset of nodes are "inputs" (that's where the net gets information), some are outputs, and there are other nodes which are called "hidden neurons". The nodes are interconnected between each other in a fashion, which is called the "topology" or sometimes "architecture" of the net. For example I-H-O is a tipical feed forward net, in which I (inputs) is the input layer, H is the hidden layer and O the output layer. All the hidden neurons connect with all the input neurons "output", and all the output neurons connect to the hidden neurons "output". The connections are called "weights", and the training adjusts the weights of all the neuron with lots of cases until the desired output is achieved. There are also algorithms and criteria to stop before the net "learns too much" and looses the ability to generalize (this is called overfitting). In particular, a net with one hidden layer and one output layer is a universal function estimator -- that is, an estimator that can model any mathematical function of the form f(x1, x2, x3, ..., xn) = y. Deep learning means you're using a feedforward net with lots of hidden layers (I think it's usually between 5 to 15 now), which apply convolution operators (hence the "convolutional" in the name), and lots of neurons (in the order of thousands). All this was nearly impossible until the GPGPUs came along, because of the time it took to train a modest network (minutes to hours for a net with a between 50 to 150 neurons in one hidden layer). This is a very shortened explanation -- if you want to read more I recommend this link[1] which gives some simple Python code to illustrate and implement the innards of a basic neural network and you can learn from the inside. Once you get that you should move to more mature implementations, like Theano or Torch to get the full potential of neutral nets without worrying about implementation. [1] http://iamtrask.github.io/2015/07/12/basic-python-network/ http://iamtrask.github.io/2015/07/12/basic-python-network/
- dharma1 11y agoWould be great to see a more in depth article about this, and maybe even some open source code?
- sarwechshar 11y agoI would be interested in this as well. So far I found a similar app called Mitzuli which is based on open source tools: http://www.mitzuli.com/en/ http://www.mitzuli.com/en/
- murbard2 11y agoI see no mention of it, but I'd be surprised if they didn't use some form of knowledge distilling [1] (which Hinton came up with, so really no excuse), to condense a large neural network into a much smaller one. [1] http://arxiv.org/abs/1503.02531 http://arxiv.org/abs/1503.02531
- poslathian 11y agoThe article mentions algorithmically generating the training set. See here for some earlier research in this area: http://bheisele.com/heisele_research.html#3D_models http://bheisele.com/heisele_research.html#3D_models
- sjwscansodoff 11y agoAll that power and capability, and the fuckers are just using it to sell me ads. Sigh.
- jaredhansen 11y agoReally? You're posting that in a thread about Google Translate? Come on.
- o_____________o 11y agoWhat happened to "don't be evil", jaredhansen?
- anantzoid 11y agoJust waiting for the paper to come out that'll detail all the transformations that were done on the training data specifically for the phone and how did they arrive at deciding to use them. > To achieve real-time, we also heavily optimized and hand-tuned the math operations. That meant using the mobile processor’s SIMD instructions and tuning things like matrix multiplies to fit processing into all levels of cache memory. Let's see how this turns out to be. I'm still skeptical if other apps might crash because of this.
- eosrei 11y agoI used this in Brazil this last March to read menus. It works extremely well. The mistranslations make it even more fun. Much faster than learning Portuguese! I took a few screen shots. Aligning the phone, focus, light, shadows on the small menu font was difficult. You must keep steady. Sadly, I ended up hitting the volume control on this best example. Tasty cockroaches! Ha! http://imgur.com/j9iRaY0 http://imgur.com/j9iRaY0
- shkkmo 11y agoI had some Brazilian roomates who didn't speak english (and I don't speak portugues). We used a combination of my poor spanish and google translate off my phone to comunicate. It worked ok (much better than nothing.) However there were a number of times when there were very large issues in the translations that created some pretty big misunderstandings. Luckily we had a friend who had fluent English and Portuguese who would translate when things got to confused. To reduce errors, you do need to be really careful to use short, complete sentences with simple and correct grammar. It's also better to use and that contain words that aren't ambiguous. (Those two sentences would probably not translate well.) e.g. Please write simple words, short phrases and simple phrases. Please write words with just one meaning. Those phrases and words are easier to translate.
- thaumasiotes 11y ago> Please write words with just one meaning. Those words are very rare and tend to only be useful in very technical contexts.
- shkkmo 11y agoFair enough. The idea that is intended to express is 'unambiguous'. I tend to try to avoid more obscure words when writing text for automatic translation, often at the expense of explicit accuracy.
- raverbashing 11y agoInteresting It seems it can't really handle context, so 'cockroaches' may have been a mistranslation of 'cheap' in some contexts, as the 'it had stopped chestnut' may have simply been 'brazil nuts'
- Uhhrrr 11y agoI don't get it. They say they use a dictionary, and they say it works without an Internet connection. How can both things be true? I'm pretty sure there's not, say, a Quechua dictionary on my phone.
- deleted 11y ago[deleted]
- mattmanser 11y agoIt doesn't come with all the languages, you have to download them.
- ori_b 11y agoAre you sure? On my desktop, the english dictionary is ~1 megabyte uncompressed, and compresses to ~250k with gzip. The download for Google Translate is somewhere around 30 megabytes.
- josu 11y agoYou have to download them beforehand, and offline translating is limited to just a few languages.
- mrigor 11y agoFor those unfamiliar with google's deep learning, this talk covers their recent efforts pretty well https://youtu.be/kO-Iw9xlxy4 https://youtu.be/kO-Iw9xlxy4 (not technical)
- hellrich 11y agoI wonder if they use some kind of (neural) language model for their translations. Using only a dictionary (as in the text) would be about 60 years behind the state of the art...
- up_and_up 11y agoThis technology has been around since 2010 and was developed by Word Lens, which was acquired by google in 2014: https://en.wikipedia.org/wiki/Word_Lens https://en.wikipedia.org/wiki/Word_Lens
- liabru 11y agoThis is great. I particularly like that they also automatically generated dirty versions for their training set, because that's exactly what I ended up doing for my dissertation project (a computer vision system [1] that automatically referees Scrabble boards). I also used dictionary analysis and the classifier's own confusion matrix to boost its accuracy. If you're also interested in real time OCR like this, I did a write up [2] of the approach that worked well for my project. It only needed to recognize Scrabble fonts, but it could be extended to more fonts by using more training examples. [1] http://brm.io/kwyjibo/ http://brm.io/kwyjibo/ [2] http://brm.io/real-time-ocr/ http://brm.io/real-time-ocr/
- JabavuAdams 11y agoGlad there's prior art on that. I had a small project where I iterated all the fonts on the system and used them to generate glyph training images. The next step was to dirty them up, but I never continued the project. More generally, I really like the idea of generating controlled synthetic images and then messing them up for regularization.
- joe_the_user 11y agoIt seems your dissertation paper is behind something password protected [1]. It would be nice to see that too. Can't get [1]https://www.dcs.shef.ac.uk/intranet/teaching/campus/projects/archive/l31011/pdf/LBrummitt_aca08lb_com3021.pdf https://www.dcs.shef.ac.uk/intranet/teaching/campus/projects...
- liabru 11y agoHmm looks like they have, well here's another link to it: https://dl.dropboxusercontent.com/u/1672291/scrabble-referee.pdf https://dl.dropboxusercontent.com/u/1672291/scrabble-referee...
- megalodon 11y agoFunny, just read an article today proposing the same feature detection algorithm (the one you called 'grid merge'). Have you tried applying these techniques on scanned/photographed documents?
- Animats 11y agoWord Lens is impressive. It came from a small startup. Google didn't develop it; it was a product before Google bought it. I saw an early version being shown around TechShop years ago, before Google Glass, even. It was quite fast even then, translating signs and keeping the translation positioned over the sign as the phone was moved in real time. But the initial version was English/Spanish only.
- modfodder 11y agoHere's a short video about Google Translate just released. https://www.youtube.com/watch?v=0zKU7jDA2nc&index=1&list=PLeqAcoTy5741GXa8rccolGQaj_nVGw76g https://www.youtube.com/watch?v=0zKU7jDA2nc&index=1&list=PLe...
- afsina 11y agoThey did this even more impressively when squeezing their speech recognition engine to mobile devices. http://static.googleusercontent.com/media/research.google.com/en//pubs/archive/41176.pdf http://static.googleusercontent.com/media/research.google.co...
- xigency 11y agoGiven the reliability of closed captions on YouTube and the frequency of errors in plaintext Google translate, I wouldn't be surprised if this service fails often, and often when you need it most.
- sytelus 11y agoThe most awesome and surprising thing about this is that the whole thing runs locally on your smartphone! You don't need network connection. All dictionaries, grammar processing, image processing, DNN - the whole stack runs on phone. I used this on my trip to Moscow and it was truely god send because it didn't need expensive international data plans (assuming you have connectivity!). English usage is fairly rare in Russia and it was just fun to learn Russian this way by pointing at interesting things.
- pschanely 11y agoDoesn't this article seem to say that the size of the training set is related to the size of the resulting network? It should be proportional to the number of nodes/layers that the network is configured for, not proportional to the number of training instances. Am I missing something?
- alok-g 11y agoThe network is sized to be able to learn the training data reasonably well (e.g. via hyper-parameter optimization). If there is too much variation in data that is not seen in the real application (like rotation of letters mentioned in the article), an appropriately sized network will still learn it, but would be an overkill for the application at hand.
- tdaltonc 11y agoAnyone want to do a $1 bet on an over/under for how long until word lens can handle Chinese?
- agazso 11y agoThere is an app called Waygo that's already capable of handling Chinese, so I guess it's not too far.
- megalodon 11y agoI generated training sets for an OCR project in JavaScript [1] a while ago using a modified version of a captcha generator [2] (practically the same technique mentioned in this article). [1] https://github.com/mateogianolio/mlp-character-recognition https://github.com/mateogianolio/mlp-character-recognition [2] https://github.com/mateogianolio/mlp-character-recognition/blob/master/captcha.js https://github.com/mateogianolio/mlp-character-recognition/b...
- birdsbolt 11y agoWhy do they need a deep learning model for this? They are obviously targeting signs, product names, menus and similar. Model will obviously fail in translating large texts. Was there any advantage of using a deep learning model instead of something more computationally simple?