11 ms·
Keras vs PyTorch
- Raf_ 8y agoAuthor here - the article compares Keras and PyTorch as the first Deep Learning framework to learn. It explores the differences between the two in terms of ease of use, flexibility, debugging experience, popularity, and performance, among others. If you have experience with learning, or teaching Deep Learning with PyTorch or Keras, we’d love to hear your thoughts about them.
- entropie 8y agoArticle timeouts.
- stared 8y agoIt seems to be a HN hug of death. Though, as I see - it loads, though sometimes with a considerable delays (5-10 sec).
- bitL 8y agoI agree, I also use Keras for stable complex models (up to 1000 layers) in production and PyTorch for fun (DRL). However, if I want to run a distributed training optimization with minimum setup, whether I like it or not, the simplest way is to use TensorFlow's Estimator model and some pre-baked environment like SageMaker. Horovod or CERNDB/Keras require a bit more setup/devops work. The issue with estimators is that once you start using some bleeding-edge things in Keras, it might be very complicated to translate them back to estimators, despite conversion from Keras model to tf.Estimator being trivial.
- jacquesm 8y ago> I also use Keras for stable complex models (up to 1000 layers) That sounds interesting, are you at liberty to say what you are doing?
- bitL 8y agoComplex computer vision classification tasks based on DenseNet/ResNet approaches; those often could be reduced in depth by some Wide ResNet technique. Keras is super easy there and you get a world-class performance after 1 hour of coding and a week of training, when you know what are you doing.
- mlthoughts2018 8y agoI mentioned in another comment [0], but also useful here: most of TensorFlow's tools for distributed model training or multi-gpu training will work out of the box directly on Keras, and distributed training is not at all a reason to directly use TensorFlow over Keras. At worst, you have to add in a tiny bit on TensorFlow code on top of the majority being in Keras, but you would still never need to write a significant amount directly in TensorFlow. I also work on production systems built around deep ResNet architecture for computer vision tasks, and my team does this using solely Keras, including when we do distributed training. Just adding this thought in case anyone mistakenly thinks you have to start out all-in using only TensorFlow because you might expect to need distributed training at some point. [0]: < https://news.ycombinator.com/item?id=17416904 https://news.ycombinator.com/item?id=17416904 >
- inputcoffee 8y agoI appreciate how you focus on just two, which are the state of the art in your opinion. I find it less useful to see comparisons of "top 50 deep learning frameworks for 2018" which include esoteric stuff that is only there for sake of completeness. This way a person branching out from Tensorflow (I assume its Tensorflow) knows which two frameworks to try out, and what to look for.
- probably_wrong 8y agoFor what it's worth, here's my experience: My adviser decided (wisely) that we all needed to learn NN, and we settled on Tensorflow. That went... poorly. I've told this before: the Seq2Seq tutorial was designed for an older version of TF, and it triggered a bug that was not fixed because that way to do Seq2Seq was deprecated and a new tutorial was coming "soon". The "tutorial" was also just a code dump with barely any comments. Eventually we had new people coming in with even less theoretic background than ours (we had read papers for at least 6 months), and that's when we realised it would not work at all. So we organised a 1-week hackathon with Pytorch, and we've been using it ever since.
- Al-Khwarizmi 8y agoSimilar story here. I got bitten by that very seq2seq "tutorial", lost a lot of time with it, and haven't used TensorFlow ever since except for reproducing other people's experiments. It's Keras, Torch, DyNet or PyTorch for me.
- gjmveloso 8y agoWhy not mxnet[1]? Or even better, Gluon[2]? [1]: http://mxnet.incubator.apache.org http://mxnet.incubator.apache.org [2]: https://gluon.mxnet.io https://gluon.mxnet.io
- stared 8y ago(Another author here). It is explicitly explained in the text. :) tl;dr: not nearly as popular (which means: less tutorials, less documentation, less examples, less integration with other systems, less community support for development or discussions) Sure, all frameworks do have some goal and once one is confident in DL, may be a good choice. As you see from the plots there - MXNet is very fast for some applications.
- novaRom 8y agomxnet is not just fast for "some" applications. It consistently outperforms many frameworks especially in realm of compute intensive convolutions.
- stared 8y agoSee charts. Likely, but still 30% speed boost is not a factor for someone learning DL (then debugging, or training wrong models, can easily give overhead of 5-20x).
- t27 8y agoGoogle webcache link if the main page doesnt load (archive.org had some rendering issues) https://webcache.googleusercontent.com/search?q=cache:https://deepsense.ai/keras-or-pytorch/ https://webcache.googleusercontent.com/search?q=cache:https:...
- Raf_ 8y agoThanks, appreciate it!
- gaius 8y agoAs Keras has pluggable back-ends it's the obvious choice (I am using CNTK with mine)
- nightcracker 8y agoI can definitely second Keras - it was the third NN library I tried and the one I actually got working the fastest producing results.
- formalsystem 8y agoThis article echoes my experience as well. I was working on some core NLP models for a larger tech company and wanted to experiment with Keras. I had my models designed within a day and training done within another and had amazing model perf. I was also told that doing it the real way using Tensorflow would be the way to go and I agree with that sentiment if my problem was Google scale which it wasn't. In fact I would argue that most workloads around the world are not Google scale and neither are most Google workloads. This attitude of "real deep learning engineers use Tensorflow" is an unhelpful way of saying "I agree that the API is unreadable but I've invested so much time in the ecosystem that I'll refuse to see its usability problems". Kind of reminds me of assembly programmers that thought C wasn't for l33t 10xx pwner programmers.
- Ar-Curunir 8y ago> Kind of reminds me of assembly programmers that thought C wasn't for l33t 10xx pwner programmers. It's funny because this the same attitude C/C++ programmers have towards developers using other languages now...
- thewizardofaus 8y agoPffft what are programming languages??? If you aren't writing code in straight up binary then you aren't a real h@k3r
- tomrod 8y agoI think there is an emacs command to convert between that and butterfly wing flaps directing cosmic rays to flip bits. (h/t xckd)
- Risord 8y agoIt’s pretty universal thing...
- hknd 8y agoFrom my opinion: Getting started with Tensor Flow, and having a model designed within a day and training within another is also possible. This mostly depends on your model and your data, and (imho) not on the framework of choice. For all, Keras/PyTorch/Tensorflow, you'll need to learn the API - but if you have any ML background, that should be straight forward.
- stared 8y agoBTW (from author - of blog post, and this library): for super-simple live training plots in Jupyter Notebook for Keras (and PyTorch): https://github.com/stared/livelossplot https://github.com/stared/livelossplot For more advanced training for business or Kaggle competitions (version controlling of code and results, advanced charts): https://neptune.ml/ https://neptune.ml/
- maaaats 8y agoI don't do much ML (this kind of ML at least), so I know I'm not the target audience for these libraries. But I'd wish they were written in some language with static typing for IDE help. The API interface/tweaks-to-be-done for some of them is enormous, and mostly undiscoverable. I mean, just looking at the "getting started, 30 seconds to Keras"[0], there are so many magic strings and options. Of course, if one is well-versed in this domain, they make sense. But it's hard to grasp, and Keras is supposed to be the high-level one. [0]: https://keras.io/ https://keras.io/
- gaius 8y agoBut I'd wish they were written in some language with static typing CNTK has a C++ API but the documentation is unfortunately just "read the header file" https://docs.microsoft.com/en-us/cognitive-toolkit/cntk-library-api https://docs.microsoft.com/en-us/cognitive-toolkit/cntk-libr... Also Python obv, or use it as a backend to Keras (in R).
- hak8or 8y agoI totally agree, the lack of documentation via types and lack of smart autocomplete (which I rely on very heavily for API discovery) is not only why I never got into tensor flow, it's also why I never got into languages like python. I even went so far as to use typescript instead of Javascript. I do believe c# has some machine learning libraries, but afaik they aren't anywhere near the level of tensor flow or keras.
- jaegerpicker 8y agoScala and Java both have access to really great ML libraries, MLLib in particular is supposed to be really good. I've used it a little but I'm not really a ML expert to judge.
- eafkuor 8y agoI've been recently following the course over at fast.ai and I'm having the same issue. The API is, as you mentioned, undiscoverable, and I don't really want to look at the source code, which is kind of unreadable anyway ([0]). The course itself is littered with poorly readable code ([1]) such as: to_np(m.ib(V(topMovieIdx))) Why, just why. Despite all this, I wholeheartedly recommend this course, it demystified DL for me. [0]: https://github.com/fastai/fastai/blob/master/fastai/model.py https://github.com/fastai/fastai/blob/master/fastai/model.py [1]: https://github.com/fastai/fastai/blob/master/courses/dl1/lesson5-movielens.ipynb https://github.com/fastai/fastai/blob/master/courses/dl1/les...
- aaronsnoswell 8y agoCompletely agree with the OP. Had the experience of learning TF first, then switching to PyTorch. Cannot recommend enough.
- zerostar07 8y agoWhile we are at it, which framework is the easiest to tweak at the low-level, e.g. create modified LSTMs etc ?
- laingc 8y agoPytorch, by a mile. If I had to summarise the frameworks in a few words, they would be: Keras for speed, Tensorflow for production, Pytorch for research.
- mlthoughts2018 8y agoI would say keras or mxnet for speed and production. PyTorch for research. By this point there are hardly any cases when it’s worth it to descend to lower TensorFlow levels.
- pdyck 8y agoKeras only offers standard layers but you can implement your own LSTM layer and use it with Keras. This way you can take advantage of all the other features
- pdyck 8y agoThat's what I did for my bachelor's thesis. I didn't take any advanced Maths classes, so I had to learn everything from scratch. Keras helped me to build an intuition for neural networks and made me more interested in learning about the formulas and how it works with TensorFlow in the background. I really enjoyed learning with this top-down approach.
- phonebucket 8y agoI don't believe I would ever discourage anyone from using any particular framework. The skills learnt from one are highly transferable, so it doesn't matter too much which framework you start with. Also, with eager execution, Tensorflow has become much more accessible to new users. Having said that, the world would likely be a better place if everyone just used PyTorch :)
- Raf_ 8y agoThese are great points, thanks.
- cttet 8y agoI just like Tensorflow better. For building new models, the graph is complex and errors are unavoidable. There is a separate compile time for Tensorflow and errors will be found before the data come in. Tried pytorch before, the error messages are usually not helpful at all and often leads to clueless debugging for hours. For trying out deep learning, or build on existing models, pytorch or keras may be easier to grasp. But when making new models that involves a lot of math, the Theano/Tensorflow is more helpful IMO.
- v4r 8y agoIf a model involves a lot of math, is it more helpful to be able to debug it? What tf lacks is an intuitive debugging tool. I think this is where pytorch excels.
- cttet 8y agoYou don't often need to debug, especially if the model can be checked on compile time. Think static types for programming. For Tensorflow all the data types and tensor dimensions are checked before loading any data, it the math is derived correctly then it is not necessary to even debug. If data and model are mixed, it often resort to line-by-line debugging to zone out the real problem, which often takes more time.
- stared 8y agoIt is an interesting perspective, but my experience is exactly the opposite. In Theano debugging was awful. TF felt like a breeze until it didn't. When I jumped on PyTorch - it TF started feeling confusing by comparison. Errors exactly in the defective lines, possibility to print everywhere (or using any other kind of feedback / logging intermediate results). For using models it may note matter that much (though, again read YOLO in TF and PyTorch and then decide which is cleaner :)). For new models which go beyond a standard ConvNet/LSTM... well, PyTorch is heaven, Theano sounds like a torture.
- cttet 8y agoI am not sure what is your programming style in Pytorch. As people recommended and in most tutorials I see the sequential approach, where a small mistake of data preprocessing would lead to clueless errors in a completely irrelevant line. YOLO is a quite standard feed-forward model in my opinion. I mean the math part, which I am more concerned with. I have never used Theano before, my idea from it is that Tensorflow followed its static graph approach.
- v_lisivka 8y agoStart with Yolo: small C program with minimum of dependencies.
- novaRom 8y agoI compared many frameworks and find mxnet is the best choice for production if you care about development time and speed of the training.
- k__ 8y agoIs TensorFlow.js a viable alternative? I don't know much about Python :/
- stared 8y agoLearning Python is way faster than learning Deep Learning, so it shouldn't be an issue. I am planning organizing a TensorFlow.js bootcamp, but here it is more difficult (as data preprocessing, and debugging in general, is way more difficult in JS than in Python).
- k__ 8y agoTrue. I just had the hope, that integrating this into existing JS code-bases would be easier with TensorFlow.js :)
- mlthoughts2018 8y agoHaving used both plain TensorFlow and Keras for some very large image processing production services, Keras wins easily, and interoperates with sprinkling in low-level TensorFlow very well. Even defining a custom deep CNN for multiple image prediction tasks (so, deep and custom architecture), Keras holds up well — and creating your own layers in Keras is very easy.
- leecarraher 8y agoFor nn's in my experience out of memory, and preprocessing tends to cause an equal number issues as the nn optimization. Which tfrecords and streaming seem to solve. Are there similar object loading facilities in pytorch? Though I have not specified models in keras, since it is now part of tf i presume the formats are compatible.
- iaml 8y agoFWIW keras is integrated into tensorflow as tf.keras [0], or at least, it should be - never tried it myself. [0] https://www.tensorflow.org/api_docs/python/tf/keras https://www.tensorflow.org/api_docs/python/tf/keras
- mark_l_watson 8y agoNice article, and I agree with the explanations of what makes Keras and TensorFlow best for specific use cases. Some history: I have used TensorFlow for years, switched to coding against the Keras APIs about 8 months ago. I wish I had more experience with PyTorch, but I just have the time right now to do more than just play with it. One suggestion to the authors: the benchmark figures are interesting, but I wish you had shown CPU only results also. At work, I have all the GPU resources I need but for my home projects, which are all NLP deep learning experiments, I usually rent a many core large memory server with no GPUs (GPUs seem to speed up RNNs less than other model types).
- Raf_ 8y agoGlad you like the article and thanks for suggesting researching the CPU usage across these frameworks - it's something worth looking into.
- visarga 8y agoFor most applications you can probably use a TCN (temporal convolutional network) instead of LSTM. TCN's are implemented in all major frameworks and work an order of magnitude faster because they are parallel. https://arxiv.org/abs/1803.01271 https://arxiv.org/abs/1803.01271 https://arxiv.org/abs/1608.08242 https://arxiv.org/abs/1608.08242
- droidist2 8y agoI've been using a QRNN in PyTorch, is this similar to a TCN? https://github.com/salesforce/pytorch-qrnn https://github.com/salesforce/pytorch-qrnn
- visarga 8y agoNo, TCN is similar to WaveNet (dilated convolutions + masking the future + residual connections). It's a plain convnet, not an LSTM with a twist. That's why it runs efficiently in parallel on GPUs, like image processing convnets.
- dekhn 8y agoI don't know about how many people external to Google know about tf.estimator, but it's where most people who aren't building complicated custom architectures should be starting. Keras is nice, it's easy to use, but I wouldn't use it to design build and run a massive productive pipeline. tf.estimator is just that.
- kajecounterhack 8y ago+1 for the estimators API.
- ulucs 8y agoHaving used Torch (the Lua library) before, the comparison between the Sequential models seems very absurd. Even the pyTorch documentation gives an almost equivalent model defintion method: # Example of using Sequential model = nn.Sequential( nn.Conv2d(1,20,5), nn.ReLU(), nn.Conv2d(20,64,5), nn.ReLU() )
- nabla9 8y agoThese sequential models are are like Fibonacci function comparisons between programming languages. They are simple and basic, difference between 5 lines of code or 20 lines of code makes no difference. You spend very little time actually coding these layers. Understanding the model, default parameters used underneath is more important. It would be nice to see some examples with skip-layers, weight sharing etc. You you have to drop sequential model to do them or not?
- nbeleski 8y agoInteresting article. I've been doing some production work with ML and I was wondering which tool would work better in my specific enviroment. Currently I've been training a CNN model in Keras with good success, and using custom scripts to port it to a TensorFlow model. The .h5 file from Keras helps a lot with this step. Next step is compiling a shared Tensorflow library so I can deploy the trained model in C++ (project requirement) and this has been a pain in the ass, regardless of framework...
- sivakon 8y agoPytorch has the best API to understand deep learning and Pytorch based Pyro is also very good for probabilistic programming (fresh take coming from Stan/PyMC3)
- minimaxir 8y agoSpeed is one thing, but the key value proposition of Keras for me that rarely comes up in these comparisons are Keras’s native utility functions, including easy and correct text tokenization/padding, easy OHE of categorical variables without using sklearn, and easy model saving/loading from an .hdf5 file. (Although I am not an expert on PyTorch and not as familiar with the ETL pipeline for that)
- rjdagost 8y agoI agree strongly with the ease of model saving / reloading in Keras. I found this basic functionality to be exasperatingly difficult and cumbersome in TensorFlow.
- m3kw9 8y agoI was able to build a deep learning OCR using CNN from scratch using Keras and runnning in an App using iOS’s coreml in 2 months without prior experience. Hard part was actually getting the data set great. 80% of the time was data massaging. Keras saved me some time. Although the results wouldn’t be world class.
- ChankeyPathak 8y agoWhat about tf.keras? https://www.tensorflow.org/versions/r1.9/programmers_guide/keras https://www.tensorflow.org/versions/r1.9/programmers_guide/k...
- d_burfoot 8y agoCan we please please please not have the kind of framework overproliferation and fragmentation in the Deep ML world that they have in the front-end web world? It's hard enough to learn ML concepts without also having to learn a new ML framework every year.
- mkirklions 8y agoI was going to disagree, but I am using laravel php (technically backend, but Ive developed a front end for testing) and.... I'm also using React Native because I dont like Apple and hopefully I can use a friends computer the moment I compile for iphone.
- jbgordon 8y agoLove Keras. If you like, you can also use the Keras API inside TensorFlow (as tf.keras). We recently published this guide w/ more info - https://www.tensorflow.org/versions/r1.9/programmers_guide/keras https://www.tensorflow.org/versions/r1.9/programmers_guide/k... - and are working on a few more examples for the v1.9 release in a couple weeks. Optionally, you can also use tf.keras in combination w/ eager execution, enabling you to write code like this: https://colab.research.google.com/github/tensorflow/tensorflow/blob/master/tensorflow/contrib/eager/python/examples/nmt_with_attention/nmt_with_attention.ipynb https://colab.research.google.com/github/tensorflow/tensorfl...
- Raf_ 8y agoThanks for sharing these guides. Excited about exploring tf.keras with Eager Execution! :)
- opwieurposiu 8y agoAnyone get pytorch to work on windows? I could not figure it out. Keras and tensorflow work great on windows.
- Raf_ 8y agoWhich version were you installing? Version 0.4 released in April added Windows support.