25 ms·
Designing a neural network is a thousand times harder than I imagined. After AlphaGo, I tasked myself with creating a neural network that would use Q-Learning
by 2bitencryption 10y ago
Designing a neural network is a thousand times harder than I imagined.
After AlphaGo, I tasked myself with creating a neural network that would use Q-Learning to play Reversi (aka Othello).
At that point, I had already utilized Q-Learning (the tabular version, not using a neural network) for some very simple and mostly proof-of-concept projects, so I understood how it worked. I read up only perceptrons, relu, the benefits/disadvantages of having more/fewer layers, etc.
Then I actually started on the project, thinking "I know about Q-Learning, I know about neural networks, now I just need to use Keras and I'll have a network ready to learn in about twenty lines of python."
Boy was that naive. Regardless of how much you understand the CONCEPTS of neural networks, actually putting together an effective one that matches the problem state perfectly is so, so difficult (especially if there are no examples to build off of). How many layers? Dropout or no, and if so, how much? Do you flatten this layer, do you use relu, do you need a SECOND neural network to approximate one part of the q-function and another to approximate a different part?
I spent MONTHS messing with the hyperparameters, and got nowhere because I'm doing this on a desktop pc without CUDA, so it takes days to train a new configuration only to find out it hardly "learned" anything.
At one point after days of training, my agent actually had a 90% LOSE rate against an opponent that played totally randomly. To this day I am baffled by this.
I went into the project thinking "I have this working with a table, the q-learning part is in place -- just need to drop in a neural net in place of the table and I'm good to go!" It's been almost a year and I still haven't figured this thing out.
- general_ai 10y agoDoing anything large on a machine without CUDA is a fool's errand these days. Get a GTX1080 or if you're not budget constrained, get a Pascal-based Titan. I work in this field, and I would not be able to do my job without GPUs -- as simple as that. You get 5-10x speedup right off the bat, sometimes more. A very good return on $600, if you ask me.
- kuschku 10y agoFor a student doing this in their free time, 600$ can be a huge sum, usually two months rent. That’s not easily paid.
- general_ai 10y agoSpending 10 months instead of 1 or 2 and not getting anywhere is also not free.
- kuschku 10y agoThe problem is that as a student I just can not, in any way, get the money for a 1080. The choice is spending 10 months, or not even starting.
- general_ai 10y agoSounds like someone is really good at coming up with excuses.
- grzm 10y agoPlease don't snipe at people. If you don't have something constructive and civil to say, please just don't comment.
- general_ai 10y agoI didn't snipe. This is life advice. Get a part time job. If you're in North America (Canada or US) and you can't squeeze out $400-600 from your budget over the period of a year, you're making excuses, pure and simple. Cook at home, drop cable subscription, don't go to Starbucks, do part time jobs, and so on and so forth. That's what students did back when I was one.
- kuschku 10y agoI’m in Germany, have a part time job, cook at home, have no cable subscription, don’t go to starbucks, buy food if possible at ALDI. Rent has been going up every year, but wages haven’t, so by now for me, rent for a small apartment is over 60% of my monthly income. The world has changed quite a bit since you were a student.
- solomatov 10y agoAnd why Pascal based Titan? Is it the best investment in terms of performance per $ spent? Also, how cost effective is it to use cloud GPUs for real world machine learning?
- general_ai 10y agoCloud GPUs are not economical if you use them 24x7x365 (which for any serious deep learning researcher or engineer is usually the case). The only scenario I can think of in which they'd be more economical than something under your desk is when you need to run a massive and embarrassingly parallel workload. I.e. try training dozens of models at the same time with different hyperparameters, and run that for a few days. You could do it cheaper, but it would take a long time and it would be a massive pain in the ass, so you pay the pretty penny and get it done in a week. For my needs I have a machine with a 2011-v3 socket, and four GTX1080 GPUs. Warms up my man cave pretty nicely in winter. I also have access to about a hundred GPUs (older Titans, Teslas, newer 1080s and Pascal Titans) at work that I share with others. Now, regarding Titan. Titan is actually not that much faster than GTX1080, so in terms of raw speed there's no reason to pay twice as much. BUT, it has 4GB more RAM, which lets you run larger models. NVIDIA rightly decided that for a $100+/hr deep learning researcher $600 is not going to be that big of a deal, and priced the card accordingly. If your models fit into 8GB, you'll be better off buying two 1080's instead. As to me, I'm thinking of replacing at least one of my 1080s with a Titan, to be able to train larger models. On a purely TFLOP or even TFLOP/watt basis, it doesn't currently make sense to buy anything that doesn't run Pascal.
- modeless 10y agoTitan XP is the maximum single chip performance you can buy right now. $1200 is well worth it if ML is part of your career. The time saved will pay for it.
- nl 10y agoCloud GPUs are cost effective if you need to either fine-tune a pretrained network (eg, use pretrained ResNet/VGG/AlexNet for custom classes, ie[1]) or for inference, or if you don't want the upfront costs. A 4GB GTX1050 is ~$180. A p2 instance on Amazon is $0.9/hour. The cost effectiveness depends on if you have a PC already. [1] https://blog.keras.io/building-powerful-image-classification-models-using-very-little-data.html https://blog.keras.io/building-powerful-image-classification...
- verbify 10y agoIf you're budget constrained, a cheaper card will still get you massive improvements. I'm on a gtx 970, and it far outstrips the cpu. Even a gtx 650 (about £65) should outperform a cpu.
- michaelgrosner2 10y agoOr set up an EC2 GPU unit - spot prices are usually in the sub $.20/hour range.
- reubenmorais 10y agoEC2 GPUs are slower to train than local hardware and more expensive long term. The upside is being able to scale much more easily, but I'd definitely recommend a good consumer grade GPU over EC2 if you're planning on using it for months as opposed to days.
- general_ai 10y agoThey can also be unceremoniously preempted in the middle of your week long training run.
- philipov 10y ago> At one point after days of training, my agent actually had a 90% LOSE rate against an opponent that played totally randomly. Nice! You should have just added a bit at the end to invert whatever answer it got, and you would have had a winner. But more seriously, I think that we will become more clever with designing genetic algorithms to evolve the neural networks as part of the training process rather than trying to build our own from scratch every time. I vaguely recall there is some research being done on that front already.
- deepnotderp 10y agoNeural architecture search with reinforcement learning. We've used an actor critic method internally with good results as well.
- mrjoeblack 10y agoThe fact that a random opponent performs better means that simply inverting the output of a bad strategy (assuming that is even possible, in cases where the output is more complex than binary it should not be) would just give you another bad strategy.
- philipov 10y ago> ...would just give you another bad strategy. That's the joke, but to be fair, inversion is not a binary concept. Negation is binary, but inversion is more general and has the 2D interpretation of reflecting something across X=Y.
- Xcelerate 10y agoYou can always apply for compute time on a large research cluster. It's rare to not be associated with a university or national lab, but I don't think there's anything prohibiting an individual from getting time with a good enough proposal application.
- deepnotderp 10y ago1. For drl,you usually don't want a ResNet actually, since reward assignment is difficult. 2. Almost always,.5 is usually good, but you can tune this. 2. You usually flatten right before the final layers. 3. Yes, relu is preferred. 4. You're referring to double q networks, although this helps, optimality bounds are even better. You'll learn eventually, don't worry. Cheers and welcome to deep learning! ;)
- Aeolos 10y ago> 2. You usually flatten right before the final layers. The newer trend appears to be fully-convolutional networks even for classification, since they appear to overfit less, compared to flattening+dropout.
- deepnotderp 10y agoYup, but he asked me about where to use flatten, not whether to use flatten. But you're right, fully convolutional is the way to go in classification.
- bthornbury 10y agoYou said it. I've been working on some ConvNets for object localization in an image over the past couple weeks and it took days to figure out why my network just seemed to be randomly guessing (50% accuracy). In the end, it was a reduction of the training rate (with SGD) that made things work in what felt like magic. I've started reading the deep learning text book from Ian Goodfellow now (http://www.deeplearningbook.org http://www.deeplearningbook.org). Hoping a solid foundation will build some intuitions for reasoning about these hyper parameters.
- deepnotderp 10y agoHaha, there's no intuition behind these. You should still read the book, but because of other reasons.
- dweekly 10y agoLet me know if you'd like an SSH login on a box with a Titan X Pascal and CUDA stack already installed. Happy to lend a hand and a few teraflops.
- lateguy 10y agoHey, Just curious is this offer for other people also? I am working on this https://github.com/deependersingla/deep_portfolio https://github.com/deependersingla/deep_portfolio (first version open sourced) and before open-sourced this https://github.com/deependersingla/deep_trader https://github.com/deependersingla/deep_trader (255 stars). I was using GTX 980-TI locally but system crashed two weeks ago. I am basically trying to find a RL agent which automatically optimize momentum strategy look_back period and generate portfolio according to optimisation one set on reward function. More can be read on project Readme. Thanks.
- karpathy 10y agoIf it makes you feel any better, I've been doing this for a while and it took me last ~6 weeks to get a from-scratch policy gradients implementation to work 50% of the time on a bunch of RL problems. And I also have a GPU cluster available to me, and a number of friends I get lunch with every day who've been in the area for the last few years. Also, what we know about good CNN design from supervised learning land doesn't seem to apply to reinforcement learning land, because you're mostly bottlenecked by credit assignment / supervision bitrate, not by a lack of a powerful representation. Your ResNets, batchnorms, or very deep networks have no power here. SL wants to work. Even if you screw something up you'll usually get something non-random back. RL must be forced to work. If you screw something up or don't tune something well enough you're exceedingly likely to get a policy that is even worse than random. And even if it's all well tuned you'll get a bad policy 30% of the time, just because. Long story short your failure is more due to the difficulty of deep RL, and much less due to the difficulty of "designing neural networks".
- deepnotderp 10y agoThis 1000x. Drl is a totally different beast than imagenet. Also listen to karpathy, he knows his stuff, rather than random me :D
- option_greek 10y agoPardon my naive question: Is there any point to RL apart from automatically generating labels to a SL network ?
- karpathy 10y agoNot sure if I understand "automatically generating labels to a SL network". I don't believe RL is used in this setting. RL is about learning expected-reward-maximizing policies for environments that you get to interact with. Common benchmarks currently mostly include games (e.g. ATARI, AlphaGo, VizDoom), physics-based animation (e.g. https://www.cs.ubc.ca/~van/papers/2016-TOG-deepRL/index.html https://www.cs.ubc.ca/~van/papers/2016-TOG-deepRL/index.html), or (simulated) robotics-like tasks (e.g. MuJoCo). But the core algorithms (such as policy gradients) can be used more generally in settings that don't necessarily look like environments as usual, but where you want to train a network with stochastic nodes, such as in hard attention, etc. RL is a funny area; A lot of AI researchers get excited about it (mostly motivated by its promise as the formalism that leads to AGI), and yet despite the hype it has so far had very little impact in the industry so far (the Google data center application possibly being an exception, though it was more "RL" than RL, with quotes). It has some promise for Robotics in the real world, but not applied directly and naively. The way that will play out is likely through behavior cloning on human demonstrations or on outputs of trajectory optimizers from simulation, or possibly RL fine-tuning in simulation transferred to real world. But it's still quite early to tell. On this topic, fun story, the most impressive robots I'm aware of right now are from Boston Dynamics and as they mentioned at this year's NIPS they use ZERO machine learning. Forget deep learning or even deep reinforcement learning. Zero Machine Learning. I gave a talk last week about some of our RL experiments @ OpenAI and someone came to me after the talk, described their (straight forward) supervised learning problem and asked me how they can apply RL to it. This, to me, is an alarming sign of damaging hype to the community. You don't use RL for your SL problems. You can if you really want to (e.g. reward = 1.0 if you guess the correct label or -1.0 otherwise), but you really don't want to. You're lucky, use your labels, business as usual.
- malandrew 10y agoFor all of you replying to this comment, what recommendations would you give for someone wanting to learn all this stuff?
- _pastel 10y agoI wonder if some of your difficulties are related to the particular structure of Reversi. The obvious heuristic in the endgame - maximize the number of your pieces - is actually reversed in the opening and midgame, where you want to minimize your number of pieces (other considerations like edge play and parity being equal). So it's possible your NN managed to do worse than random because it learned an endgame heuristic and generalized it improperly. If so, you may want to consider explicitly wiring in the turn counter as an input.
- thomasahle 10y agoDoesn't reversi have a number of stones exactly equal to the turn? It seems strange if a network wouldn't be able to pick that up..
- dlss 10y ago> At one point after days of training, my agent actually had a 90% LOSE rate against an opponent that played totally randomly. To this day I am baffled by this. This is indeed baffling. My best guess: perhaps your error function was reversed? Getting to 90% loss sounds like it would require training.
- Al-Khwarizmi 10y agoAnother post of the "If it makes you feel any better" type: I'm a relatively established researcher in NLP, having worked with a variety of methods from theoretical to empirical, publishing in the top venues with decent frequency, and still I'm having a really hard time to get into the deep learning (DL) stuff. I'm training a sequence-to-sequence model and have been tuning hyperparameters for the last 2-3 months. I'm making progress, but painfully slowly due to the large time it takes to train and test models (I have a local Titan X and some Tesla K80's in a remote cluster, to which I can send models expecting a latency of 3-4 days of queue and a throughput of around 4 models running simultaneously on average - probably more than many people can get, but still feels slow for this purpose) and the fact that hyperparameter optimization seems to be little more that blind guessing with some very rough rules of thumb. The randomness also doesn't help, as running the same model with different random seeds I have noticed that there is huge variance in accuracy. So sometimes I tweak a parameter and get improvements, but who knows if they are significant or just luck with the initialization. I would have to run every experiment with a bunch of seeds to be sure, but that would mean waiting even more for results and my research would be old before I got to state of the art accuracy. Maybe I'm just not good at it and I'm a bit bitter, but my feeling is that this DL revolution is turning research in my area from a battle of brain power and ingenuity to a battle of GPU power and economic means (in fact my brain doesn't work much in this research project, as it spends most of the time waiting for results for some GPU - fortunately I have a lot of other non-DL research to do in parallel so the brain doesn't get bored). In the same line, I can't help but notice that most of the top DL NLP papers come from a very select few institutions with huge resources (even though there are heroic exceptions). This doesn't happen as much with non-DL papers. Good thing that there is still plenty of non-DL research to do, and if DL takes over the whole empirical arena, I'm not bad at theoretical research...
- Aeolos 10y ago> The randomness also doesn't help, as running the same model with different random seeds I have noticed that there is huge variance in accuracy. Set all your random seeds to something predefined, such as 42. Even though the exact randomness is OS-specific, this will at least rule out lucky runs from real hyperparameter improvements.
- raverbashing 10y agoI tried the Cats/Dogs exercise on Kaggle Getting TS to work was hard. Ended up opting for Keras/Theano, TS would be slower for the same problems CPU only Then getting all tensor shapes to "fit together", to run, then to converge, getting the right non-linearity for each layer, then finding out overfitting is a B! with a capital B It is hard
- gambler 10y agoA cynic in me says this is one of the reasons large companies like deep learning so much. Unlike some other branches of AI, deep learning is not something that can be easily used by individuals or smaller companies. It does not scale down very well if at all. So it makes a great basis for some expensive cloud service. They can develop it without fear that someone new will use it to disrupt their business.
- tossaway322 10y agoThis. Why the broadspread interest in a technology so unpredictable, whose payoff is so little and that requires so much investment in hardware/developer time? Wouldn't you be better off learning tools you can understand and that can be used to build reliable and predictable programs? Why not warn graduate students: "You'll be working on this project for about two years. We don't know what to tell you about how to solve the problems involved. We don't have any good general guidelines; this field is changing all the time; nobody knows how these things work except in very broad terms (there is no explanatory power). Sometimes it takes years to find something useful. Luck is the key - if you're unlucky, you're screwed." "Nothing useful will come of your work when you're done. At completion you'll be dumber than when you started, because you will not have learned anything useful except possibly 'patience', a trait not valued in our field and sometimes viewed as equivalent to 'stupidity' or 'stubbornness'. You will, as a result of working with deep learning/NN, forget that sometimes you must cut your losses and quit exploring a particular solution path. At the end, you _may_ get a Master's degree, if you can show you made a good effort or even some progress. Three months after graduation, everything you have done will be obsolete and no firm will hire you: those familiar with NN because your knowledge is obsolete; those not in NN because they see no particular value in your training."
- anigbrowl 10y agoThis reminds me a bit of trying to make 'noodles' on modular synthesizers - a term for a piece that keeps playing and evolving by itself without further tweaking or patching, and preferably without falling into a simple periodic attractor. It's more of an intuitive skill than a science, though I've found Integrated Information Theory useful and relevant.