10 ms·
Unreasonable Ineffectiveness of Machine Learning in Computer Systems Research
- mcguire 9y agoBranch predictions are an interesting use, although I'm wondering how expensive a misprediction really is. But this: "Another example is the use of regression techniques from machine learning to build models of program behavior. If I replace 64-bit arithmetic with 32-bit arithmetic in a program, how much does it change the output of the program and how much does it reduce energy consumption? For many programs, it is not feasible to build analytical models to answer these kinds of questions (among other things, the answers are usually dependent on the input values), but if you have a lot of training data, Xin Sui has shown that you can often use regression to build (non-linear) proxy functions to answer these kinds of questions fairly accurately." I'm not sure whether I am fascinated or horrified.
- pm215 9y agoWikipedia reckons 10-20 cycle penalty for a branch mispredict (which sounds plausible given it fills the pipeline with useless junk). Given that branches are quite common, that's painful enough to want to avoid, but definitely not so painful that you'd want to devote as much silicon to solving it as you do to, say, L1 cache. I do recall a bit of research (published by a Nokia R&D team I think) that reckoned you could get a mostly-ok performance estimate by tracking about half a dozen indicators including instructions executed, cache misses, tlb misses and brancb mispredicts and weighting them appropriately. The trouble there is nobody wants a performance model that's right 90% of the time but significantly wrong 10% of the time with no way to tell if the workload you want to test is in the 10%. But it's an indication of the importance of branch prediction still, I think.
- 60654 9y ago> Driverless long-haul trucks are apparently just a few years away, and the main worry now is not so much the safety of these trucks but the specter of unemployment facing millions of people currently employed as truck drivers. No, no they're not. We have some lane tracking in good weather etc., but we are still decades (or more) away from full level-5 autonomy that would make drivers behind the wheel unnecessary. But it only goes to show that not even computer science experts are immune to marketing hype and well funded PR campaigns. :) As for the unreasonable ineffectiveness - it's not just in systems research. ML can be very effective in some areas (especially when there is a ton of training data), but many areas of human endeavor are hard to model via function approximation techniques like those used in most of ML.
- therpe1 9y agoWe are certainly not decades away. It seems to be a classic human error: people always seem to overestimate how far we've come and how far we've to go... and also underestimate the rate of progress/change. That's why technologies like the iPhone seem to show up "out of the blue" and "change everything." That said, people also seem to never take into account the politics. Even if the technology was ready tomorrow it's not clear that certain special interests and elected officials would allow driverless 40-ton trucks to roam free on the highways. There would probably be a lot of push-back -- for the kids of course. Transportation is the definition of a highly-regulated industry and it's not clear that technologists will be able to move in and disrupt it so easily.
- seanmcdirmid 9y ago> but we are still decades (or more) away from full level-5 autonomy that would make drivers behind the wheel unnecessary. Do you have a citation for that? I often see this asserted, and I see many assertions the other way. Whenever I personally see self-driving cars in the wild, I'm astounded at their effectiveness on uncontrolled access roads. What makes you think controlled access roads will be much more difficult? Or to be more specific, why is industry consensus that we are decades (or more) away from this goal? Or is this just your personal opinion?
- xfs 9y ago> none of us in the automobile or IT industries are close to achieving true Level 5 autonomy - Gill Pratt, Toyota Research Institute http://spectrum.ieee.org/cars-that-think/transportation/self-driving/toyota-gill-pratt-on-the-reality-of-full-autonomy http://spectrum.ieee.org/cars-that-think/transportation/self... > It will be 25 years before self-driving cars take off in America - Bill Gurley, Uber investor http://www.cnbc.com/2017/04/06/bill-gurley-uber-investor-self-driving-cars-25-years-away-in-us.html http://www.cnbc.com/2017/04/06/bill-gurley-uber-investor-sel...
- sroussey 9y agoAt the point of level 4, the roads and conditions will change to better accommodate, likely pushing cities and villages farther apart (particularly in America).
- andreasvc 9y agoCute title but the post didn't really address either reasonableness or effectiveness, but mostly claimed that the potential has not yet been realized. It's a pet peeve of mine to see these hackneyed joke titles referencing famous papers, "considered harmful" is another case in point. Let's just stick to descriptive titles.
- visarga 9y agoIt imitates an old paper title: "The Unreasonable Effectiveness of Mathematics in the Natural Sciences" by Eugene Wigner (1960)
- deleted 9y ago[deleted]
- backpropaganda 9y agoI thought Wigner was actually paying homage to Karpathy.
- visarga 9y agoFew know that Karpathy was using a previous meme.
- gipp 9y agoYes, that's exactly what the parent was referencing -- bloggers in the CS space looove referencing that one and the "Goto Considered Harmful" paper, and it's been really, really annoying and not clever at all for years.
- coliveira 9y agoIt is a modern disease of CS and related areas. They think it is better to be "cute" than descriptive, as a way to attract attention.
- gumby 9y agoIt's an element of all disciplines and, more charitably, one aimed at achieving two functions. One is to say "this is a paper from someone embedded in the discipline and who speaks the same vocabulary as you". The other is to say "and this is intended to expound on a subject parallel to X (or following on from X, as the case may be)" You can argue that this class of jargon is exclusionary or not. Agre, at UCLA, for example took the position that jargon was inherently a tool of dividing up groups into "in and out" (ironically he himself, with an MIT PhD, was very much in the in group). I tend to consider jargon just a tool like any other, typically valuable because you can save a lot of time and gain clarity by saying "O(n^2)" or "trie" and assume your reader understands it.
- elvinyung 9y agoTo contrast this opinion, a promising research project from CMU that uses an RNN to manage a database based on "forecasted" workloads (I know it's not quite architecture, but still): http://pelotondb.io/ http://pelotondb.io/
- colorincorrect 9y agoPerhaps this paper provides an explanation? https://arxiv.org/pdf/1608.08225.pdf https://arxiv.org/pdf/1608.08225.pdf "The exceptional simplicity of physics-based functions hinges on properties such as symmetry, locality, compositionality and polynomial log-probability, and we explore how these properties translate into exceptionally simple neural networks approximating both natural phenomena such as images and abstract representations thereof such as drawings."
- pizza 9y agoThat was an interesting read - although I guess it would have been cool had they delved deeper into the brain/cognition aspect of the introduction a bit more..
- xfs 9y agoThis is really good. Deep learning right now is giving off a kind of illusion of domain-independent general intelligence that can solve any problem, so it would be really helpful to have some theoretical characterization of the specific problem domains it's good at and ones it's not good at.
- irascible 9y agoWhat if the unreasonable effectiveness is the singularity discovering limited time travel and incrementally pushing back the onset of the singularity, in order to understand what happened before the big bang?
- 0verride 9y ago"[...] I mean something that causes people to wonder whether computer architects, programmers, compiler writers or operating systems builders will soon be joining truck drivers and limo drivers in the unemployment line! " How comparable are these jobs?
- mfreed 9y agoKeith Winstein & collaborators have good work about using ML to train TCP's congestion control in different scenarios: http://web.mit.edu/remy/ http://web.mit.edu/remy/
- austincheney 9y agoPerhaps the biggest hurdle in this regard is the approach to machine learning. Nearly everything I have seen on machine learning is a primer on big data followed by a series of algorithms on making the best and smartest decision upon that mountain of data. This is completely the wrong approach. Machine learning can be done on a dime, provided the proper nurturing and environment, but you have to be willing to make some concessions. First and for most you have to be able to write a program that can make a decision. A simple "if" condition is sufficient. Secondly, that decision is open to modification by asserting the evaluation (the "if" condition) against its result. In this regard the logic is fluid opposed to a series of static conditions written by humans hoping to devise organic decisions. Finally, the decision is allowed to be completely wrong. Wrong decisions are better than either no decision or the same decision without deviation. This is how humans learn and it should be no surprise that computers would benefit from the same approach. The key to getting this right is bounds checking and simplicity. A decision must find a terminal point in which to stop improving upon its outcome, and a narrow set of boundaries must be affirmed to prevent unnecessary deviation. It is perfectly acceptable if some grand master must occasionally prod the infantile program in the right direction. This is also something that people do to other people who are learning. If you can do that you have machine learning. You don't need big data to get this. You certainly don't need complex transportation machines or voice activated software to validate it. AI on a dime. If you can do it on a dime you can certainly do it with a multi-billion dollar budget and thousands of developers.
- rattray 9y agoYou're advocating an evolutionary approach, correct? Doesn't such an approach need lots of examples to trial, before it generalizes broadly? "Big data" is often shorthand for "lots of examples", no?
- bmh100 9y agoThe poster is referring to control theory (often seen in ML as reinforcement learning), while also touching on the explore-exploit tradeoff in optimization more generally.
- dkarapetyan 9y agoI think this is because the kinds of problems that arise in system design are logical and symbolic in nature and the current crop of "AI" has no symbolic reasoning capabilities. All the current hype is about pattern matching. Very good pattern matching but just pattern matching nonetheless. Whereas when constructing a compiler or a JIT it's more like what mathematicians do by setting down some axioms and exploring the resulting theoretical landscape. None of the current hype is about theorem proving or the kinds of inductive constructions that crop up in the process of proving theorems or designing compilers and JITs. For an example of the kind of logical problem optimizers solve you can take a look at: https://github.com/google/souper https://github.com/google/souper. So I don't see how you can take the current neural nets and get them to design a more efficient CPU architecture or a better JIT.
- bmh100 9y agoIt would actually be very straightforward to do so if the costs of testing solutions weren't so high. CPU architecture and JIT code can both be represented as unstructured (non-tabular) data. I even recall a circuit having been optimized by a genetic algorithm a while ago in an experiment. I also recall using LSTM to generate valid code from IIRC examples in Linux. Superiptimizatiom is also a relevant topic. We just need better simulation tools or more resources.
- jacquesm 9y agoGenetic algorithms are not neural nets though.
- xapata 9y agoAnd they're generally worse than simulated annealing.
- dkarapetyan 9y agoI have also seen genetic algorithms used for these kinds of optimization problems. In fact there is a module in postgres that uses genetic algorithms to optimize query plans (https://www.postgresql.org/docs/9.6/static/geqo-pg-intro.html https://www.postgresql.org/docs/9.6/static/geqo-pg-intro.htm...). But I don't put genetic algorithms in the same bucket. Genetic algorithms are a different breed of optimization algorithm compared to neural nets and gradient descent which is what the modern crop of AI is basically all about.
- ganfortran 9y agoI guess computer is much more deterministic than what is required for ML to be useful. ML, in a very inaccurate way, can be seen as: 1.We have observations and conclusions. 2.We don't know exact those observations leads to the conclusions. 3.The assumed procedure that leads the observations to conclusions is called model. 4.With enough pairs of (observation, conclusion), we can train a good model that is good enough to make good decision on future observations. Problem for traditional computer science is that, the system is so deterministic that we know EXACTLY how it works on instruction level, while ML is good at dealing problem that is inherently probabilistic.
- saosebastiao 9y agoIndirectly they have helped quite a bit. Some of the most advanced mathematical and symbolic solvers (MIP, IP, LP, CP, SAT, SMT) have slowly been incorporating machine learning to advance their capabilities. Their use cases in these solvers include: branch prediction, branch selection, constraint evaluation order, solver type selection, search strategy selection, cost estimations for column generation strategies, problem classification, solve time estimation, etc. And since advancements in our abilities at solving symbolic and mathematical problems have directly enabled the current research in PLT and formal systems, I see no reason to discount the impact ML has had in pushing that frontier.
- justicezyx 9y agoSince when the belief that "a silver bullet exists for computer science research" becomes a thing?
- Animats 9y agoThe perceptron scheme for branch prediction (full paper) [1] probably works because it uses far more memory for branch history than the usual approaches. It's not doing a better job with comparable resources. [1] https://www.cs.utexas.edu/~lin/papers/hpca01.pdf https://www.cs.utexas.edu/~lin/papers/hpca01.pdf
- csours 9y agoOT: How and why is this page overriding my control key, and how can I stop it from doing that? I use ctrl+scroll wheel to zoom and it is very annoying when that behavior is overridden.
- leecarraher 9y agoThe latter part of the article seems to focus on deep methods not influencing lower level system architecture and design. Perhaps the reason is that those systems and problems are fundamentally different than the dynamics systems that NNs are finding so much success in. In short, compiler, programming and architecture are very formal exact systems, while driving, tts, stt, image association, etc are nowhere near as controlled of environments.