7 ms·
When I was in school and first leaning about programming I assumed that code written in C or Java would eventually be ported to hand tuned assembler once enough
by johnwatson11218 7y ago
When I was in school and first leaning about programming I assumed that code written in C or Java would eventually be ported to hand tuned assembler once enough people were using it. Then I got in to the industry and realized that we just keep adding layer after layer until we end up at the point this article talks about.
I remember once reading that IBM was going to implement an XML parser in assembler and people were like "Why? If speed is needed then you shouldn't use XML anyway." I thought that concern was invalid because these days XML ( or JSON ) is really non-negotiable in many scenarios.
One idea that I've been thinking about lately is some kind of neural network enabled compiler and/or optimizer. I have heard that in the javascript world they have something called the "tree shaking" algorithm where they run the test suite, remove dependencies that don't seem to be necessary and repeat until they are getting test failures.
I'm thinking why not train a LSTM to take in http requests and generate the http response? Of course sometime the request would lead to some sql, which you could then execute and feed the results back into the LSTM until it output a http response. Then try using a smaller network until something like your registration flow, or a simple content management system was just a bunch of floating point numbers in some matrices saved off to disk.
- MaxBarraclough 7y ago> I'm thinking why not train a LSTM to take in http requests and generate the http response? Why? With responses generated according to what? Are you really just suggesting using neural networks in the compiler's optimiser? > Then try using a smaller network until something like your registration flow, or a simple content management system was just a bunch of floating point numbers in some matrices saved off to disk. Why? What's the advantage over just building software?
- johnwatson11218 7y agoI'm suggesting that you take an existing system and build up a corpus of request/response pairs. Then you use the LSTM to build a prediction model so that given a request it will tell you that the current production system will produce the following sql statement and this http response. Once the LSTM's output is indistinguishable from your current production system , for all use cases, then you replace the production system with the LSTM and a thin layer that can listen on the port, encode/decode the data, and issue sql queries. Why would I want to do this? I'm not 100% sure ... I think it would be super fast once you got it working. I think it would avoid many security bugs. You wouldn't have to read that "oh drupal 3.x has 20 new security bugs" better go patch our code. I think when I had this idea I was thinking about it terms of a parallel system that could catch hacking by noting when actual http responses diverged too much from the predicted response. The main idea being that for a given input the output really is 100% predictable, assuming your app doesn't use random numbers like in a game or something. To link this idea to the article, I think things like XML parsers could be written this way .... I can't prove it but I suspect that they would be very fast and not come with all the baggage that the article complains about. I started thinking along these lines after reading stuff like this https://medium.com/@karpathy/software-2-0-a64152b37c35 https://medium.com/@karpathy/software-2-0-a64152b37c35
- capitalsigma 7y agoWhat if your app has literally any mutable state? Registering accounts, posting comments, etc. Also I'll bet you that your neural net is > 100x slower than straight line code.
- johnwatson11218 7y agoMutable state in the sense of database writes would be part of the network's output and just passed on to a regular db. Mutable state in the sense of variables that the application code uses while processing a request? Well LSTM networks can track state like that. For session based variables? Not sure, either it all becomes stateless and the code has to read everything from storage for each reqeuest .... or maybe the lstm is able to model something like an entire user session and remember the stuff that the original app would have put in the session. That Andrej Karpathy article that I linked to two comments above ... he pointed out, in a different blog post, that regular neural networks can approximate any pure function. Recurrent neural networks like the LSTM can approximate any computer program. It is because they can propagate state from step to step that allows them to do this. As far as it being 100X slower, well at a certain point I will be willing to take your money :)
- jodrellblank 7y agoThe main idea being that for a given input the output really is 100% predictable [..] I think it would be super fast once you got it working. I imagine it would be fast, then you realise you've made a static content caching layer out of a neural network and replace it with Varnish cache and it would be hyper fast.
- johnwatson11218 7y agoI don't think a caching layer would work. One example would be an online mortgage estimator. You input the loan amount, interest rate, length of loan etc. all as http input parameters. I'm suggesting that the LSTM can eventually figure out that those variables are being used by the application code to go in to a formula. That application code and its formula would all be replaced by the LSTM. I just don't know how you can achieve that with static cache ... only if somebody else requested that exact mortgage calculation before and it is still in the cache. Also, my idea of the "given input" from the earlier comment would have to include results of sql queries that would form the entire input to the LSTM. But honestly I think over trained auto encoders can be used as hash maps. That would be an application more in line with what I think you are saying.
- ncmncm 7y agoXML actually is parsed with assembly code, now, using vector instructions that split up bytes, bits 0 of 128 or more bytes in XMM0, bits 1 of the same bytes in XMM1, et al., and doing bitwise operations on the registers to recognize features. Imagine how bad it would be if not!
- pedagand 7y agoThat's interesting: would you happen to have any reference implementing/describing such a parser, by any chance? In the crypto world, this is called "bitslicing": https://www.bearssl.org/constanttime.html#bitslicing https://www.bearssl.org/constanttime.html#bitslicing
- Pete_D 7y agoIt's not NN-based, but you might be interested in https://en.wikipedia.org/wiki/Superoptimization https://en.wikipedia.org/wiki/Superoptimization.