5 ms·
Not having any particular domain experience here, I've idly wondered whether or not there's any role for neural net models in translating code for other archite
by peatmoss 4y ago
Not having any particular domain experience here, I've idly wondered whether or not there's any role for neural net models in translating code for other architectures.
We have giant corpuses of source code, compiled x86_64 binaries, and compiled arm64 binaries. I assume the compiled binaries represent approximately our best compiler technology. It seems predicting an arm binary from an x86_64 binary would not be insane?
If someone who actually knows anything here wants to disabuse me of my showerthoughts, I'd appreciate being able to put the idea out of my head :-)
- qsort 4y agoYou would need a hybrid architecture with a NN generating guesses and a "watchdog" shutting down errors. Neural models are basically universal approximators. Machine code needs to be obscenely precise to work. Unless you're doing something else in the backend, it's just a turbo SIGILL generator.
- throw10920 4y agoThis is all true - machine code needs to be "basically perfect" to work. However, there are lots of problems in CS that are easier to check the answer to a solution than to solve in the first place. It may turn out to be the case that a well-tuned model can quickly produce solutions to some code-generation problems, that those solutions have a high enough likelihood of being correct, that it's fast enough to check (and maybe try again), and that this entire process is faster than state-of-the-art classical algorithms. However, if that were the case, I might also expect us to be able to extract better algorithms from the model - intuitively, machine code generation "feels" like something that's just better implemented through classical algorithms. Have you met a human that can do register allocation faster than LLVM?
- classichasclass 4y ago> turbo SIGILL generator This gave me the delightful mental image of a CPU smashing headlong into a brick wall, reversing itself, and doing it again. Which is pretty much what this would do.
- Someone 4y ago> It seems predicting an arm binary from an x86_64 binary would not be insane? If you start with a couple of megabytes of x64 code, and predict a couple of megabytes of arm code from it, there will be errors even if your model is 99.999% accurate. How do you find the error(s)?
- hinkley 4y agoI think we are on the cusp of machine aided rules generation via example and counter example. It could be a very cool era of “Moore’s Law for software” (which I’m told software doubles in speed roughly every 18 years). Property based testing is a bit of a baby step here, possibly in the same way that escape analysis in object allocation was the precursor to borrow checkers which are the precursor to…? These are my inputs, these are my expectations, ask me some more questions to clarify boundary conditions, and then offer me human readable code that the engine thinks satisfies the criteria. If I say no, ask more questions and iterate. If anything will ever allow machines to “replace” coders, it will be that, but the scare quotes are because that shifts us more toward information architecture from data munging, which I see as an improvement on the status quo. Many of my work problems can be blamed on structural issues of this sort. A filter that removes people who can’t think about the big picture doesn’t seem like a problem to me.
- brookst 4y agoI'm a ML dilletante and hope someone more knowledgeable chimes in, but one thing to consider is the statistics of how many instructions you're translating and the accuracy rate. Binary execution is very unforgiving to minor mistakes in translation. If 0.001% of instructions are translated incorrectly, that program just isn't going to work.
- Symmetry 4y agoMany branch predictors have traditionally used perceptrons, which are sort of NN like. And I think there's a lot of research into involving incorporating deep learning models into doing chip routings.
- saagarjha 4y agoPeople have tried doing this, but not typically at the instruction level. Two ways to go about this that I’m aware of are trying to use machine learning to derive high-level semantics about code, then lowering it to the new architecture.