6 ms·
It’s not clear that all these dialects are actually helping things, or that the design of something from which it’s easy to make new dialects is that helpful.
by subharmonicon 3y ago
It’s not clear that all these dialects are actually helping things, or that the design of something from which it’s easy to make new dialects is that helpful.
Why?
I attended an MLIR online meeting months back and there was someone presenting who had a diagram that literally had a dozen different dialects and converters that were being written between different ones (but not MxN, so it was a strictly relatively arbitrary set of combinations that were being converted between, and mostly one-directional if I recall correctly). This was the “solution” to being able to make various things interoperate, but from a distance it looks a lot like a relatively obvious problem caused by the very nature of the design of MLIR.
- j2kun 3y agoThe main reason to make a new dialect is so that a specific type of optimization is easy to write in that dialect. For example, the affine dialect explicitly exists so that one can implement polyhedral loop optimizations. These were possible in, say, GCC, but it requires "raising" an arbitrary low-level program—full of loops with mixed in memory access, weird control flow, and even GOTO—to identify the structural patterns needed to then implement polyhedral optimizations. It took many engineers over many years, even re-designing the entire GCC IR (into something called "GIMPLE"), to enable this, and IIUC it's still considered one of the more challenging parts of the GCC codebase to deal with. Then with MLIR you want to do ML optimizations, which require even _higher_ level constructions like identifying linear algebra ops, so you can do stuff like tiling optimizations and lower to dedicated hardware instructions. THe high level MLIR dialects enables that information to be preserved as long as possible so that it's easier to implement all the magic.
- subharmonicon 3y agoYes, I completely understand the reason for the dialects. I’ve been working on compilers for decades, and what’s interesting about MLIR is that it’s a bit of an anachronistic approach. People used to write optimizers with multiple levels of IR, often based on the same data structures (but sometimes not), and they were effectively different dialects in the same sense as MLIR. That fell out of favor due to the fact that you end up having to choose and commit to the phase ordering fairly early on and writing separate lowering steps to convert between dialects. So the tide turned toward having a single mid-level IR (and sometimes a single high-level IR for things like specific loop optimizations that was then lowered to that mid-level IR).
- j2kun 3y agoI was unaware! Thanks for that context.
- mathisfun123 3y ago> So the tide turned toward having a single mid-level IR (and sometimes a single high-level IR for things like specific loop optimizations that was then lowered to that mid-level IR). You realize this is only feasible if you have one team working on a compiler for one domain right? Eg Rust's MIR is probably a good target for a systems language like rust but a bad target for a SQL like language. >phase ordering fairly early on and writing separate lowering steps to convert between dialects. I don't see how a single IR solves the phase ordering problem? LLVM IR is a single IR (not talking about backends) and yet you still have phase ordering problems.
- subharmonicon 3y ago> You realize this is only feasible if you have one team working on a compiler for one domain right? It sounds like you think I’m advocating for something. I’m not. These are all just engineering trade offs that depend on your goals. Regarding phase ordering: A single IR allows you to freely reorder passes rather than having to reimplement them if you want to move them earlier or later in the phase order.
- mathisfun123 3y ago> A single IR allows you to freely reorder passes rather than having to reimplement them if you want to move them earlier or later in the phase order. ...i'm not trying to be rude here but... that's not how any of this works... you can scholar.google.com "phase ordering llvm ir" to find thousands of papers that demonstrate.
- subharmonicon 3y agoI have no idea what you mean by “how any of this works”. I didn’t bring up solving the phase ordering problem, you did. I’m simply pointing out that if you have a compiler where you have multiple IRs or dialects of IR, and you have a pass that is written to work on IR “X”, and then at some point after that pass you translate to IR “Y”, if you want to move your pass after that point of translation, you either need to rewrite your pass so that it operates on “Y”, or you need to translate back to “X” again.