3 ms·
Just for receipts I span up a basic test for XLA: Can you figure out that 16 matrix vector multiplications into a concatenate is the same as concatenating firs
by kingstnap 2mo ago
Just for receipts I span up a basic test for XLA:
Can you figure out that 16 matrix vector multiplications into a concatenate is the same as concatenating first into a matrix matrix operation which can go on the GEMM. This is like the most basic thing you can imagine doing.
Turns out no it doesn't and there is a 23% performance difference by moving the concatenate up in the python code in my specific test.
Not to say that it didn't recognize it. I had an agent look at the XLA and the graph actually does a partial fusion into a sum and stack. But does not realize the whole thing is just a GEMM.
In more complicated examples the differences you can get can be much larger.
- deleted 2mo ago[deleted]
- Gangway0829 2mo agoIs this surprising? My experience with compilers from the old days of Fortran is that they care about correctness first, performance second. I used to spend plenty of time rewriting algorithms so they ended up in a form the compiler would like. I think this is a great area for LLMs. You write the correct physics code, the LLM analyzes intent and goes back and forth with the compiler and your correct code to rewrite it in something that emits performant code