4 ms·
A very unfortunate choice of comparison on their end. People use numpy, they don't write matrix multiplications in pure python. This feels like an own-goal, bec
by carbocation 3y ago
A very unfortunate choice of comparison on their end. People use numpy, they don't write matrix multiplications in pure python. This feels like an own-goal, because bad benchmarks reduce trust.
- andy99 3y agoMy theory is there is a "persona" they are targeting with their marketing and it's not someone who knows about optimizing ML code.
- foooorsyth 3y agoThat’s just insulting to the intelligence of their potential users (presumably, smart ML engineers that can see through that gimmick of a benchmark)
- andy99 3y agoThat's the reason for my comment, if they were targeting smart ML engineers that can see through the gimmick, you'd expect them to market differently.
- polygamous_bat 3y agoHow many people fall into the intersection of “wants to do matrix multiplication really fast, and willing to pay for it”, “only knows pure python and can’t touch numpy”? Is that segment of the market worth investing $100M into?
- deleted 3y ago[deleted]
- jillesvangurp 3y agoActually, the point here is that with Mojo, you no longer need to use C/C++ libraries for things that are performance critical since it will just compile down to native speed if you need it to. So you can write nice python code (well with the Mojo additions) and it can be fast and optimized and use things like GPUs, multiple cpu cores and it won't be bogged down by things like the global interpreter lock. Kind of nice. The point of this benchmark is making it clear that things you wouldn't dream of doing with python are fine in mojo. People use numpy because python is stupidly slow. Mojo isn't. You can still use numpy of course. But you don't have to.
- carbocation 3y agoI see your point but respectfully disagree with the conclusion. Comparing to numpy, not native python, would be the right move. Again, I don't write matmuls in pure python. I do compute matmuls in numpy. I don't actually know whether numpy is 90,000x faster than pure python, so this benchmark is not useful for me. If they want to show all 3: mojo, numpy, and pure python, then that might be the best of all worlds. They could brag about being 90,000x faster than python, while at the same time showing the actual slowdown of using a pure python-like language (mojo) compared to a compiled numpy library. Let's say the mojo code ends up being 0.5x as fast as numpy; that would still be a pretty great tradeoff for being able to do it all in one language. If you're a python programmer and want to do something that isn't possible with the existing compiled libraries, this would be a good sell. To me, that's still the real comparison of interest.
- deleted 3y ago[deleted]
- vmchale 3y ago> People use numpy because python is stupidly slow. People use Python because numpy isn't slow! It works quite well for its domain.
- gyrovagueGeist 3y agowell kinda... Numpy is rather slow relative to what you can build with a pure C/C++/Fortran/Rust pipeline. Its API usually prevents memory reuse, putting memory allocations on the critical path. You can't fuse operators. And, all of its operations aside from calls into BLAS (for example it's ufunc elementwise processing) are single threaded. People use numpy bc of the python ecosystem and all the domain specific libraries it is compatible with. It very fast relative to Python and provides aa stable, easy to use, array API.
- foooorsyth 3y agoThat's cool, but this runs contrary to the drop-in interop they're touting. It's a pretty confusing initial benchmark to show off. On the Mojo Overview page, they show some numpy+Python using np.max, then they...rewrite np.max using their SIMD parallel map magic? So is this meant to replace vanilla python matmul (which nobody uses IRL)? No? That was just a benchmark to show off their compiler tricks? Okay, how does it fare against numpy (which is actually used)? Well, it's faster? But I have to write little wrappers around basic functions like np.max to parallelize them myself? Shouldn't that just be in your std lib / invisible to the programmer? I thought it was supposed to be a drop-in speed improvement... I don't really get it. Am I stupid? Maybe I'm stupid.