6 ms·
How Mojo gets a speedup over Python – Part 2
- Kelteseth 3y agoSo TL;DR: Using SIMD and multithreading is faster than doing no optimization in python. The only real comparison here is when not doing any optimization is: > The above code produced a 90x speedup over Python and a 15x speedup over NumPy as shown in the figure below: Am I missing something?
- melling 3y agoGetting >10x speed up isn’t exciting enough for many people? I’ll take it. This is all pretty impressive if I can take my unmodified (slightly modified?) Python code and get that sort of improvement.
- mathisfun123 3y ago> This is all pretty impressive if I can take my unmodified (slightly modified?) Python code and get that sort of improvement. it'll never work as smoothly as they advertise. just hands down, beyond a shadow of a doubt, their claims about supporting "unmodified" Python code are startup hype. how do i know? i could give you a bunch of technical reasons about Python as a language and CPython as the de facto implementation (thereby informing tons of code already written, re extensions) but there's a much simpler way to reason about it: because there are already >10 attempts at this and no one has been able to do it. there's no magic here that any number of dollars or brains could pull off. instead each such project picks a point on the pythonic<->performant design-space tradeoff curve and then asks/expects you to live with that choice. and taking ^ into consideration, mojo is not that special. only thing going for it is chris lattner isn't bad at designing languages so maybe, on its own, it'll be a nice language (but it needs to be open to get any traction on its own).
- Takennickname 3y ago> i could give you a bunch of technical reasons about Python as a language and CPython as the de facto implementation Please do. I'm very interested.
- nvm0n2 3y agoIt's not 10x but GraalPy can speed up unmodified Python by 3.4x on average: https://www.graalvm.org/python/ https://www.graalvm.org/python/ And they've not been going at it that long. A few years at most.
- mathisfun123 3y agograalpy does not fully support C extensions and will have just as hard a time extending support as anyone else. maybe even the hardest because they're plumbing through the JVM which, notoriously, has bad C FFI (at least until recently?).
- nvm0n2 3y agoIt's incomplete but it does support C extensions and can run code with NumPy and other science modules. Their approach is unique which is why it can work (they proved out the idea with ruby already). They compile the modules with LLVM and then extend the Python interpreter/JIT compiler with support for LLVM bitcode. So the JITC compiles both Python and C extensions together as one unit. The interpreter API is then virtualized so that code that looks like a structure read or method call from C is compiled directly down to the optimized machine code being used by the rest of the JITC. In this way the interop overhead can be optimized out. This is all separate tech that goes well beyond a normal FFI. JNI doesn't even get involved at all.
- mathisfun123 3y agofrom reading, some of this isn't quite correct (it's graalvm that supports bitcode), but i have to say i didn't realize that that's what they were doing (compiling python and llvm bitcode to graalvm both). interesting but okay you have to admit that's a fairly "beyond scope" approach - ie they solve the problem of C extensions not by compiling python to C but by compiling both to JVM. anyway thanks.
- anonymouse008 3y agoSo does this mean Swift and Metal offers the same if not better performance enhancements? SIMD is very much a first class citizen as a type there
- ElectronCharge 3y agoNo, Lattner learned from Swift and is avoiding anything except zero-cost abstractions. Also, Swift isn’t very interesting outside the Apple ecosystem, and Metal doesn’t exist outside the Apple ecosystem. Mojo has a real shot at widespread, general-purpose, language adoption!
- brucethemoose2 3y ago> no optimization in python Well, isn't that most Python? If Mojo can pave over the slow interpreted bits I repeatedly dig up in Python profilers, even well maintained projects, with no code changes, that would be huge.
- pantsforbirds 3y agoGood blog post. I do wonder how it would do compare to an implementation of pycuda.
- two_handfuls 3y agoA Python with easy-to-use SIMD and multithreading sounds awesome!
- frakt0x90 3y agoAt least they included numpy in this one. On their last post, after all their optimizations, numpy.matmul() produced almost the exact same throughput as their most optimized example. Would still need to dig in to see if this one has issues. Benchmarks are always such a minefield.
- Certhas 3y agomatmul is a wrapper for BLAS. If you're faster than BLAS you're beating handwritten assembler code specialized per CPU architecture.
- aidenn0 3y agoBut people use numpy for matrix multiplies in Python. Unless they are claiming to be 35k times faster on general-purpose code, the 35k number is absurd.
- balencpp 3y agoA lot of ugly,unreadable code has come into existence because of the need to twist it into NumPy calls. If you can replace these with good old for loops and achieve similar performance, then you've already won. Besides that, there are a lot of code that involves looping that isn't matrix multiplication or covered by NumPy.
- archgoon 3y agoRight; but the point is that the optimizations didn't require an entirely new language; you just take the core logic and write it in an existing language that has decades of optimizations. If you're doing math; there's likely a natural, well defined interface that can be used, so you just call that interface from Python, which has historically always been the point of 'glue' languages :)
- deepsquirrelnet 3y agoI don’t understand this from a goals perspective. What is an “AI compiler” - and why aren’t they comparing benchmarks with technologies more commonly used in AI? I think I should be impressed, but I feel like I’m missing the point.
- bjourne 3y agoI guess the point is that getting the same performance in most other languages requires hundreds of lines of code. Here they are ostensibly achieving that performance using very succinct code. That is pretty nice especially if it integrates well with Python.
- dandiep 3y agoI don't understand the play here for Modular. If this is a worthwhile improvement that is broadly applicable, won't it at some point make it's way into Python, numpy, etc? In Java land we had a bunch of other JVMs over the years offering better performance. Most important things got absorbed into what is now OpenJDK, and the other JVMs, if they even exist at all, are niche players. Performance is a huge focus in Python and ML lands right now, so why would this be any different?
- sebzim4500 3y agoThey aren't just speeding up existing python code, they are making a superset of python which has additional performance features. I guess it's possible that these features will be introduced into cpython etc. but I doubt it.
- gojomo 3y agoIf they've figured out how to deliver performance that Python might get around to in 5-10y, shouldn't they tout that, for people who might want that now? Ultimately promoting the possibility for better performance, & current contrast, is good for prodding other languages/runtimes like Python to match these options. The "important things [get] absorbed" process you mention relies on teams making some "play for" alternatives, to create the impetus to get new things integrated.
- dandiep 3y agoTotally, just trying to understand why this is a $100MM of VC money investment. Is the market that big for this? (Honest question)
- zengid 3y agoI'm pretty excited about Mojo and have been keeping an eye on it's development. I feel like the team has learned a lot from their experience, and are taking the best from languages like Python, Rust, Swift, Hylo (Formerly known as Val), and are taking a really nice pragmatic approach in implementing them so that the language is approachable, but also very safe and fast. Once it's out, I hope someone sits down and makes a SwiftUI-like cross platform UI library with it ;).
- barnabee 3y agoYeah, I've been following and am interested too. Actually more interested in things like UIs, quick API servers, stuff like that than the AI/ML use cases. The idea of most of the ease and approachability of Python, a proper type system, and access to the entire ecosystem of Python libs in a compiled language is pretty compelling.
- zengid 3y agoI agree, I'm excited to use it as a General Purpose language, and see how far the Autotuning feature can go for just normal old apps and servers.
- dist-epoch 3y agoCool, but it has very little to do with Python, except some similar looking syntax. So for a Python programmer with a performance problem, it doesn't look like a solution.
- barnabee 3y agoThey are also building in pretty serious Python interop. You should be able to at least somewhat mix the two or migrate gradually, and still use Python libs for less performance critical code (or if the libs do their performance critical stuff in C++ or whatever and are therefore fast enough).
- leecarraher 3y ago35Kx speedup is not scaled speedup. Throw this, naively parallelizable task at a bigger computer and get 70kx speedup, etc. While i think there are tons of optimizations to be done for python (looking at you GIL) giving access to low level cpu primitives is not one I think that will be broadly adopted by the python community. That's one of the joys of python: system agnostic, looks pretty close to pseudocode, coding. If you want speed, glue together a bunch of compiled code calls, and hope the call overhead isn't too large. Or write cpu intensive operations in numba, or pyrex. At the end of the day, mojo's pay to play programming language harkens back to the early 90's Borland days.
- ElectronCharge 3y ago> 35Kx speed up is not scaled speed up. Right. However, this is a comparison versus Python and the GIL, which can’t do that at all. > While i think there are tons of optimizations to be done for python (looking at you GIL) giving access to low level cpu primitives is not one I think that will be broadly adopted by the python community. It doesn’t need to be, any more than writing Numba or Pyrex is done on a large scale. > That's one of the joys of python: system agnostic, looks pretty close to pseudocode, coding. If you want speed, glue together a bunch of compiled code calls, and hope the call overhead isn't too large. Or write cpu intensive operations in numba, or pyrex. At the end of the day, mojo's pay to play programming language harkens back to the early 90's Borland days. The appeal is having a high level language that compiles to efficient machine (and GPU!) code. One can “drop down” to Python for non performance intensive parts. I think this will be much more of a draw for people coming from C++, Fortran and other older, jankier languages. It looks to hit a sweet spot for real time embedded development VERY well, especially given Rust-like memory safety! Mojo will also be a worthy competitor to Julia in the HPC scientific arena I think…we’ll see!
- catgary 3y agoHave you played with Mojo? It really doesn’t feel high level. I feel like JAX has been eating Julia’s lunch lately, making me think that there’s a real market for a small functional differentiable programming language with good Python interop - like a more polished Dex or Futhark.
- pjmlp 3y agoStill waiting if all of this will be another Swift for Tensorflow, or actually make a difference.
- deleted 3y ago[deleted]
- spencerchubb 3y agoWhy is this a language superset of python rather than a python library? Genuinely asking and not trying to bash
- mrfox321 3y agoThat sounds intractable. How would you differentiate mojo code from vanilla python without a ton of boilerplate at language boundaries.
- thebigspacefuck 3y agoThey lost me with the emoji for file extension. That’s not a world I want to live in.
- ElectronCharge 3y agoYou don’t have to, “.mojo” is equivalent.
- queuebert 3y agoBut I use DOS...
- Takennickname 3y agoYou could get a vps on the cloud
- pjmlp 3y agoHow to access it over Netbios?
- Shorel 3y agoThis, while being an apparently superfluous complaint, would be important for eventual enterprise adoption. Other languages have failed for less visible reasons.
- queuebert 3y agoAs a high-performance computing person, I'm usually I/O bound, not compute bound. I wish someone would come up with a 10x speed up for disk and network I/O.
- CoreyFieldens 3y agoI'm really interested in Mojo not for its AI applications, but as an alternative to Julia for high performance computing. Like Julia, Mojo is also attempting to solve the two-language problem, but I like that Mojo is coming at it from a Python perspective rather than trying to create new syntax. For better or for worse, Python is absolutely dominating in the field of scientific computing, and I don't see that changing anytime soon. Being able to write optimizations at a lower level in a Python-like syntax is really appealing to me. Furthermore, while I love Julia the language, I'm disappointed in how it really hasn't taken off in adoption by either academia or industry. The community is small and that becomes a real pain point when it comes to tooling. Using the debugger is an awful experience and the VSCode extension that is recommended way to write Julia is very hit-or-miss. I think it would really benefit from a lot more funding that doesn't actually seem to be coming. It's not a 1-to-1 comparison, but Modular has received 3 times the amount of funding as JuliaHub despite being much younger.
- pjmlp 3y agoThey already failed once with Swift for Tensorflow, so I am currently curious if there will be some lessons learned from that effort. For the time being, my chips are still on the Julia horse.
- ElectronCharge 3y agoI’m a huge Julia fan, you can take a look at my posting history. I love Julia’s syntax, and some of its language ideas. …BUT… For my personal tastes, Mojo’s lack of garbage collection, Rust-like memory safety, and attention to ahead-of-time compilation put it way ahead. The vast pool of Python developers who can easily pick it up if interested is a big plus. Julia is aimed at a somewhat different space, but there’s also a huge overlap. Let’s hope for good interoperability between the two, it seems fairly straightforward…
- pjmlp 3y agoLets see how it plays out, given that they are focused only on AI workloads, and somehow those VCs want their money back, which doesn't appeal to everyone. I acknowledge that there is finally pressure in the Python community to tackle down performance, but don't see Mojo being the solution unless there is something that it will make it go wild. Right now, I see that more likely with Facebook, NVidia, Intel and Microsoft efforts.
- laweijfmvo 3y agonit: The text says 743x but the graph (Figure 3) shows 527x
- brrrrrm 3y agoI just want to see real un-hyped benchmarks. Comparing random Python native code makes no sense and seems dishonest, deterring me from actually trying out the tool. I want a Python that can statically plan underlying GPU allocations, avoids CUDA kernel dispatch overhead and enables a multi-GPU API that isn't some multiprocessing abomination.
- erichocean 3y agoMojo needs to demonstrate Hugging Face's AI libraries with Mojo acceleration. Nothing else will have the kind of impact that would have. Throw a half dozen engineers at it, develop a deployment plan for SD XL, profit. You'll get a ton of open source developers working on improving the Mojo versions even further once you release it, researchers developing extensions, etc. GO TO WHERE THE DEVELOPERS ARE. Stable Diffusion is crazy compute heavy, so if Mojo is what it's purported to be, it should be possible to get speedups.