6 ms·
Making Cython as easy as Python
- vog 11y agoThis strategy should be more common! In the source tarball, there should be one main program which you start. The compiling should be an implementation detail of the main program. The same for web applications. I never liked the idea of reintroducing a separate build (concat/minify/compress) process whenever the sources change. Instead, the production mode should be as simple as the debug mode: When the main page is requested per HTTP, first rebuild the client JS app as needed, then deliver it to the browser. Caching the build result should be no different from any other caching of automatically-generated resources. I use this approach for almost all new software that I write and I'm quite satisfied with that. The main program (e.g. "bin/mytool") is a wrapper shell script or Python script that runs the build as needed, then starts the program. In the simplest case, this is a wrapper around "make && ./run_tool". Of course, the program should indicate the build with a small message to stderr, especially if the build takes some seconds or more. Moreover, it should be possible to trigger the build individually such that distros can build packages from that. On the other hand, running the program with "-h" or "--help" is a kind of build-only process, too.
- icebraining 11y agoIt's practical, but on the other hand, that goes against the security practice of not allowing the executable zone (be it memory or filesystem) to be writeable. I personally prefer to precompile stuff, even running "python -m compileall" to create .pyc files, to avoid having to keep that path writeable by the user running the program.
- vog 11y agoGood point! However, if you go down that route in a web application, don't forget to pre-compile templates, and do the same for all other cached things that are "code-like". More generall, I find W^X hard to enforce in typical web applications. Sure you can mount directory read-only, disable access to /tmp and observe what happens. But then they might use the database as template cache, or whatever. In most (web) application I'd be happy if they had a central directory where they write stuff into (bonus points for making that directory configurable), instead of scattering generated files all across the source tree.
- dkersten 11y agoI like to write my code to assume that it will run the same on dev and prod, but then precompile/minify prod code as part of the deployment. That is, I like to set it up like you say, but then not rely on it in production. Mainly because I want as few moving parts as possible on prod (especially if it means I don't need dev tools/libs on the production server).
- SixSigma 11y agoBetter to use a file system based file change monitor and a makefile. Statting your file system for changes on every request adds to latency.
- deleted 11y ago[deleted]
- vegabook 11y agoCython, Numpy, Numba (and others) are what make me skeptical that any of the numerical computing competitors to Python (Julia, or to a lesser extent Lua/JS/Clojure even compiled Scala etc) can displace it at. Why would you abandon this wonderfully friendly, malleable language with easily the largest set of options for anything you might like to do, for something nominally faster on an artificial benchmark (which inevitably ignores the above 3 libraries), when by dipping into any of these you get C-speed in such an accessible and friendly way, most of the time beating the competitors easily on performance. The only real threats to Python to me come would from a properly parallelized language that would use the GPU/multicore natively in vector form. I don't see that anywhere yet.
- zzleeper 11y agoI've seen some very impressive Julia libraries that perform almost at the level of C code, so I'm not sure if the gap that Julia needs to fill is that large. Also, a quick question: which one is preferred nowadays, Cython or Numba?
- dagw 11y agoCython is much more mature, works on basically all python code and basically never leads to slowdowns compared to pure python. Unfortunately non of that can really be said about Numba. That being said, when Numba works it's great and is much easier to work with than cython.
- synparb 11y agoI'll echo @dagw's comments. Cython has been rock solid for me for a long time and it is my go to for any sort of external c/c++ library interfacing. That said, I find myself using Numba more and more in the places that I can as Numba has gotten significantly better in the last 6 months or so. It's still not a complete replacement for cython (and I don't think it ever will or intends to be), but for hot spot numerical calculations it's really nice since it's much faster to test things out since it doesn't involve the boiler plate required by Cython, and doesn't require the (often slow) compilation times. In Numba 0.21.0, on-disk caching of jit'd code was also introduced, which was one of the major sticking points for us to put Numba into production in areas where we needed faster start-up times. Before we could really only use Cython because we required the start-up times available only from AOT compilation. That all said, Numba has been a bit buggy for me at times, although these get squashed pretty quickly. I've only found a single bug in Cython in all of the years I've been using it, and it's amazing how quickly Robert Bradshaw or Stefan Behnel respond and fix things considering Cython is not their full time job.
- rplnt 11y agoSeeing how this uses mktemp, I assume the results of runcython build are not reused. Is that correct? If sou, wouldn't it make sense to save the results? Maybe with a switch? Also feature requests: Windows support for run/make would be nice. Cross-compile options for make would be awesome.
- python_que 11y agoI am using Python to solve combinatorial problems, and so my code relies heavily on the itertools library. My question is if I would still get a considerable speed-up if I rewrite some of my code in Cython, given my reliance on itertools?
- dagw 11y agoThe calls to itertools won't be sped up significantly. However if then you loop over the results of a call to itertools and do something to each element, then that loop might end up being much faster.
- versteegen 11y agoAs dagw said, if your inner loops aren't inside a .pyx file, then Cython can do little to help. Even if there are, if it involves lots of fiddling with data structures like dicts and lists, or making use of dynamic or high level features of Python like generators or classes or list comprehensions, then Cython would provide only a small speed increase (say, 20%), because almost all the running time will be spent inside the Python runtime. To get the biggest speed ups out of Cython you need to write C-like code with type annotations. E.g. "for x in range(...):" loops.
- hogu 11y agoit may be worth checking out cytoolz
- mathnode 11y agoYup. Cytoolz. Some notes on performance of Cytoolz: http://matthewrocklin.com/blog/work/2014/05/01/Introducing-CyToolz/ http://matthewrocklin.com/blog/work/2014/05/01/Introducing-C...
- python_que 11y agogreat thanks both for the tip.
- cbsmith 11y ago"Runcython aims to simplify the process of using Cython without sacrificing scalability." I know this is being a semantic weenie, but I hate when people use scalability when they mean efficiency.
- chuckbot 11y agoAnd I hate when people use efficiency when they mean speed. Python -> Cython speeds up you program but does not change the complexity of your algorithm.
- oliwarner 11y agoHow do you suppose it speeds it up then? I'll give you "the complexity of your algorithm" likely doesn't change, but how that algorithm is processed is usually more efficient when compiled through C, rather than running through a Python interpreter, if only because of how it's running.
- deleted 11y ago[deleted]
- chuckbot 11y agoI agree it is faster by running it through compiled C code. Saying "my code is more efficient than your code" might be acceptable, but saying "my code is efficient" drives me mad when it is used as a synonym for "my code is fast". Take a look at [1] and the note that this is not about optimization. Efficient algorithms are usually algorithms with close to optimal time or space complexitiy. https://en.wikipedia.org/wiki/Algorithmic_efficiency https://en.wikipedia.org/wiki/Algorithmic_efficiency
- williamstein 11y agoSometimes it does. For example if you write a simple O(n) for loop in Python, then convert it to Cython, then compile it, the compiler may replace the for loop by an equivalent O(1) algorithm. I use this to surprise students when benchmarking Cython code for teaching purposes in my SageMath course.
- willvarfar 11y agoIts an aside, but recently Python 3 has got annotations. It would be nice if cython supported annotation syntax, so that type-annotated cython was as much as possible valid python than can also be run through an interpreter. I understand a lot of the nuances, particularly the lack of in-function annotations etc. My own foray into Python type annotations is obiwan https://pypi.python.org/pypi/obiwan/ https://pypi.python.org/pypi/obiwan/
- halosghost 11y agoThis is great to see! I don't know if it was inspired by runhaskell, but it does appear to be similar. I have to say, this kind of tool is something I have come to feel is incredibly helpful and important for my dev workflow. I spend most of my time developing in statically-typed, compiled-to-machine-code languages; obviously, one of the major advantages of more dynamic or interpreted languages is the ability to rapidly prototype. This type of tool allows me to quickly prototype something without losing all the power, control and other benefits of my languages of choice (in particular, Haskell these days). Things like this make me more amenable to python by the day! Keep up the good work!