4 ms·
One of my biggest points of criticism of Python is its slow cold start time. I especially notice this when I use it as a scripting language for CLIs. The startu
by randomtoast 10mo ago
One of my biggest points of criticism of Python is its slow cold start time. I especially notice this when I use it as a scripting language for CLIs. The startup time of a simple .py script can easily be in the 100 to 300 ms range, whereas a C, Rust, or Go program with the same functionality can start in under 10 ms. This becomes even more frustrating when piping several scripts together, because the accumulated startup latency adds up quickly.
- nickjj 10mo ago> The startup time of a simple .py script can easily be in the 100 to 300 ms range I can't say I've ever experienced this. Are you sure it's not related to other things in the script? I wrote a single file Python script, it's a few thousand lines long. It can process a 10,000 line CSV file and do a lot of calculations to the point where I wrote an entire CLI income / expense tracker with it[0]. The end to end time of the command takes 100ms to process those 10k lines, that's using `time` to measure it. That's on hardware from 2014 using Python 3.13 too. It takes ~550ms to fully process 100k lines as well. I spent zero time optimizing the script but did try to avoid common pitfalls (drastically nested loops, etc.). [0]: https://github.com/nickjj/plutus https://github.com/nickjj/plutus
- tlyleung 10mo agoJust a guess - but perhaps the startup time is before `time` is even imported?
- williadc 10mo ago`time` is a shell command that you can use to invoke other commands and track their runtime.
- zahlman 10mo ago> I can't say I've ever experienced this. Are you sure it's not related to other things in the script? I wrote a single file Python script, it's a few thousand lines long. It's because of module imports, primarily and generally. It's worse with many small files than a few large ones (Python 3 adds a little additional overhead because of needing extra system calls and complexity in the import process, to handle `__pycache__` folders. A great way to demonstrate it is to ask pip to do something trivial (like `pip --version`, or `pip install` with no packages specified), or compare the performance of pip installed in a venv to pip used cross-environment (with `--python`). Pip imports literally hundreds of modules at startup, and hundreds more the first time it hits the network.
- fwip 10mo agoAnd it's worse if your python libraries might be on network storage - like in a user's homedir in a shared compute environment.
- dekhn 10mo agoExactly this. The time to start python is roughly a function of timeof(stat) * numberof(stat calls) and on a network system that can often be magnitudes larger than a local filesystem.
- zahlman 10mo agoI do wonder, on a local filesystem, how much of the time is statting paths vs. reading the file contents vs. unmarshaling code objects. (Top-level code also runs when a module is imported, but the cost of that is of course highly module-dependent.)
- dekhn 10mo agoMaybe you could take the stat timings, the read timings (both from strace) and somehow instrument Python to output timing for unmarshalling code (or just instrument everything in python). Either way, at least on my system with cached file attributes, python can startup in 10ms, so it's not clear whether you truly need to optimize much more than that (by identifying remaining bits to optimize), versus solving the problem another way (not statting 500 files, most of which don't exist, every time you start up).
- nickjj 10mo agoMakes sense, most of my scripts are standalone zero dependency scripts that import a few things from the standard library. `time pip3 --version` takes 230ms on my machine.
- maccard 10mo agoThat proves the point, right? `time pip3 --version` takes ~200ms on my machine. `time go help` takes 25, and prints out 30x more lines than pip3 --version.
- randomtoast 10mo agoHere is a benchmark https://github.com/bdrung/startup-time https://github.com/bdrung/startup-time This benchmark is a little bit outdated but the problem remains the same. Interpreter initialization: Python builds and initializes its entire virtual machine and built-in object structures at startup. Native programs already have their machine code ready and need very little runtime scaffolding. Dynamic import system: Python’s module import machinery dynamically locates, loads, parses, compiles, and executes modules at runtime. A compiled binary has already linked its dependencies. Heavy standard library usage: Many Python programs import large parts of the standard library or third-party packages at startup, each of which runs top-level initialization code. This is especially noticeable if you do not run on an M1 Ultra, but on some slower hardware. From the results on Rasperberry PI 3: C: 2.19 ms Go: 4.10 ms Python3: 197.79 ms This is about 200ms startup latency for a print("Hello World!") in Python3.
- zahlman 10mo agoInteresting. The tests use Python 3.6, which on my system replicates the huge difference shown in startup time using and not using `-S`. From 3.7 onwards, it makes a much smaller percentage change. There's also a noticeable difference the first time; I guess because of Linux caching various things. (That effect is much bigger with Rust executables, such as uv, in my testing.) Anyway, your analysis of causes reads like something AI generated and pasted in. It's awkward in the context of the rest of your post, and 2 of the 3 points are clearly irrelevant to a "hello world" benchmark.
- maccard 10mo agoA python file with import requests Takes 250ms on my i9 on python 3.13 A go program with package main import ( _ "net/http" ) func main() { } takes < 10ms.
- dotdi 10mo agoThis is not an apples-to-apples comparison. Python needs to load and interpret the whole requests module when you run the above program. The golang linker does dead code elimination, so it probably doesn't run anything and doesn't actually do the import when you launch it.
- maccard 10mo agoSure it's not an apples to apples comparison - python is interpreted and go is statically compiled. But that doesn't change the fact that in practice running a "simple" python program/script can take longer to startup than go can to run your entire program.
- dotdi 10mo agoStill, you are comparing a non-empty program to an empty program.
- tuhgdetzhh 10mo agoEven if you actually use the network module in Go, just so that the compiler wouldn't strip it away, you would still have a startup latency in Go way below 25 ms from my experience with writing CLI tools. Whereas with Python, even in the latest version, you're already looking at atleast 10x the amount of startup latency in practice. Note: This is excluding the actual time that is made for the network call, which can of course also add quiete some milliseconds, depending on how far on planet earth your destination is.
- maccard 10mo agoYou're missing the point. The point is that python is slow to start up _because_ it's not the same. Compare: import requests print(requests.get("http://localhost:3000").text) to package main import ( "fmt" "io" "net/http" ) func main() { resp, _ := http.Get("http://localhost:3000") defer resp.Body.Close() body, _ := io.ReadAll(resp.Body) fmt.Println(string(body)) } I get: python3: 0.08s user 0.02s system 91% cpu 0.113 total go 0.00s user 0.01s system 72% cpu 0.015 total (different hardware as I'm at home). I wrote another that counts the lines in a file, and tested it against https://www.gutenberg.org/cache/epub/2600/pg2600.txt https://www.gutenberg.org/cache/epub/2600/pg2600.txt I get: python 0.03s user 0.01s system 83% cpu 0.059 total go 0.00s user 0.00s system 80% cpu 0.010 total These are toy programs, but IME that these gaps stay as your programs get bigger
- baq 10mo agoit depends somewhat on what you import, too. some people would sell their grandmothers to get below 1s when you start importing numpys and scikits.
- zbentley 10mo agoThe upcoming lazy import system may help with startup time…but if the underlying issue wasn’t “Python startup is slow” but rather “a specific program imports modules that take a long time to low”, it’ll only shift the time consumption to runtime.
- kortex 10mo agoThat's totally fine, because many CLI tools are organized like `mytool subcommand --params=a,b...`, and breaking out those subcommands into their own modules and lazy loading everything (which good CLI tools already know to do) means unused code never gets imported. You can already lazy import in python, but the new system makes the syntax sweeter and avoids having to have in-function `import module` calls, which some linters complain about.
- TudorAndrei 10mo agoAre you comparing the startup time of an interpreted language with the startup time of a compiled language? or you mean that `time python hello.py` > `( time gcc -O2 -o hello hello.c ) && ( time ./hello )` ?
- randomtoast 10mo agoI'm referring to the startup time as benchmarked in the following manner: https://github.com/bdrung/startup-time https://github.com/bdrung/startup-time
- maccard 10mo agoHere's the thing - I don't really care if its' because the interpreter has to start up, or there's a remote http call, or we scan the disks for integrity - the end user experience on every run is slower.
- smartmic 10mo agoYes, that is also my feeling. But comparing an interpreted language with a compiled one is not really fair. Here is my quick benchmark. I refrain from using Python for most scripting/prototyping task but really like Janet [0] - here is a comparison for printing the current time in Unix epoch: $ hyperfine --shell=none --warmup 2 "python3 -c 'import time;print(time.time())'" "janet -e '(print (os/time))'" Benchmark 1: python3 -c 'import time;print(time.time())' Time (mean ± σ): 22.3 ms ± 0.9 ms [User: 12.1 ms, System: 4.2 ms] Range (min … max): 20.8 ms … 25.6 ms 126 runs Benchmark 2: janet -e '(print (os/time))' Time (mean ± σ): 3.9 ms ± 0.2 ms [User: 1.2 ms, System: 0.5 ms] Range (min … max): 3.6 ms … 5.1 ms 699 runs Summary 'janet -e '(print (os/time))'' ran 5.75 ± 0.39 times faster than 'python3 -c 'import time;print(time.time())'' [0]: https://janet-lang.org/ https://janet-lang.org/
- curiousgal 10mo agoWell python is also compiled technically.
- syrusakbary 10mo agoCompletely agree on this. Regarding cold-starts, I strongly believe V8 snapshots are perhaps not the best way to achieve fast cold starts with Python (they may be if you are tied to using V8, though!), and will have wide side effects if you go out of the standards packages included on the Pyodide bundle. To put some perspective: V8 snapshots are storing the whole state of an application (including it's compiled modules). This means that for a Python package that is using Python (one wasm module) + Pydantic-core (one wasm module) + FastAPI... all of those will be included in one snapshot (as well as the application state). This makes sense for browsers, where you want to be able to inspect/recover everything at once. The issue about this design is that the compiled artifacts and the application state are bundled into one piece artifact (this is not great for AOT designed runtimes, but might be the optimal design for JITs though). Ideally, you would separate each of the compiled modules from the state of the application. When you do this, you have some advantages: you can deserialize the compiled modules in parallel, and untie the "deserialization" from recovering the state of the application. This design doesn't adapt that well into the V8 architecture (and how it compiles stuff) when JavaScript is the main driver of the execution, however it's ideal when you just use WebAssembly. This is what we have done at Wasmer, which allows for much faster cold starts than 1 second. Because we cache each of the compiled modules separately, and recover the state of the application later, we can achieve cold-starts that are a magnitude faster than Cloudflare's state of the art (when using pydantic, fastapi and httpx). If anyone is curious, here is a blogpost where we presented fast-cold starts for the application state (note that the deserialization technique for Wasm modules is applied automatically in Wasmer, and we don't showcase it on the blogpost): https://wasmer.io/posts/announcing-instaboot-instant-cold-starts-for-serverless-apps https://wasmer.io/posts/announcing-instaboot-instant-cold-st... Note aside: congrats to the Cloudflare team on their work on Python on Workers, it's inspiring to all providers on the space... keep it up and let's keep challenging the status quo!
- ndr 10mo agoIt might not be the fastest but I suspect something weird is happening with python resolution. For instance `uv run` has its own fair share of overhead. $ hyperfine --warmup 10 -L py "uv run python,~/.local/bin/python3.14,/usr/local/bin/python3.12,~/.local/share/uv/python/pypy-3.11.13-macos-aarch64-none/bin/pypy3.11" "{py} -c 'exit(0)'" Benchmark 1: uv run python -c 'exit(0)' Time (mean ± σ): 58.4 ms ± 19.3 ms [User: 26.4 ms, System: 21.7 ms] Range (min … max): 48.2 ms … 138.0 ms 50 runs Benchmark 2: ~/.local/bin/python3.14 -c 'exit(0)' Time (mean ± σ): 13.3 ms ± 6.9 ms [User: 8.0 ms, System: 2.5 ms] Range (min … max): 9.9 ms … 53.7 ms 174 runs Benchmark 3: /usr/local/bin/python3.12 -c 'exit(0)' Time (mean ± σ): 16.4 ms ± 7.6 ms [User: 8.9 ms, System: 3.7 ms] Range (min … max): 12.2 ms … 65.2 ms 152 runs Benchmark 4: ~/.local/share/uv/python/pypy-3.11.13-macos-aarch64-none/bin/pypy3.11 -c 'exit(0)' Time (mean ± σ): 18.6 ms ± 7.4 ms [User: 10.0 ms, System: 5.0 ms] Range (min … max): 14.4 ms … 63.5 ms 138 runs Summary ~/.local/bin/python3.14 -c 'exit(0)' ran 1.23 ± 0.86 times faster than /usr/local/bin/python3.12 -c 'exit(0)' 1.40 ± 0.92 times faster than ~/.local/share/uv/python/pypy-3.11.13-macos-aarch64-none/bin/pypy3.11 -c 'exit(0)' 4.40 ± 2.72 times faster than uv run python -c 'exit(0)'
- dilawar 10mo agoReminds me of mercurial cvs!!
- yegle 10mo agoYes it's bad enough that there's a chg to (barely) improve the command laten y. (Side note this is why jj is awesome. A `jj log` is almost as fast as `ls`).
- mixmastamyk 10mo agoBig packages shouldn’t be imported until the cli has been parsed, and handed off to main. There’s been work to do this automatically, but it’s good hygiene to avoid it anyway. A modern machine shouldn’t take this long, so likely something big is being imported unnecessarily at startup. If the big package itself is the issue, file it on their tracker.
- dekhn 10mo agoRun strace on Python starting up- you will see it statting hundreds if not thousands of files. That gets much worse the slower your filesystem is. On my linux system where all the file attributes are cached, it takes about 12ms to completely start, run a pass statement, and exit.
- rcarmo 10mo agoYou can run .pyc stuff “directly” with some creativity, and there are some tools to pack “executables” that are just chunked blobs of bytecode.
- az09mugen 10mo agoI don't know why people care so much about a few hundreds of milliseconds for python scripts versus compiled languages that take just ten times less. Real question : what would you do more with the spared time ? You are that in a hurry in your life ?
- paulddraper 10mo agoIt’s not Python runtime startup, but loading all your modules. Use lazy/dynamic imports and you will see it drop .