3 ms·
I'm excited to see more profiling tools for Python! This sounds like it does peak memory, which is critical for batch jobs, since that's the bottleneck. Memory
by itamarst 4y ago
I'm excited to see more profiling tools for Python!
This sounds like it does peak memory, which is critical for batch jobs, since that's the bottleneck. Memory is fundamentally different than performance in that it's a limited resource, instead of cumulative cost; making any part of the program faster almost always helps speed up the program (at least a little, or at least reduces CPU load), but optimizing non-peak memory has no impact. You have to be able to identify the peak in order to reduce memory usage.
If you want peak memory profiling for Python that also runs on macOS, check out https://pythonspeed.com/fil/ https://pythonspeed.com/fil/ (ARM support has some issues, but once I unpack my new Mac Mini I plan to fix it.)
Ways memray is better than Fil:
- Native callstacks.
- More kinds of reports, and ability to do custom post-processing of data.
- Much lower overhead (but not always, see reply).
- Subprocess support.
Fil I suspect has better flamegraphs: https://pythonspeed.com/articles/a-better-flamegraph/ https://pythonspeed.com/articles/a-better-flamegraph/
And if you're running Python batch jobs, and want both peak memory and performance profiling in production, check out Sciagraph: https://pythonspeed.com/sciagraph/ https://pythonspeed.com/sciagraph/
(You can probably cobble together something like Sciagraph with py-spy + memray, but you won't e.g. get timeline reports designed with batch jobs in mind.)
- pablogsal 4y agoIt does much more than that! It tracks every single allocation and dumps it to a file that can later be analysed in many ways. Currently our reporters report peak memory (and leaked memory at the end of the execution) but technically any other reporter can be used. For example, we plan to allow to generate flame-graphs at arbitrary points in the execution and much more!
- itamarst 4y agoBTW I am starting a Slack for devs working on profilers, would be great to have you all join, I'd love to hear more about the ELF patching technique (Fil uses LD_PRELOAD and macOS equivalent).
- itamarst 4y agoCool! Fil also tracks every allocation too, although it doesn't dump that at the moment, just the resulting report.
- pablogsal 4y agoNice! Fil looks like an awesome profiler and is fantastic that works in other platforms. I am super excited to see more cool features and https://pythonspeed.com/sciagraph/ https://pythonspeed.com/sciagraph/ looks fantastic :)
- pablogsal 4y agoOne thing to note is that we support tracking forked/child processes as well! >> - Much lower overhead, sounds like. I actually think Fil has the potential to be faster in some situations because it seems that aggregates the flamegraph in memory. Memray needs to do a ton of I/O to disk holding a lock so if is under very heavy pressure it will be a bit slow. Here are some non-very scientific tests running the "test_list" file from the CPython test suite: * With pymalloc active (not a lot of heavy presure): $ fil-profile run -m test test_list ... Total duration: 278 ms $ memray3.10 run -m test test_list ... Total duration: 128 ms * With pymalloc not active (heavy presure): $ PYTHONMALLOC=malloc fil-profile run -m test test_list ... Total duration: 278 ms $ PYTHONMALLOC=malloc memray3.10 run -m test test_list Total duration: 344 ms So as you can see Fil is 20% faster than memray in this scenario. This means that Fil is doing a fantastic job! We spent a lot of time optimizing memray and the fact that Fil can beat it is a testament to Fil's quality :)
- foota 4y ago>> but optimizing non-peak memory has no impact. You have to be able to identify the peak in order to reduce memory usage. In some sense this is only true if you're the end user of a platform, if you're trying to pack jobs onto machines then you actually do care about the utilization at any given time, since you can oversubscribe based on someone's max usage. E.g., you can give everyone a limit based on their peak memory, but then bin pack based on their actual usage (and evict when you're wrong)
- pid-1 4y agoUnrelated, but thank you for your awesome writings!