10 ms·
Python consumes a lot of memory – how to reduce the size of objects?
- necovek 7y agoThis is a pretty neat and useful comparison of how much memory different structures use in Python to achieve roughly the same goals.
- mmrezaie 7y agoI am a C/C++ programmer mostly. Is there good documentation for python development when I care about performance and memory footprint out there as a book or something that anyone can recommend? It was fun reading this but I like to know more.
- BerislavLopac 7y agoThis article is a solid start: https://rushter.com/blog/numba-cython-python-optimization/ https://rushter.com/blog/numba-cython-python-optimization/
- myth_drannon 7y agoHigh Performance Python is a good book. https://www.oreilly.com/library/view/high-performance-python/9781449361747/ https://www.oreilly.com/library/view/high-performance-python...
- drej 7y agoHigh Performance Python by Ian Ozsvald and Micha Gorelick - fairly outdated at this point, but still quite relevant (I believe the authors are working on a new edition).
- BerislavLopac 7y agoThey have started work on the new edition if I'm not mistaken.
- agumonkey 7y agoNo draft online ?
- BerislavLopac 7y agoI don't think so. This is quoted from Ian's latest mailing list [0] posting: Second Edition of High Performance Python in the works I'm very pleased to say that Micha and I have started work on a Second Edition of High Performance Python with O'Reilly, planned for early 2020. This book will use Python 3.7 and will add tools that barely existed 4 years ago including Dask and probably some Tensorflow, amongst lots of other goodies. I'll let you know how the book progresses via this list. So far I've updated the Profiling chapter, added some advice on 'being a highly performant developer' and have rebuilt the Cython/Numba/PyPy code for the Compiling chapter. [0] https://ianozsvald.com/data-science-jobs/ https://ianozsvald.com/data-science-jobs/
- ausjke 7y agoit's written in 2014 so not that old I think
- sixplusone 7y agoI thought the usual mantra for performance-critical functions is to implement them in C and call from py.
- hermitdev 7y agoNot always necessary. I do a lot of ETL type work in python, and generally, the idea is to never pull everything into memory at once if you don't need to. This means leveraging generators quite a bit. Read a row, process a row, generate a row, pass it along the pipeline. I've written scripts that can process a 1k line csv file with the same memory footprint as a 10M line csv. If you have to read everything into memory, because you're doing some sort of transforms like a list to a dict or such, explicitly deleting variables can help, as well, rather than waiting for them to go out of scope, but this should be pretty rare.
- pjmlp 7y agoOr just port to Julia, Scheme or Common Lisp, while enjoying the power of dynamic languages with compilation to native code.
- njharman 7y agoI'm a Python "snob", going on 24 years and If I ever "I care about performance and memory footprint" I'd not use Python. Python is good enough 90%. You get faster 2x 10x, etc code by picking better algorythms or solutions. In Python, you should not be caring about 10% or 30% speed improvement. It's not worth it, It's not Python's strength. When you need faster you go to C based libs (the python libs that need to go fast are already in C), https://en.wikipedia.org/wiki/Cython https://en.wikipedia.org/wiki/Cython https://en.wikipedia.org/wiki/Numba https://en.wikipedia.org/wiki/Numba etc.
- jerven 7y agoOr go for [1]pypy or [2]graalpython. I know that equivalent code in graal java ee produces equivalent assembly/performance to C+current gcc, which really shows that yes changing to a JIT can pay off. The python in graalvm is early stages but it shows that the old lesson of just write the hot spot in C is no longer true. Chris Seaton shows that for Ruby/C the language Ruby interpreter/C switch is expensive. I think the same is true for Python interpreter/C switches. Something that GraalVM+Python|Ruby can optimize out. [1] https://www.pypy.org https://www.pypy.org [2] https://www.graalvm.org/docs/reference-manual/languages/python/ https://www.graalvm.org/docs/reference-manual/languages/pyth...
- jashmatthews 7y agoLuaJIT has great performance through the C FFI. It’s more that traditional JITs can’t optimize across FFI boundaries. Hopefully Chris will correct me if I’m full of shit. There are other approaches which can work too, like compiling C to LLVM IR using Clang so a function can be inlined by an LLVM based JIT at runtime.
- edflsafoiewq 7y agoLike with JS, you might not have a choice. I mostly use Python for extending programs that use it as a embedded scripting language. Performance is still important.
- zerkten 7y agoSee Mike Mueller's tutorials from PyCon at https://www.youtube.com/watch?v=DGrS0uwMuHY https://www.youtube.com/watch?v=DGrS0uwMuHY. That's the 2018 one, but I'm guessing there is a similar one. There always seems to be a high perf tutorial of some kind.
- kbirkeland 7y agoI feel like the last two are cheating a bit by explicitly using 32 bit integers where the other examples seemed to use 64 bit.
- auscompgeek 7y agoNo, the fields that take up 8 bytes are pointers to PyObject. (I guess this article assumes a 64-bit memory model.)
- wil421 7y ago>A significant reduction in the size of a class instance in RAM is achieved by eliminating __dict__ and__weakref__. This is possible with the help of a "trick" with __slots__: What is the downside to the __slots__: trick? In what cases do you need the dict and weakref?
- weberc2 7y agoIf you are writing code that adds or removes properties at runtime. I.e., when you want to write shoddy code.
- SEJeff 7y ago"when you want to write shoddy code" - This so much!
- hermitdev 7y agoI have written code like that in the past, and would consider doing so again in the future. When, it's appropriate, it can be hugely beneficial. When I last did it, I was wrapping a C++ API that I needed to compare against another dataset for a merge. The C++ API (which, admittedly, I also wrote), didn't have equality or hash operators defined on the Python objects. So, I monkey-patched them in for the keys I needed. It was actually the most elegant solution I could come up with as I could then naturally use the objects in sets and easily get differences to update the appropriate dataset. As an aside, when I wrote the python wrapper around my C++ API, I purposely didn't define equals and hash operators for the objects, despite having full information to do so on the natural keys, because I wanted the flexibility to do the monkey-patching and override how the objects were compared depending upon circumstance.
- drej 7y agoThe article doesn't mention a hidden gem in Python's standard library - typed arrays! Those are in package `array` and they are a very barebones version of numpy's ndarray - with just one (explicit) dimension and no overloaded operators. But if you just want to keep a bunch of numbers in a contiguous array, they can save you tons of memory. (I know the purpose of the article is to describe a more complex data structure, but arrays can still get you very far.)
- vbezhenar 7y agoThat's how you deal with that problem with Java as well: arrays of primitives. If you need to operate on million of points, you can either use Point[] which will incur something like 16 MB additional memory or use two int[] arrays which won't incur any extra overhead short of few bytes. Your code won't be pretty, but it'll be fast and you can always hide this weirdness behind pretty API.
- 0815test 7y ago> That's how you deal with that problem with Java as well: arrays of primitives. It just goes to prove that you can indeed write FORTRAN in any language!
- llukas 7y agoThis is actually a feature - python array slicing syntax is very similar to fortran.
- cesarb 7y agoAlso known as the "Struct of Arrays" approach.
- wtallis 7y ago"Struct of Arrays", to be contrasted with "Array of Structs", because in Java it's actually "Array of pointers to Struct". In languages where "Array of Structs" is actually possible, the decision of which to use is less clear-cut and depends on how big the Struct is, what the access patterns are like, and whether you're trying to perform SIMD operations on multiple Structs at once.
- miohtama 7y agoMany workloads do not require more than 4GB addressable RAM per process. Linux offered a 32-bit user space with 64-bit instruction set: https://en.m.wikipedia.org/wiki/X32_ABI https://en.m.wikipedia.org/wiki/X32_ABI As effectively many Python workloads are usually object oriented business logic and objects are mostly pointers, setting up x32 user space "halved" the memory usage. It also made execution performance faster because of better CPU cache utilisation. Sadly, x32 was "very custom" and very hard to support. Last I heard x32 is being phased out from Linux kernel.
- miohtama 7y agoLooks like this little gem was in the Wikipedia article > The best results during testing were with the 181.mcf SPEC CPU 2000 benchmark, in which the x32 ABI version was 40% faster than the x86-64 version.
- chrisseaton 7y agoAnother option is to use the standard 64 bit ABI, but store object pointers compressed into 32 bits when on the heap. This lets you address perhaps 32 GB in a 32 bit value.
- srean 7y agoCould you elaborate more on how the 'compress' part works. Quite curious. I can imagine working with a base pointer and 32 bit offsets.
- miohtama 7y agoAlso curious about this. Does 64-bit instruction set provide some segments or functionality form this? How about "native" pointers coming from glib and such? If there has to be base + offset translation on every pointer access it is way too slow. I would also assume JavaScript VMs in browsers would be already utilising this, as web page workloads are not gigabytes (hopefully).
- deleted 7y ago
- erdewit 7y agoIn the pure Python cases the 8 bytes per attribute are just pointers. The x, y and z are themselves full-blown objects with all the extra memory overhead that comes with it and this is not counted in the article. For example, an int object uses 28 bytes, so three of them already use up more than each of the described container objects. The Cython and Numpy cases directly store the actual data and this has the larger effect to reduce memory.
- spott 7y agoOn the other hand, the first 256 integers are singletons in python, so they aren't duplicated.
- hermitdev 7y ago0-255 (inclusive) are singletons, just to clarify.
- mswtk 7y agoSlightly pedantic correction: This is a performance optimization in CPython. I wouldn't be surprised if other implementations have something similar, but to my best knowledge, this behaviour isn't part of the standard.
- merlincorey 7y ago> ... but to my best knowledge, this behaviour isn't part of the standard. To the best of my knowledge, Python isn't one such language with a standard, is it?
- omnimkar69 7y agothis functions u will also see in the case of JAVA s well.this problem arises due to large no. of objects are active in RAM durig the execution of a progrm especially if there ae restrictions on the total amount of availablle memory
- kazinator 7y agoTXR Lisp, 64 bit: 1> (defstruct blank ()) #<struct-type blank> 2> (pprof (new blank)) malloc bytes: 16 gc heap bytes: 32 total: 48 milliseconds: 0 #S(blank) 32 bit: 2> (pprof (new blank)) malloc bytes: 8 gc heap bytes: 16 total: 24 milliseconds: 0 #S(blank) The structure instance has a pointer to its type, followed by a numeric ID (which is also in the type, but is forwarded to the instance for faster access). The ID is combined with a slot symbol to perform a cache lookup to get the offset of a slot. The numeric ID is a fixnum, which leaves a few spare bits for a couple of flags: struct struct_inst { struct struct_type *type; cnum id : sizeof (cnum) * CHAR_BIT - TAG_SHIFT; unsigned lazy : 1; unsigned dirty : 1; val slot[1]; }; If someone wanted to shrink this, they could patch the code so that the dirty flag support, and lazy instantiation of structs is made optional (as in compiled out), and so is the forwarding of the inst->type->id to inst->id. This struct type is not known outside of struct.c, which is only some 1700 lines long. If you take out the id member from struct_inst, the C compiler will find all the places that have to be fixed up; literally a 15 minute job. I can also think of a more substantial refactoring that would eliminate the type pointer also. There is no room in the heap object handle to store it directly: heap handles have four words in them; the COBJ ones used for structs have the type field, a class symbol, a pointer to an operations structure, and a pointer to some associated object (in this case struct_inst). All structures share the same operations structure. However, if we dynamically allocate that operations structure for each struct type, we could stick the type pointer in there, taking it out of the instance. Thus an instance size could literally just be sizeof(pointer) * no-of-instance-slots.
- Deimorz 7y agoIt's not always an option, but simply using PyPy can massively reduce memory usage without needing to change your code at all. The PyPy site links to this blog post (from 10 years ago) with some info: https://morepypy.blogspot.com/2009/10/gc-improvements.html https://morepypy.blogspot.com/2009/10/gc-improvements.html And a quick search found this relatively recent post that does some measurement of it: https://dev.nextthought.com/blog/2018/08/cpython-vs-pypy-memory-usage.html https://dev.nextthought.com/blog/2018/08/cpython-vs-pypy-mem...
- intellimath 7y agoIn the article consider aproach with the `dataobject` from `recordclass` library as a base object in class definition. This seem can produce less memory than PyPy.
- alkonaut 7y agoAny type that you have "a lot" of, is a bad candidate for a class/object. Your OO program should typically never instantiate thousands of heap allocated objects at once. A triangle mesh is a good object candidate. One triangle or vertex is not. A device, a render system API handle and an image is a good candidate, one pixel is not, and so on. So don't use a lot of objects. In C# you use a struct and arrays of them to avoid creating an object on the heap per array entry. In java you have to resort to SoA, in Python it's the same. Just because you have objects doesn't mean everything should be an object. Even languages that follow an "everything is an object"-design usually have an escape hatch such as plain value arrays.
- lacampbell 7y agoIt's not about using objects or not, it's how you compose them. A C# array of integers is an object. Each integer in that array is also an object (though granted, it's a special case). Numpy arrays and javascript typed arrays as well.
- alkonaut 7y agoPrimitives and value types are not objects in C#. The array is one object, that’s it.
- lacampbell 7y agoIt has methods and fields, it implements interfaces, and descends from the 'Object' type. The fact that it's unboxed is an implementation detail. https://docs.microsoft.com/en-us/dotnet/api/system.int32?view=netframework-4.8 https://docs.microsoft.com/en-us/dotnet/api/system.int32?vie...
- alkonaut 7y agoWhat’s relevant for performance is whether an array of 1000 “things” require 1 or 1001 allocated objects, whether accessing thing N in the array requires dereferencing one or two pointers, and whether the item has a storage of 32 bits without overhead for headers. Which ones to call “objects” is semantics. For efficient access, the array must be a consecutive array of primitives. This is the case in both C# and java for integers. In C# it’s also the case for an array of Vector2 with 2 primitives each, which isn’t the case in java. My point is this: avoid heap allocating many things in collections. They must be raw (primitive, consecutive) data without per instance overhead and of course without heap alloc/GC cost.
- skykooler 7y agoSeems to be an error in editing: > ...which received a rating of [stackoverflow] (https://stackoverflow.com/questions/29290359/existence-of-mutable-named-tuple-in https://stackoverflow.com/questions/29290359/existence-of-mu... -python / 29419745). In additioobjects liken, it can be used to reduce the size of objects in RAM...
- syn0byte 7y agoWhile everyone haggles about the internals I have an interesting anecdote about the costs of tools libraries. A small service that landed in my lap needed to read(only) a data source that was roughly 10k lines of yaml. No way that was going to be in any way efficient so I asked for suggestions. All the work-a-day devs(I am not) instantly said the same thing without a single real thought about it: Make it a database duh! Long story slightly less long, Loading up the libraries to interface with a database ate between 2 and 3 times the memory(depending on the DB and lib) that simply loading the entire 10k line yml ate and offered slower performance and required more code. SQLite was pretty darn close but in the interest of saving developers from themselves vis a vis parameterized queries, or the need to queries all together for that matter, increased the required code for zero benefit. The service still hums along with a 10k line yaml in memory. "Worse is better" indeed.
- novok 7y agoYou could of improved it with a protobuf or a json file then, but with the size of your data set, it shouldn't really matter what your using. It can be hard to beat an in memory data structure when your data set is small enough, true.
- pm 7y agoThe quickest change is no change, naturally. However, "better" is dependent on the situation. If the data is unlikely to change, stay the course. If not, then while a database might be more maintainable over the long term, even if not as efficient.
- rleigh 7y agoWell, the SQLite might have been more scalable if you needed to deal with much larger YAML files. I've been in a similar situation with XML. It's fine to hold in memory until you end up running out of memory, at which point you need to look at other approaches. There's definitely an overhead to changing, but it might be worth paying if the tradeoffs make sense.
- mlthoughts2018 7y agoI always hate how these things are phrased. The idea of “using a lot of memory” doesn’t exist in a vacuum. “A lot” relative to what? Do you require run-time dynamic typing features? Do you want to leverage the Python data model? If yes, then this is the memory cost of your desires. Advice in many of the comments about other languages sounds so tone deaf. The question is not about using less memory. It’s about getting _exactly_ the feature set of Python while using less memory.
- worik 7y agoI am paying US$5 a month for python's memory hogging. Running Mailman it will not run in 1GB of ram (the US$5 VPS) so I had to give it 2GB. 2GB? For a mail list server? Had the same problem running motioneye. It will not run on a Raspberry PI Zero. It is advertised to run on a Zero (Apparently I could squeeze it on by doing some Python magic....) What total crap is Python! What waste, what hubris, what technical failure! I suspect the problem is actually Django - what hubris! NIH! lighttpd is running sweetly, with some rust templating in a acceptable fraction of the memory.
- ggm 7y agoBut moving up your stack of desire, what alternate to mailman would you live with? Five a month is $60 a year. Assuming you value your own time at professional wages, this job is worth an hour of your time at most before the opportunity cost of complaint is cheaper than replacement.
- ip26 7y agoopportunity cost of complaint No, what he actually bought for his $60/year was license to complain about mailman & python :)
- ggm 7y agoThe python complaint licence fee was the best deal I ever made.
- worik 7y agoYou have completely missed the point. The US$5/month is not a hardship. What is a hardship is taking what used to be done in 40k Perl and require 2G. It is a affront! What is a waste is writing a poor quality greedy web server when Apache/Lighttpd/Nginx all exist and are much much better. What really cost me was the two weeks I spent trying to make motioneye (python) and motioneyeos (not) work as advertised. As advertised! It is not only that the python is too bloated that stops it working, it does but there are other reasons. Very clearly no body has really tested it. It is a symptom of python culture. Shoddy and arrogant. I feel terribly burnt by it. Twice. So my stack of desire in this domain is quite small and simple: Never to hear again from python! A very personal POV, be happy in your python hacking. I saw the opportunity to have a public rant and I took it!
- ggm 7y agoIs this trading space for speed or is there actually a speedup as well in some cases? With interpreters it's possible you be small and fast if you become perhaps obscure or more rigid in your structured definitions
- floatingatoll 7y agoOP, if you're reading this, your article has damaged syntax at the word `additioobjects`.
- intellimath 7y agoThe author was fixed this.