5 ms·
There's just a huge amount of waste in many cases which is very easy to fix. For example, if we have a list of fractions (0.0-1.0): * Python list of N Python f
by itamarst 4y ago
There's just a huge amount of waste in many cases which is very easy to fix. For example, if we have a list of fractions (0.0-1.0):
* Python list of N Python floats: 32×N bytes (approximate, the Python float is 24 bytes + 8-byte pointer for each item in the list)
* NumPy array of N double floats: 8×N bytes
* Hey, we don't need that much precision, let's use 32-bit floats in NumPy: 4×N
* Actually, values of 0-100 are good enough, let's just use uint8 in NumPy and divide by 100 if necessary to get the fraction: N bytes
And now we're down to 3% of original memory usage, and quite possibly with no meaningful impact on the application.
(See e.g. https://pythonspeed.com/articles/python-integers-memory/ https://pythonspeed.com/articles/python-integers-memory/ and https://pythonspeed.com/articles/pandas-reduce-memory-lossy/ https://pythonspeed.com/articles/pandas-reduce-memory-lossy/ for longer prose versions that approximate the above.)
- adamsmith143 4y agoOk now I have 100s of columns. I should do this for every single one in every single dataset I have?
- staticassertion 4y agoYes?
- itamarst 4y agoIt takes like 5 minutes, and once you are in the habit it's something you do automatically as you write the code and so it doesn't actually cost you extra time. Efficient representation should be something you build into your data model, it will save you time in the long run. (Also if you have 100s of columns you're hopefully already benefiting from something like NumPy or Arrow or whatever, so you're already doing better than you could be... )
- maerF0x0 4y ago> It takes like 5 minutes, and once you are in the habit it's something you do automatically as you write the code and so it doesn't actually cost you extra time. This is the argument I've been having my whole career with people who claim the better way is "too hard and too slow" . I'm like "gee, funny how the thing you do the most often you're fastest at... could it be that you'd be just as fast at a better thing if you did it more than never?" .
- dahfizz 4y agoHey, programmer time is expensive. It is our duty to always do the easiest, most wasteful thing. /s
- maerF0x0 4y agoFuture me's time is free to today me. :wink:
- kllrnohj 4y agoBut premature optimization is the root of all evil! I'm a better programmer for actively ignoring these optimizations! /s
- maerF0x0 4y agoBut if I change this code, I have to change them all! Good thing the status quo requires no evidence, but any change we want to propose? Impossibly high standards.
- thrwyoilarticle 4y agoAs an individual contributor you have an incentive to approach a problem in the way that teaches you the most for your career - then you can pretend it's the approach that's the best effort to risk ratio.
- eru 4y agoIt doesn't have to necessarily teach you the most. Eg at Google, if it gets you promoted, that's also good. Promotable projects at Google even have (or at least used to have) a complexity requirement. You can guess where the incentives lead.
- chaps 4y agoHah, I'd love to work with the datasets you work with if it takes five minutes to do this. Or maybe you're just suggesting it takes five minutes to write out "TEXT" for each column type? The data I work with is messy, from hand written notes, multiple sources, millions of rows, etc etc. A single point that's written as "one" instead of 1 makes your whole idea fall on its face.
- itamarst 4y agoFor pile-of-strings data, there are still things you can do. E.g. in Pandas, if there are a small number of different values, switch to categoricals (https://pythonspeed.com/articles/pandas-load-less-data/ https://pythonspeed.com/articles/pandas-load-less-data/ item 3). And there's a new column type for strings that uses less memory (https://pythonspeed.com/articles/pandas-string-dtype-memory/ https://pythonspeed.com/articles/pandas-string-dtype-memory/).
- chaps 4y agoTried that in the past, but it's really slow. Pandas is effectively removed from my workflows because of issues like this. But, I have workarounds for these issues by loading everything into postgres under TEXT columns in a "raw" schema, then do some typecast tests in a descending list of types to get the smallest possible type to transfer to a new table in a "prod" schema. It's read-only data, so it's not a big deal to run it once, and builds out a chain of changes from csv -> sql. Something like this could be done with pickling to avoid having to re-type every time I run the code (and I've done that for some past projects, but it's... ehhh).
- bee_rider 4y agoIs enough data generated from handwritten notes that the memory cost is a serious problem? I was under the impression that hundreds of books worth of text fit in a gigabyte.
- chaps 4y agoIt can be! I have a 50m row dataset of handwritten notes that's large, but when loaded into pandas, it blooms wayyyy larger than its underlying files.
- adamsmith143 4y ago5 minutes per column or 5 minutes per dataset? If per column then that is hopelessly slow. 500+ minutes per dataset and I may have dozens of datasets.
- deleted 4y ago[deleted]
- dvfjsdhgfv 4y agoYou'll need to decide on a case by case basis. Many datasets I work with are being generated by machines, come from network cards etc. - these are quite consistent. Occasionally I deal with datasets prepare by humans and these are mediocre at best, and in these cases I spend a lot of time cleaning them up. Once it's done, I can clearly see if there are some columns can be stored in a more efficient way, or not. If the dataset is large, I do it, because it gives me extra freedom if I can fit everything in RAM. If it's small, I don't bother, my time is more expensive than potential gains.
- nomel 4y agoAssuming your data is not ephemeral, and you have some way to ingest the data, from a full precision data store, why not? Store at full precision, process at fractional precision, a story as old as time.
- deckard1 4y agointeresting. Python doesn't use tagged pointers? I would think most dynamic languages would store immediate char/float/int in a single tagged 32-bit/64-bit word. That's some crazy overhead.
- acdha 4y agoThis has been talked about for years but I believe it's still complicated by C API compatibility. The most recent discussion I see is here: https://github.com/faster-cpython/ideas/discussions/138 https://github.com/faster-cpython/ideas/discussions/138 Victor Stinner's experiment showed some performance regressions, too: https://github.com/vstinner/cpython/pull/6#issuecomment-656135551 https://github.com/vstinner/cpython/pull/6#issuecomment-6561...
- nneonneo 4y agoAbsolutely everything in CPython is a PyObject, and that can’t be changed without breaking the C API. A PyObject contains (among other things) a type pointer, a reference count, and a data field; none of these things can be changed without (again) breaking the C API. There have definitely been attempts to modernize; the HPy project (https://hpyproject.org/ https://hpyproject.org/), for instance, moves towards a handle-oriented API that keeps implementation details private and thus enables certain optimizations.
- deleted 4y ago[deleted]
- justinlloyd 4y agoI love Python as a language, and all its packages, and have been using it since the late 90's, but Python's legacy decisions are one step away from causing the language to being found face down in a dirty ditch after an all night bender.
- eru 4y agoI do appreciate the backwards incompatible changes that Python 3 brought. Those were (mostly) good, and it was brave of them to go for it. I wish they'd had another go. Call it Python 4. But getting most people away from Python 2 to Python 3 already took a really long time.
- BLanen 4y agoYou're describing operations done on data in memory to save memory. That list of fractions still needs to be in memory at some point. And if you're batching, this whole discussion goes out of the window.
- rcoveson 4y agoWhy would the whole original dataset need to be in memory all at once to operate on it value-by-value and put it into an array?
- BLanen 4y agoIf the whole original dataset doesn't need to be in memory all at once, there isn't even an issue to begin with.
- saltcured 4y agoI think the point is that you can use a streaming IO approach to transcode or load data into the compact representation in memory, which is then used by whatever algorithm actually needs the in-memory access. You don't have to naively load the entire serialization from disk into memory. This is one reason projects like Twitter popularized serializations like json-stream in the past, to make it even easier to incrementally load a large file with basic software. Formats like TSV and CSV are also trivially easy to load with streaming IO. I think the mark of good data formats and libraries is that they allow for this. They should not force an in-memory all or nothing approach, even if applications may want to put all their data in memory. If for no other reason, the application developer should be allowed to commit most of the system RAM to their actual data, not the temporary buffers needed during the IO process. If I want to push a machine to its limits on some large data, I do not want to be limited to 1/2, 1/3 or worse of the machine size because some IO library developers have all read an article like this and think "my data fits in RAM"! It's not "your data" nor your RAM when you are writing a library. If a user's actual end data might just barely fit in RAM, it will certainly fail if the deep call-stack of typical data analysis tools is cavalier about allocating additional whole-dataset copies during some synchronous load step...
- forrestthewoods 4y agoWhat an unhelpful post. The realization that modern servers can easily persist multiple terabytes of data is profound. The fact that some datasets are just floats and you can quantize some floats from 32-bits down to 8-bits is true but not a helpful observation. I also don’t know where you get “Python float is 24 bytes + 8-byte pointer for each item in the list”. Wat.
- tomatotomato37 4y agoThere's still a difference between terabytes of data spread between RAM sticks multiple inches away from the CPU die and the megabytes of cache data close enough to the compute silicon to experience quantum effects. A modern CPU can burn through hundreds of instructions in the time it takes to resolve a thrashing cache
- forrestthewoods 4y agoYes and? Latency (hand wavey) L1: 1 ns L2: 2.5 ns L3: 10 ns RAM: 50 ns SSD: 50,000 ns / 50 us Those are very approximate and specifics vary. RAM is 5 to 50 times more latent than cache. SSD is ~1000 times slower than RAM. And a spinning HDD is I think ~100 times slower than an SSD. If your database fits in RAM you’re likely in a happy place. If your DB is so massive it needs to spill to disk you wind up with a mountain of complexity. Multiple machines, sharding, hot/cold etc. The point of the article is “modern servers have a lot of RAM and you might be able to delete a lot of complexity if you throw money at a server with 4 terabytes of RAM. This option is more practical than you might have realized!”
- minitech 4y ago> I also don’t know where you get “Python float is 24 bytes + 8-byte pointer for each item in the list”. Wat. Not sure how many ways there are to reword that. A CPython float takes 24 bytes of memory, and storing them in a list means 8 bytes per item for the pointer. So in CPython, a list of n floats takes 32n bytes of memory. >>> sys.getsizeof(1.0) 24 >>> l = list(map(float, range(62_500_000))) # memory use goes up by >2 GB >>> del l # memory use goes down by >2 GB (no need to go straight to NumPy to avoid this when relevant, though – array.array is built in.)
- ikiris 4y agoI mean, its python. If you're expecting it to be efficient, I don't really know what to tell you.
- justinlloyd 4y agoI was hired by SONY at one point to help optimize a piece of video editing software that the outsourced team had created and could not figure out how to make it a) go faster and b) use less RAM. This is back when 16GB was on the upper end of what workstation class machines could handle. The desktop application, written in an early 2004-ish version of C# and .NET regularly brought the workstations to their knees and thrashed memory and pegged the CPU when large 1080p images were being loaded in or moved around. Each RGBA image, stored as a PNG on the HDD, was loaded in, each RGBA 32-bit unsigned integer was unpacked into its attendant R, G, B, and alpha components into 32-bit unsigned integers themselves, which were then stored in an Int (boxed the native to a non-native type) which were then appended to a dynamically allocated at read time ArrayList. Deep copies were made of each image any time one of them was resized or manipulated, keeping the original untouched image in RAM in case it was needed. A copy of the image before transformation was stored on the Undo stack. And the 1080p working surface of the screen was super-sampled at 8x resolution to support defringing of images when layered. All stored as 8-bit RGBA components in boxed 32-bit integers in a dynamically allocated ArrayList.
- eckza 4y agoDon't leave us hanging - what did you _do?_ Did you fix this?
- Jorge1o1 4y agoNot the parent so I can’t tell you exactly what he did, but some suggestions: - The PNG specification contains plenty of information including width and length that allow you to use 2d arrays rather than ArrayLists - The RGBA 32bit integer doesn’t need to be unpacked.. you can just perform shifts and transformations of each of the 8bit channels. - If you do unpack, you should use native 8bit unsigned ints - You don’t need an unmodified copy.. that’s your file - Undo stack should contain deltas, not the whole image over and over again. Edit: formatting
- prewett 4y agoThe obvious thing to do would be to store the image in memory as just an array of 8-bit of native unsigned integers (RGBARGBARGBARGBA...) and see where that got you. I assume C# has the equivalent of a Java byte[], except that fortunately it is unsigned in C# instead of signed in Java. I expect that you would get quite a performance boost with just that.
- shapefrog 4y ago> which is very easy to fix go on amazon and buy another stick of RAM