6 ms·
It takes like 5 minutes, and once you are in the habit it's something you do automatically as you write the code and so it doesn't actually cost you extra time.
by itamarst 4y ago
It takes like 5 minutes, and once you are in the habit it's something you do automatically as you write the code and so it doesn't actually cost you extra time.
Efficient representation should be something you build into your data model, it will save you time in the long run.
(Also if you have 100s of columns you're hopefully already benefiting from something like NumPy or Arrow or whatever, so you're already doing better than you could be... )
- maerF0x0 4y ago> It takes like 5 minutes, and once you are in the habit it's something you do automatically as you write the code and so it doesn't actually cost you extra time. This is the argument I've been having my whole career with people who claim the better way is "too hard and too slow" . I'm like "gee, funny how the thing you do the most often you're fastest at... could it be that you'd be just as fast at a better thing if you did it more than never?" .
- dahfizz 4y agoHey, programmer time is expensive. It is our duty to always do the easiest, most wasteful thing. /s
- maerF0x0 4y agoFuture me's time is free to today me. :wink:
- kllrnohj 4y agoBut premature optimization is the root of all evil! I'm a better programmer for actively ignoring these optimizations! /s
- maerF0x0 4y agoBut if I change this code, I have to change them all! Good thing the status quo requires no evidence, but any change we want to propose? Impossibly high standards.
- thrwyoilarticle 4y agoAs an individual contributor you have an incentive to approach a problem in the way that teaches you the most for your career - then you can pretend it's the approach that's the best effort to risk ratio.
- eru 4y agoIt doesn't have to necessarily teach you the most. Eg at Google, if it gets you promoted, that's also good. Promotable projects at Google even have (or at least used to have) a complexity requirement. You can guess where the incentives lead.
- chaps 4y agoHah, I'd love to work with the datasets you work with if it takes five minutes to do this. Or maybe you're just suggesting it takes five minutes to write out "TEXT" for each column type? The data I work with is messy, from hand written notes, multiple sources, millions of rows, etc etc. A single point that's written as "one" instead of 1 makes your whole idea fall on its face.
- itamarst 4y agoFor pile-of-strings data, there are still things you can do. E.g. in Pandas, if there are a small number of different values, switch to categoricals (https://pythonspeed.com/articles/pandas-load-less-data/ https://pythonspeed.com/articles/pandas-load-less-data/ item 3). And there's a new column type for strings that uses less memory (https://pythonspeed.com/articles/pandas-string-dtype-memory/ https://pythonspeed.com/articles/pandas-string-dtype-memory/).
- chaps 4y agoTried that in the past, but it's really slow. Pandas is effectively removed from my workflows because of issues like this. But, I have workarounds for these issues by loading everything into postgres under TEXT columns in a "raw" schema, then do some typecast tests in a descending list of types to get the smallest possible type to transfer to a new table in a "prod" schema. It's read-only data, so it's not a big deal to run it once, and builds out a chain of changes from csv -> sql. Something like this could be done with pickling to avoid having to re-type every time I run the code (and I've done that for some past projects, but it's... ehhh).
- bee_rider 4y agoIs enough data generated from handwritten notes that the memory cost is a serious problem? I was under the impression that hundreds of books worth of text fit in a gigabyte.
- chaps 4y agoIt can be! I have a 50m row dataset of handwritten notes that's large, but when loaded into pandas, it blooms wayyyy larger than its underlying files.
- adamsmith143 4y ago5 minutes per column or 5 minutes per dataset? If per column then that is hopelessly slow. 500+ minutes per dataset and I may have dozens of datasets.