3 ms·
For pile-of-strings data, there are still things you can do. E.g. in Pandas, if there are a small number of different values, switch to categoricals (https://py
by itamarst 4y ago
For pile-of-strings data, there are still things you can do. E.g. in Pandas, if there are a small number of different values, switch to categoricals (https://pythonspeed.com/articles/pandas-load-less-data/ https://pythonspeed.com/articles/pandas-load-less-data/ item 3). And there's a new column type for strings that uses less memory (https://pythonspeed.com/articles/pandas-string-dtype-memory/ https://pythonspeed.com/articles/pandas-string-dtype-memory/).
- chaps 4y agoTried that in the past, but it's really slow. Pandas is effectively removed from my workflows because of issues like this. But, I have workarounds for these issues by loading everything into postgres under TEXT columns in a "raw" schema, then do some typecast tests in a descending list of types to get the smallest possible type to transfer to a new table in a "prod" schema. It's read-only data, so it's not a big deal to run it once, and builds out a chain of changes from csv -> sql. Something like this could be done with pickling to avoid having to re-type every time I run the code (and I've done that for some past projects, but it's... ehhh).