3 ms·
I kind of feel like pandas and ipython are anti patterns here. They seem super convenient but I've found that writing a lower level analysis with independent s
by micro_cam 12y ago
I kind of feel like pandas and ipython are anti patterns here.
They seem super convenient but I've found that writing a lower level analysis with independent scripts linked via make or similar saves massive amounts of time in the long run.
IE the first few steps should retrieve and process the data until it is just arrays of numbers (or whatever your actual analysis needs) that can be handled with efficient numpy code.
You can use pandas for this but a real database works too. After this step pandas becomes irrelevant and just leads to lots of unnecessary allocations from recasting data on the fly.
The problem with ipython is that it leaks like a sieve and you end up with all sorts of copies of the data, worker processes and no longer used variables which will eat all of your ram even on a big cloud instance.
It is much nicer to have each analysis script start with a clean stack and release its memory when done. Plus you can use non python utilities like grep, vowaple rabbit etc as intermediate steps.
I've found practices like these significantly lower the memory requirements of analysis and allow one to tackle bigger datasets on single machines.
Each to their own though.
- numlocked 12y agoI find that IPython NB is a great tool for initial exploration -- to sketch out the ETL steps that are needed, what types of computations I'm going to be doing, etc. As soon as those ideas are roughed out, I switch to writing Python scripts that I can run through completely, write tests for, git commit and get reasonable diffs, and more. I'll generally keep an ipynb open in a tab for experimentation, but I'm doing all of the work in PyCharm or Spyder. This also forces me to think about the engineering implications of the analysis a bit sooner. Inevitably these projects are not one-offs, and will need to be at least repeated regularly, if not outright productionalized -- and ipynb files do not lead to production-ready code. I also agree that using a real database is often a better option than pandas. Many folks avoid it because it's less comfortable to spin up postgres + a new DB instance than it is to sit in the comfort of IPython, but it's totally worth it. Exploring and previewing the data via SQL is just so much faster and more intuitive than spinning it around with pandas.
- jimbokun 12y agoAnd note seccess's comment about using in memory SQLite db to eliminate the need to configure postgres: https://news.ycombinator.com/item?id=8930488 https://news.ycombinator.com/item?id=8930488
- bsg75 12y agoWhy anti-patterns? Pandas makes a great companion to a database. Aggregate / reduce large datasets (>RAM) efficiently in a DB, then pivot in Pandas, and perform non-SQL-support analyses in Python. Same can be said for R. iPython is just a shell, a very good one, with options to share work for presentation and reproducibility.
- micro_cam 12y agoBecause they drive memory usage up a ridiculous amount making people think they need big data solutions when they don't. They are great for some things but I think the idea they should always be used is unfortunately prevalent making them anti patterns.
- bsg75 12y ago> the idea they should always be used is unfortunately prevalent making them anti patterns Yes, this is often too true - also to include Hadoop, NoSQL, ...