10 ms·
Also openrefine (formally google refine) http://openrefine.org/ is like a GUI version of csvkit. It can do external look ups, fuzzy matching, and has its o
by denimboy 12y ago
Also openrefine (formally google refine)
http://openrefine.org/
is like a GUI version of csvkit.
It can do external look ups, fuzzy matching, and has its own programming languages Jython and GREL.
- buckie 12y agoI tried out OpenRefine previously on the large CSV's I have to deal with at work largely on the recommendation of HN comments. On small stuff to medium size stuff, it's pretty nice. But for larger sets (+1GB) it starts to slow down and eventually will fail to commit changes when the set is big enough. The only answer I've found to consistently work quickly, with the ability to explore the data, is the killer combo of iPython Notebook + pandas' read_csv[0] + a lot of RAM -- 10GB CSV on disk becomes ~20GB in memory (don't know why yet). When I say quick, I mean 10GB CSV un-cached disk to RAM in <5Min including fuzzy parsing on dates. The nice part is, when you have things figured out, you can enable a chunked reading to get back in-core on machines of lesser specs. Further, you can dump the pandas DataFrame to HDF, thereafter having ludicrous-speed IO & 'where' queries. Still though, OpenRefine is much more turn key and feature rich. [0]http://pandas.pydata.org/pandas-docs/version/0.13.1/generated/pandas.io.parsers.read_csv.html http://pandas.pydata.org/pandas-docs/version/0.13.1/generate...
- codygman 12y agohttp://hackage.haskell.org/package/csv-conduit http://hackage.haskell.org/package/csv-conduit http://hackage.haskell.org/package/cassava http://hackage.haskell.org/package/cassava These libraries should be able to work with data that large, though I can't say whether they meet your requirements yet. I'm not sure what exactly "10GB CSV un-cached disk to RAM in <5Min including fuzzy parsing on dates". Namely, I don't know what you mean by date fuzzy parsing or what your output looks like after. Perhaps I need to open ipython notebook and import pandas ;)
- buckie 12y agoI've been working to get Haskell approved at my place of employment for 1.5 years but getting the US Gov't to change is rather hard; I've had to sit the "tech" people down and explain that javascript != java... with that baseline, explaining 'Why Haskell' is non-trivial. Come September, after 1.5 years of effort, I should have 'all the FOSS'. Until then I have to wait. It's worth noting that iPython Notebook + pandas vs. cassava + conduits (even with iHaskell Notebook) serve very different ends. If I need to explore how to do something, I'd use Haskell. But I'm still in phase 2 (phase 1: collect underpants) and I've yet to find anything as powerful and flexible as the iPython Notebook + pandas + hdf5 stack that also just works. I can just move faster with that stack than anything else I've ever seen. That being said, I'm knowingly deferring bugs to the runtime -- 'tis the cost of python. If you're unfamiliar with pandas, the "quick vignette" here[0] is decent enough. The reason pandas is awesome, IMO, isn't actually because pandas is awesome (which it is) but because it's embedded in a full language. Julia, R, etc... can do the same stuff (maybe faster), but I wouldn't also want to program, say, a production web-app in them (though I have high hopes for Julia). Fuzzy parsing on dates: pandas by default uses dateutil[1] which is both awesome and slow. 10GB CSV... : yeah... it's "fast for python" but pandas is admittedly doing a lot in that time, namely putting it into a data structure that is very friendly to time series analysis. [0] http://pandas.pydata.org/ http://pandas.pydata.org/ [1] https://labix.org/python-dateutil https://labix.org/python-dateutil
- codygman 12y agoWow! I wasn't expecting such a response! Awesome. I'll definitely have to check this out. About dateutil... actually I used it at work today. As for Haskell and fuzzy dates: http://hackage.haskell.org/package/dates http://hackage.haskell.org/package/dates I'm curious what you mean by cassava + conduits with Haskell notebook serve different ends. For instance where would you use it in place of pandas/python/hdf5?