4 ms·
How does everyone feel about a JSON dicts separated by newlines format? (Like the JSON format that mongoimport can accept.) Each sample of the type of data tha
by compare 13y ago
How does everyone feel about a JSON dicts separated by newlines format? (Like the JSON format that mongoimport can accept.)
Each sample of the type of data that I'm often dealing with tends to be nested in nature. Yes, I do have a script that can flatten out the nested dicts into a regular table, but that always results in a blowup into hundreds of columns.
Nice suggestion to share the raw data. I've never seen a researcher do that, I think many don't even save the raw data to disk before extracting what they want, but I always try to.
- edraferi 13y agoI've been using this format recently and appreciate the additional fidelity from simple rows. The flexibility makes it good for the raw stage, but I usually have to extract tidy subsets for real analytic work. I have written many little scripts to pull arbitrarily deep keys out of these structures and produce tidy tables for further analysis. actually, this format is also nice because iterating over the lines of a file is very similar to running through a mongo cursor. that makes it easy to reprise choose to work with both inputs.
- edraferi 13y agoI've been using this format recently and appreciate the additional fidelity from simple rows. The flexibility makes it good for the raw stage, but I usually have to extract tidy subsets for real analytic work. I have written many little scripts to pull arbitrarily deep keys out of these structures and produce tidy tables for further analysis. actually, this format is also nice because iterating over the lines of a file is very similar to running through a mongo cursor. that makes it easy to reprise choose to work with both inputs.
- pallandt 13y agoIt's rare (unfortunately), but Dr. Jeffrey Leek isn't the 1st proponent of sharing raw data, or otherwise a believer in reproducible results. Dr. Eamonn Keogh has made very important contributions in data mining and is also a huge advocate of reproducibility. See for example 'Why the lack of reproducibility is crippling research in data mining and what you can do about it' dated 2007 @ http://dl.acm.org/citation.cfm?id=1341922 http://dl.acm.org/citation.cfm?id=1341922