3 ms·
A reply from the man himself! Thanks for the link. I'll have a go. I do like the look of the dplyr library a lot. Combining functions like select and group_by
by ldp01 10y ago
A reply from the man himself! Thanks for the link. I'll have a go.
I do like the look of the dplyr library a lot. Combining functions like select and group_by with the pipe operators creates code that is reminiscent of SQL- very nice for readability.
- pedrosorio 10y agoYou're going to love Spark if you haven't tried it yet.
- ldp01 10y agoAnother powerful tool I will keep in mind! I think this thread illustrates a kind of tension between those coming from an IT/big-data/web oriented background and the more traditional statistics/science/engineering side. The IT side bring a lot of very powerful and scalable tools to the table. However there are aspects of traditional work which I suspect are lost on some big-data people. For example, in my line of work (physical asset mgmt) we deal with a lot of very small datasets, very poor quality datasets (e.g. some guy's favourite spreadsheet) and also cultural issues (some engineers are inherently averse to changing systems, and spending decisions are inherently political). In this situation, there is a limit to the benefit of more powerful/scalable tools, and it is advantageous to use tools which are considered high quality and vetted by the community. R is in a good position here as it has the pedigree of being accepted by the academic stats community, as well as actually being a great tool.