8 ms·
Dataframes – Julia, R, Python
- smu3l 12y agoIn Julia, you can create symbols with e.g. :user_id as in Ruby, which looks a lot nicer than symbol("user_id") and doesn't require mapping over an array of strings.
- ajinkyakale 12y agoOfcourse you can! I dont know why I didnt think of that
- Fede_V 12y agoThere was a very interesting design discussion by JMW on the julia-dev forum about nullable arrays, and column dtypes: https://groups.google.com/forum/#!topic/julia-dev/hS1DAUciv3M https://groups.google.com/forum/#!topic/julia-dev/hS1DAUciv3... Even so - right now Pandas is miles ahead of the Julia equivalent. Pandas was the brain child of Wes McKinney - an amazing coder, who really, really cared about speed (who recently also made a lot of money selling his start up to Cloudera - good for him!). The things you can do in Pandas with multi-index selects, joining dataframes on multiple axis, etc, are outright incredible.
- eggy 12y agoI am learning J for about a year now, and I will try out the examples in it to see how it goes. Memory-mapped files, and quick array operations. It is APL based, and J has been around since the 80s, and is open source. Wes McKinney seems to think it is a good way to go: https://twitter.com/wesmckinn/status/341317411607293953 https://twitter.com/wesmckinn/status/341317411607293953
- baldfat 12y agoGood showing of what are three good data languages. Strange I was a Python guy for a long time. Pandas just looks strange to me now since I switched to R two years ago. Seems like I need to dive into Julia again. Haven't for over a year.
- deleted 12y ago[deleted]
- peatmoss 12y agoFor dataframe-like operations, I've started wondering why more languages don't take the dplyr approach and simply default to using something like SQLite under the hood. Granted, I are no super data genius, but every time I start cracking a little into the internals of a dataframe implementation, I get the sinking feeling that SQL databases have already done the hard work of indexes and efficient data structures.
- pfh 12y agoYou often want to: * add or remove columns from a frame * perform vector operations on columns in a frame This would be awkward in SQLite. However, maybe a Column Oriented DBMS exists that would be workable.
- hadley 12y agoI thought that before writing dplyr, but now I see that there a big differences. Relational databases are designed to work with large datasets on disk, and to accept changes very rapidly. The demands for in memory data analytics are quite differnt. Columnar data stores are a better fit, but it's pretty easy to bang out efficient code for in memory data; it's much harder to work with out of memory data.
- peatmoss 12y ago> large datasets on disk I saw this benchmark a while back comparing Pandas to SQLite in-memory databases. While Pandas did edge out SQLite in several areas, it was by well under an order of magnitude: http://wesmckinney.com/blog/?p=414 http://wesmckinney.com/blog/?p=414 Pretty solid performance plus the ability to work with large datasets on disk seemed like a pretty big win to me. I could imagine a set of SQLite extensions (a la spatialite) that could further optimize for various data.frame use cases. As an added bonus, the same libraries would be very portable between different languages--even languages that don't currently have something like dataframes. EDIT: What I don't know about is memory efficiency. Perhaps SQLite isn't, but I'd not bet against?
- IndianAstronaut 12y ago
- IndianAstronaut 12y agoInteresting. Although after spwnding a lot of time on data frames, I have grown to like the csv parsers in postgres where I can do a lot of the same things as data frames, but with clean sql instead of the sometimes odd data frame syntax. R also stands out because it is so easy to run a wide variety of statistical methods easily on a data frame.
- mapcar 12y agoIt's amazing to think that R (or S) had data frames since the 70s and only now are other languages implementing them. There are some quirks of course, and pandas introduced some convenient features. But the R community has also provided its own improvements in the way of data.table, and now, dplyr.
- ajinkyakale 12y agoData.table package by matt dowle definitely deserves a mention! Its fast and I like the indexing functonalities it provides. The benchmark timings are pretty impressive.
- arun_sriniv 12y ago@ajinkyakale, thanks. What'd be also interesting is to benchmark memory usage in addition to runtime.
- ajinkyakale 12y agoI should have mentioned you (arun_sriniv) as the co-developer of data.table! Thanks for all the hard work. And yes, memory usage will be interesting as that is the bottleneck when it comes to large dataset. I am working on something on those lines. Will post something soon :)
- arun_sriniv 12y agoNo worries :-). And glad to hear you're working on it! Let me know if I can be of any help.
- sebastianavina 12y agoI've seen a lot about Julia on the last months, it seems like a good language (performance and kind of a nice syntax), For me, what makes R a very good choice is because of RStudio. Being able to play there with your data and save it all for later is one of the biggest reasons to use RLang. The Python equivalent would be emacs org-mode, which is great, but not as graphical as RStudio. Julia seems like a good language, maybe someday i will jump on it, but for the evil mind out there planning to write another language. Please stop, we already have great languages! I can't keep up with the learning! and is so damn difficult to even start a project with so many choices!
- alecdbrooks 12y agoIf you're looking for a nice graphical way to play with data in Python, may I suggest IPython Notebook [0]? It's not always easy to configure, but it's maturing fast and lets you have Python code, Markdown, and graphs in one place, not to mention the 21 other languages available natively or as add-ons[1]. [0]: http://ipython.org/notebook.html http://ipython.org/notebook.html. [1]: https://github.com/ipython/ipython/wiki/IPython%20kernels%20for%20other%20languages https://github.com/ipython/ipython/wiki/IPython%20kernels%20... (I'm counting the two Perl kernels as one language and not counting Calico or the example kernel.)
- cschmidt 12y agoThere is a version of iPython that works with Julia too
- kyllo 12y agoYeah you can run IPython Notebook with a Julia kernel.
- sebastianavina 12y agoI just checked ipython notebook. Looks like Matlab Notebook, but in the browser instead of MS Word (I don't know if Matlab Notebook still exists, it's been a long time since the last time I used Windows), anyway, you should try org mode on emacs for python, is way more versatile, compiles to latex, html, and has almost all the features of the ipython notebook.... still Rstudio has a lot of graphical incentives, like data edition, environment variables, easy to follow documentation. Both org-mode and RStudio are very powerful tools, I was mostly ranting about how many options we have for making almost the same thing.
- jzwinck 12y agoThe example which claims to get "all the rows from 50th row to the 55th row" is broken, since Python is zero-based whereas Julia and R are one-based. The 50th row in Python is at index 49, so the code is not equivalent between the examples.
- kldavenport 12y agoGlad they added .query() to Pandas in .13. In general I find the methods in Pandas/NumPy much more consistent with general programming constructs than most of what I see in R. No doubt Hadley has bolted on a lot of great functionality to R, but the Rcpp dependency/GPL license is a turn off.
- hadley 12y agoWhat's wrong with Rcpp? And what's wrong with GPL? Unless you're planning on distributing your code, you shouldn't even need to think about it.
- urlwolf 12y agoAnyone doing R comparisons should use data.table instead of data.frame. More so for benchmarks. data.table is the best data structure/query language I have found in my career. It's leading the way in The R world, and in my way, in all the data-focused languages.
- hadley 12y agoData tables are extremely fast but I think their concision makes it harder to learn and code that uses it is harder to read after you've written it. It's very reminiscent of APL.
- arun_sriniv 12y agodata.table's `DT[i, j, by]` is quite consistent actually and is comparable to SQL's - i = where, j = select | update and by = group by. This form is always intact. For example: require(data.table) DT = data.table(x=c(3:7), y=1:5, z=c(1,2,1,1,2)) DT[x >= 5, mean(y), by=z] ## calculates mean of y while grouped by z on ## rows where x >= 5 DT[x >= 5, y := cumsum(y), by=z] ## updates y in-place with it's cumulative sum ## while grouped by z on rows where x >= 5 "Harder to read after you've written it" and "harder to learn" are all very subjective and pointless. One could make very similar observations about `dplyr`, but I'll refrain from it here. I implore the readers to take a look at over 100+ reviews on crantastic: http://crantastic.org/packages/data-table http://crantastic.org/packages/data-table from users of the package. Keeping `i`, `j` and `by` operations together allows optimising for speed and more importantly memory usage (altogether under a consistent syntax) - which are two very important aspects especially working on really huge data sets (10-100GB in RAM or more). Here's a detailed benchmark (only on grouping so far) on 10 million (in MB) to 2 billion rows (100GB): https://github.com/Rdatatable/data.table/wiki/Benchmarks-%3A-Grouping https://github.com/Rdatatable/data.table/wiki/Benchmarks-%3A...
- ajinkyakale 12y agoI agree to what Hadley said in some ways. It takes a bit more time to get used to the [i, j, by] notation and I personally feels its unlike most of the R syntax. But I dont see that stopping me from using something as fast as data.table.