4 ms·
I loved the idea of Julia but R and specifically the tiddyverse https://www.tidyverse.org/ https://www.tidyverse.org/ Just makes everything else seem not as el
by baldfat 6y ago
I loved the idea of Julia but R and specifically the tiddyverse https://www.tidyverse.org/ https://www.tidyverse.org/ Just makes everything else seem not as elegant to my humble eyes.
- phillc73 6y agoHave you tried Query.jl or DataFramesMeta.jl?
- baldfat 6y agoI don't like working with DATA TABLES UNLESS it is a HUGE data frames. Then if it is huge I'll go towards sparks. I normally am working with under a million objects which with today's computers is not that big. Edit I meant to say that DATA TABLES library in R reminds me more of Query.jl then tiddyverse
- wodenokoto 6y agoWhat are you doing in the tidyverse that is not related to dataframes (or tibbles as the subclass of data.frame, that tidyverse uses is called)?
- phillc73 6y agoQuery.jl supports two different paradigms, one inspired directly by LINQ, the other by dplyr.[1] I actually prefer data.table in R, over dplyr, and Query.jl is really quite different. [1] http://www.queryverse.org/Query.jl/stable/ http://www.queryverse.org/Query.jl/stable/
- cwyers 6y agoVery much not the parent, but as a heavy R user, I don't think either of them quite nail the way dplyr and the tidyverse work. The thing about dplyr is... it's just functions. Okay, so, it's functions that leverage features R has (notably lazy evaluation and non-standard evaluation). But it's just functions. All you need is a function that takes a data frame and returns a data frame. So you can take a function out of the R standard library, you can take a function from a package written before dplyr came around, you can take a function from a recent non-tidyverse package, you can write your own function... it's all just functions. In DataFramesMeta.jl, though, you have a macro, and everything runs inside that macro. So if you want to take something that isn't a part of DataFramesMeta.jl... here's an example. Let's say you want to take the popular mtcars dataset, and get the five cars with the best gas milage. In dplyr, that goes mtcars %>% arrange(mpg) %>% head(5) arrange is a function from the dplyr package, head is a function from the standard library, but they both work seamlessly together. DataFramesMeta.jl lets you work in a pipe-forward fashion, but (at last I knew, at least, it's been a while since I played with it), you couldn't use the Julia head function within a DataFramesMeta.jl pipeline. You have to do your data transformations, assign to a variable, and then get the head of that variable. Which, okay, probably doesn't sound like a big deal. But I think it gets at the heart of what efforts to do something Tidyverse-like in other languages (Python and Julia, mostly) really miss. The key value proposition of the Tidyverse in R is that it is very composable and very extensible. That means, if you are trying to solve something in a Tidyverse way, you can probably find something that works for you. If you are doing financial analysis? Get tidyquant. If you're doing time series analysis, the tidyverts packages are for you. And it all works because there is so little friction involved in writing your own functions that extend the functionality of Tidyverse packages. Yes, dplyr is a useful querying DSL in its own right, but you can find a bunch of SQLish query languages, and they're all some degree of fine. Query.jl or DataFramesMeta.jl might expose a useful querying DSL for data frames, but they don't seem to me to be built to support building a whole ecosystem like dplyr and the Tidyverse are.
- phillc73 6y agoThat's a really good point that I'd not really thought about. I'd never really considered the difference between calling just functions versus macros. Thinking about Query.jl and DataFramesMeta.jl, and I am for sure not an expert in either, I can't specifically speak to your `head` example, but other base functions can be combined with macros. For example, see the LINQ examples from DataFramesMeta.jl[1] where `mean` is being used. Or again the LINQ style examples in Query.jl[2], where `descending` is used in the first example, or `length` later in the Grouping examples. Is that the kind of thing you meant? For whatever reason, with the way my brain is wired, the LINQ style of query just works for me. I have never directly used LINQ, but do have some SQL experience. In fact, I wrote some dinky little wrapper functions[3] around duckdb[4] so I could directly query R dataframes and datatables with SQL using that backend, rather than sqldf[5]. [1] https://juliadata.github.io/DataFramesMeta.jl/stable/#@linq-and-other-chaining-macros https://juliadata.github.io/DataFramesMeta.jl/stable/#@linq-... [2] https://www.queryverse.org/Query.jl/stable/linqquerycommands/ https://www.queryverse.org/Query.jl/stable/linqquerycommands... [3] https://github.com/phillc73/duckdf https://github.com/phillc73/duckdf [4] https://duckdb.org/ https://duckdb.org/ [5] https://cran.r-project.org/web/packages/sqldf/index.html https://cran.r-project.org/web/packages/sqldf/index.html
- diarrhea 6y ago> tiddyverse That is not what you meant to say.
- fishmaster 6y agoYou can use RCall to use R from Julia: https://github.com/JuliaInterop/RCall.jl https://github.com/JuliaInterop/RCall.jl
- ku-man 6y agoGood idea, and since I am on it I should use rpython to, in turn, call python from R. Julia provides awesome solutions.
- tfehring 6y agoI use R and the Tidyverse extensively. Exploratory data analysis is definitely clunkier in Julia - you need the `@pipe` macro to patch up some limitations in native pipes, there's no `dbplyr` equivalent that I know of, and despite Julia's better metaprogramming in general, the lack of built-in equivalents to scoped `select`/`mutate`/`summarize` is a real drag. But Julia's type system, substantially better date/time system and utilities, explicit vectorization with `.`, and the use of functions instead of scoped expressions in functions like filter are all real benefits over R. If you write a lot of Rcpp, Julia's performance without dropping down into a lower-level language is also a significant advantage. It's easy, bordering on trivial, to performantly implement a generic join (i.e. `join(f, df1, df2)`) in Julia; `dplyr` still doesn't have those at all, `data.table` only sort of does, and I believe the canonical R implementation (AFAIK) in the `fuzzyjoin` package requires holding the Cartesian product of the dataframes in memory, which is obviously not great.