Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
arun_sriniv
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
arun_sriniv
12y ago
No worries :-). And glad to hear you're working on it! Let me know if I can be of any help.
2.
▲
by
arun_sriniv
12y ago
ajinkyakale, "harder to learn" doesn't expose the fact that data.table provides so many features that, for example, dplyr just doesn't. And in addition, it is fast and memory efficient. Rolling joins for example are slig
3.
▲
by
arun_sriniv
12y ago
Unfortunately the datasets in that benchmark less than 3MB each in size - it fits entirely in cache. It doesn't give a good indication of how well the function/implementation scales on bigger data sizes that really matter (in term
4.
▲
by
arun_sriniv
12y ago
@ajinkyakale, thanks. What'd be also interesting is to benchmark memory usage in addition to runtime.
5.
▲
by
arun_sriniv
12y ago
data.table's `DT[i, j, by]` is quite consistent actually and is comparable to SQL's - i = where, j = select | update and by = group by. This form is always intact. For example: require(data.table) DT = data.table(x=c(3:7),
6.
▲
by
arun_sriniv
12y ago
Here's a benchmark Matt recently did comparing data.table, dplyr and pandas on 50GB and 100GB: https://github.com/Rdatatable/data.table/wiki/Benchmarks-%3A...
7.
▲
by
arun_sriniv
12y ago
We're in the process of adding more detailed vignettes illustrating more clearly the philosophy behind data.table's `i, j, by`. Should make things lot easier for beginners - https://github.com/Rdatatable/data.