Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ritchie46
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
61.
▲
by
ritchie46
5y ago
> (as usual project is called select...) Yeah.. this confusion is in the API as well (you can pass projection to IO readers). we used `select` because SQL. In the logical plan we make the correct distinction between selection and projec
62.
▲
by
ritchie46
5y ago
The benchmarks are hosted by H2oAI, not by the polars team. Vaex is not in that benchmark. I don't believe Vaex would be faster though. They aim at larger than RAM data processing, not maximum in-memory performance like we do.
63.
▲
by
ritchie46
5y ago
The embarrassingly parallel is aimed at the expression API. This allows one to write multiple expressions, and all of them get executed parallel. (So embarrassingly, meaning they don't have to communicate and use locks).
64.
▲
by
ritchie46
5y ago
Note that the compile times of julia are not included in the benchmarks. If you read the website, you'd seen that the grapsh show the first (excluding the compilation) and the second run (with hot cache). Also in the second run, julia
65.
▲
by
ritchie46
6y ago
Thanks, I feel so too. There is still a hope work to do. I hope that I can also bridge the gap regarding utility and documentation.
66.
▲
by
ritchie46
6y ago
There is definitely a gap, and I don't think that Arrow tries to fill that. But I don't think that its wrong to have multiple implementations doing the same thing right? We have PostgresQL vs MySQL, both seem valid choices to me.
67.
▲
by
ritchie46
6y ago
Hi Author here, Polars is not an alternative to PyArrow. Polars merely uses arrow as its in-memory representation of data. Similar to how pandas uses numpy. Arrow provides the efficient data structures and some compute kernels, like a SUM,
68.
▲
by
ritchie46
6y ago
Author here. These examples and docker-composes files are heavily outdated. Please take a look at the docs for up to date examples. P.S. I do what I can to keep things up to date, but only have the time I have.
69.
▲
by
ritchie46
6y ago
Right.. So testing, testing, testing it is.
70.
▲
Ask HN: Which hash function for join algorithms
2 points
by
ritchie46
6y ago
|
2 comments
71.
▲
Polars: Rust DataFrames Based on Apache Arrow
(github.com)
1 points
by
ritchie46
6y ago
|
1 comments
72.
▲
by
ritchie46
6y ago
As a hobby project I tried to build a DataFrame library in Rust. I got excited about the Apache Arrow project and wondered if this would succeed. After two months of development it is faster than pandas for groupby's and left and inner