Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wesm
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
wesm
7y ago
Too big of a discussion for Hacker News! Come on dev@arrow.apache.org if you want to talk about it
32.
▲
by
wesm
7y ago
Dask does not have "experimental Arrow integration". It supports using Arrow to read Parquet files but no Arrow-based computational functionality.
33.
▲
by
wesm
8y ago
I'm looking at the code I linked, and you are serializing in the general case, it is not zero copy. Unpacking a bitmap is not free.
34.
▲
by
wesm
8y ago
Note: Vaex has its own memory model. If you input Arrow, it converts to the Vaex data representation. Details here: https://github.com/vaexio/vaex/blob/master/packages/vaex-arr... One of the primary
35.
▲
by
wesm
8y ago
RAPIDS is partly powered by Apache Arrow. So we are all collaborating on a common next-generation computation ecosystem.
36.
▲
by
wesm
8y ago
In general we are talking about O(n) algorithms, and the gains are due to better CPU cache utilization and fewer instructions per value, which LLVM helps do
37.
▲
by
wesm
8y ago
Arrow Flight is still being developed and may be while before it drops to real user code. See ARROW-249. Indeed having a open standard runtime memory format for tabular / columnar data sets is key to improved performance
38.
▲
by
wesm
8y ago
I'm not the main maintainer anymore (Jeff Reback has that honor), as I've moved on to work on the "greater pandas ecosystem" (of which my work on Apache Arrow is a part of this long development arc)
39.
▲
by
wesm
9y ago
I've been working on innovating core computational and IO infrastructure for pandas (and projects like pandas) -- much of this work has been happening in other codebases. See: http://wesmckinney.com/blog/apache-ar
40.
▲
by
wesm
9y ago
It's true -- as Jeff (lead core dev/maintainer the last several years) often says "Wes gets the kudos, I get the hate-mail"
41.
▲
by
wesm
9y ago
Hadley and I are also friends and collaborators! I think we're going to see a lot more interesting collaborations between the R and Python communities in the future, since at the end of the day we're solving a lot of the same prob
42.
▲
by
wesm
9y ago
Posted here: http://wesmckinney.com/blog/arrow-columnar-abadi/
43.
▲
by
wesm
9y ago
Apache Arrow is not competing with Apache Parquet or Apache ORC.
44.
▲
by
wesm
9y ago
My objective, in fact, is to see the data science world unify data frame representations around Apache Arrow, so code written for Python, Julia, R, C++, etc. will all be portable across programming languages. See https://www.yout
45.
▲
by
wesm
9y ago
One of the lead Arrow developers here ( https://github.com/wesm ). It's a little bit disappointing for me to see the Arrow project scrutinized through the one-dimensional lens of columnar storage for database systems --
46.
▲
by
wesm
9y ago
I'm sorry to be pedantic but can you add "(incubating)" to the title of the post?
47.
▲
by
wesm
9y ago
I see Weld as an embeddable component in systems like pandas, not a replacement.
48.
▲
by
wesm
9y ago
The first talk about Blaze was in November 2012. It was marketed as many things over the years, including "pandas for Big Data". My understanding is that Anaconda (fka Continuum Analytics) is no longer working on it.
49.
▲
by
wesm
9y ago
Right, but I don't think it's fair to say that R's support for categorical data is "excellent" if only strings can be category labels/levels. Categories (aka dictionary-encoded data in other systems) semantical
50.
▲
by
wesm
9y ago
I disagree with your premise that production systems require API stability in all thirdparty dependencies.
51.
▲
by
wesm
9y ago
Disagree. Factors (categorical data) in R only support strings
52.
▲
by
wesm
9y ago
I recommend my JupyterCon keynote to explain how it fits into the picture: https://www.youtube.com/watch?v=wdmf1msbtVs Data processing systems need runtime memory formats. Arrow is an efficient one for analytical data proce
53.
▲
by
wesm
9y ago
Yikes! I disagree with you -- these are assertions made on no evidence. This software is suitable for production systems as long as you are OK with occasional API changes / deprecations as the software evolves. We are not far off from
54.
▲
by
wesm
9y ago
Wes here. From what I understand, pandas is the middleware layer powering about 90% (if not more) of analytical applications in Python. It is what people use for data ingest, data prep, and feature engineering for machine learning models. T
55.
▲
by
wesm
9y ago
we're due to add a "Powered By" page to the website, the list of users is growing pretty quickly (e.g. Spark has Arrow support now)
56.
▲
by
wesm
9y ago
Hi, Wes here. R data frames have limited to no query planning and are not multithreaded in general, so Problems 10 and 11 are problems in R also.
57.
▲
by
wesm
9y ago
> Would somebody have some benchmarks against pandas for some standard operations ? pandas creator here. Numba is a complementary technology to pandas, so you can and should use them together. It is designed for use with NumPy arrays and
58.
▲
Software patents are evil, but BSD+Patents is probably not the solution
(wesmckinney.com)
1 points
by
wesm
9y ago
|
0 comments
59.
▲
by
wesm
9y ago
I see they are supporting Apache Arrow ( http://arrow.apache.org/ ) in their result set converter, which will be nice for interoperability with other data system that use Arrow: https://github.com/mapd/ma
60.
▲
by
wesm
10y ago
I was also a bit offended by the premise that TV (or video games) has any causal relationship with achievement. I think it's the other way around, some people are not ambitious and spend their idle time watching TV. But watching TV d
More ›