4 ms·
given the sheer number and scope of Apache projects, it would be nice to have the title mention what it is (eg, "Apache Arrow - in memory analytics - 1.0.0") es
by neurobashing 6y ago
given the sheer number and scope of Apache projects, it would be nice to have the title mention what it is (eg, "Apache Arrow - in memory analytics - 1.0.0") esp when the linked page doesn't directly say.
- j45 6y agoI find the same, but normally most Apache projects have an overview link at the top of the project.
- wenc 6y agoNormally I would agree but Apache Arrow already has significant name-recognition to people who are likely to use it -- data engineers/scientists etc. It's a library that provides an in-memory columnar format, and supplies a data engine currently used in Apache Spark, Pandas, and libraries like turbodbc, which helps these tools achieve high performance on operations on tabular data. Having a single high-performance in-memory format means different programs can read/write from the same source without serializing/copying/deserializing. For instance, if you wanted to pass a huge table of data from R to Java to Python (because your tools span different languages), normally you'd have to copy and serde (Protobuf? JSON?) to pass the data, which equals huge overheads. With Arrow, each of those languages can directly interact with the same copy of in-memory data, in-process -- with the highest possible performance. You also get the performance of columnar databases without implementing your own columnar data structure. But of course, no harm adding a short description to the title to broaden its audience. Arrow is truly something amazing and the more people know about it the better.[1] Folks who program against traditional databases might not know about it, and I think they should, especially if they need to generate analytics (i.e. fast filtering/aggregation for dashboarding or for data pipeline tasks). [1] Overview: https://arrow.apache.org/overview/ https://arrow.apache.org/overview/
- ghshephard 6y agoI spend 8+ hours a day working with/ingesting data. Had no clue. Checked with my data-using colleagues just to see if any of them had heard of Arrow - about 50% had. These are people who spend their entire day working/ingesting data in various pipelines. So, while I won't comment on whether the subtitle would be useful, I'm pretty confident that the (perhaps not large) majority of people who would/could use Apache Arrow had not hear of it. A lot of people just stay in their lane. Having a solid description, like the one you provided, would be super useful. Perhaps a tool-tip feature of HN, even.
- irontinkerer 6y agoAnd it has a broad-level of support through it's libraries! "Libraries are available for C, C++, C#, Go, Java, JavaScript, MATLAB, Python, R, Ruby, and Rust."
- casion 6y agoI'm a potential user and had no clue. Not everyone keeps up with trends. In fact, if my coworkers (educators) are any indication, few people do. Even 2 sentences on this page would've helped a lot.
- tgb 6y agoThis is just a mismatch with how HN posts pages and how projects expect their release pages to be read. Release pages are for users to know what changed. But HN likes to link to original sources and not articles about the source. In cases like this, we end up pretending the release page is meant to be read by a wider audience than the writers ever intended it for - not Apache's fault! Apache Arrow's homepage does have that one-sentence description on it ('A cross-language development platform for in-memory analytics').
- m0dest 6y agoA blog post announcing version 1.0 is a good opportunity for a one-liner to explain what the heck your project is :) Just in case your audience happens to be expanding instead of contracting...
- buryat 6y agowhere is it used in production?
- axegon_ 6y agoApache Arrow? First thing the top of my head(given that I had to work with this these past few days), the python snowflake connector depends on it if you intend to fetch the results from snowflake as a pandas or dask dataframe. In other words - lots of places.
- wenc 6y agoHere's a list (and as mentioned in my comment, Apache Spark, Pandas, turbodbc) https://arrow.apache.org/powered_by/ https://arrow.apache.org/powered_by/
- fotta 6y agoSpark uses it under the hood to interop between PySpark and JVM
- agacera 6y agoDremio is a really nice product built using arrow. https://www.dremio.com https://www.dremio.com
- vertexclique 6y ago(contributor here) At my company our proprietary database layout is based on Arrow. Our workload is analytical and we are using it for read mostly workloads.
- aeontech 6y agoThis is a really nice summary, thank you!
- uberman 6y agoJust to be clear, Arrow is not an "in memory analytics" solution but rather an "in memory columnar data store" that might be useful when doing analytics.
- wesm 6y agoThis isn't accurate -- there are multiple query engine subprojects within Apache Arrow.
- epistasis 6y agoIs there some documentation for this on the Arrow website somewhere? I've been looking for info on the "compute engine" that's mentioned in this 1.0 announcement but haven't found much. In general, where's the best place to learn more about Arrow? I've approached it several times, and can find a lot about how to integrate it into other products, but none of the tools like the query engines that I would find very useful.
- uberman 6y agoI am eternally indebted to you for Pandas. Many thanks for that. Are you talking about there being support for multiple language libraries like PyArrow or about there being multiple Apache projects that utilize Arrow like Parquet and Spark? If not, I'm not following what sub-projects you are speaking about. As far as I know, Arrow is principally the Arrow Columnar Format and Arrow Flight with some other potentially interesting interfaces for compute kernels and CUDA devices. Am I missing something?
- andygrove 6y agoThe Arrow project contains implementations in multiple languages. Some of these languages contain code that can evaluate expressions against Arrow data, or even execute full queries. The C++ and Rust implementations contain query capabilities, and the Java implementation contains the Gandiva library that can delegate to C++ via JNI to evalulate expressions, for example.