Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
andygrove
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
31.
▲
by
andygrove
6y ago
Thrift is a serialization format. Arrow is a memory format.
32.
▲
by
andygrove
6y ago
You should be able to find this info in the specification: https://arrow.apache.org/docs/format/Columnar.html
33.
▲
by
andygrove
6y ago
The Arrow project contains implementations in multiple languages. Some of these languages contain code that can evaluate expressions against Arrow data, or even execute full queries. The C++ and Rust implementations contain query capabiliti
34.
▲
Cyclon Distributed Data Framework Using Arrow
(cylondata.org)
3 points
by
andygrove
6y ago
|
0 comments
35.
▲
Apache Spark GPU Accelerator from Nvidia
(github.com)
2 points
by
andygrove
6y ago
|
0 comments
36.
▲
Making Spark Fly: Nvidia GPU Acceleration for Apache Spark
(blogs.nvidia.com)
2 points
by
andygrove
6y ago
|
0 comments
37.
▲
Rapids Accelerator for Apache Spark
(github.com)
2 points
by
andygrove
6y ago
|
0 comments
38.
▲
by
andygrove
6y ago
I think it makes sense for software that is intended to run in Docker and frameworks like Kubernetes that use Docker.
39.
▲
by
andygrove
6y ago
Thanks! That really does seem to be the issue and I wouldn't have known about this, had I not asked. I will try this out and will update the blog post in ~8 hours time.
40.
▲
by
andygrove
6y ago
I get that a lot!
41.
▲
Why does musl make my Rust code so slow?
(andygrove.io)
154 points
by
andygrove
6y ago
|
72 comments
42.
▲
Rust Bindings for Apache Spark
(github.com)
3 points
by
andygrove
6y ago
|
0 comments
43.
▲
Rust Bindings for Apache Spark
(github.com)
1 points
by
andygrove
6y ago
|
0 comments
44.
▲
Rust DataBase Connectivity (RDBC)
(andygrove.io)
167 points
by
andygrove
7y ago
|
23 comments
45.
▲
Rust DataBase Connectivity (RDBC)
(github.com)
1 points
by
andygrove
7y ago
|
0 comments
46.
▲
Announcing Rust DataBase Connectivity (RDBC)
(github.com)
1 points
by
andygrove
7y ago
|
0 comments
47.
▲
by
andygrove
7y ago
Each executor within the cluster would use Arrow in-memory design. If you have enough cores and memory on a single node then potentially you wouldn't need a cluster.
48.
▲
by
andygrove
7y ago
I will write up some guidance in the next few days for those looking to contribute!
49.
▲
by
andygrove
7y ago
This is exactly the situation I am in. We have workloads that I think we could deliver with much lower TCO using Rust/DataFusion/Ballista. We could probably even use a single node for some of our workflows just using DataFusion di
50.
▲
by
andygrove
7y ago
Yes contributors are welcome. I will write up some guidance in the next few days for those looking to contribute!
51.
▲
by
andygrove
7y ago
These are great questions and topics I plan on addressing in a future blog post. SQL is a great convenience for simple analytical queries and it can be nice to be able to mix and match SQL and other access patterns (this is one thing I like
52.
▲
by
andygrove
7y ago
Some good points. Some incredible engineering has gone into Spark to work around the fact that it runs on the JVM. Memory overhead of Spark particularly (not just JVM) is very high. In some cases close to 100x more memory than equivalent qu
53.
▲
by
andygrove
7y ago
That's kinda what I'm trying to do here. Ballista is definitely very much inspired by Apache Spark but not a direct port. Spark has some design choices that were very much driven from being a JVM-first platform e.g. the way lambda
54.
▲
by
andygrove
7y ago
I've only been using Kubernetes for a couple months so far and am still learning, but I am very impressed so far. I love the way it facilitates dev and devops collaborating and the fact that it is cloud agnostic (I can even run a Kuber
55.
▲
by
andygrove
7y ago
Some great points there. I have multiple goals here. Goal #1 is to demonstrate how Rust can be effective at data science / analytics and evangelize a bit. Goal #2 is to have fun trying to build some of this myself and improve my skills
56.
▲
by
andygrove
7y ago
I agree. Rust has a very steep learning curve compared to JVM languages. My hope in building this platform is that it can provide value to other languages (especially JVM) by taking query plans and executing them efficiently.
57.
▲
by
andygrove
7y ago
Benchmarks are rarely fair. They are often biased in favor of those producing the benchmarks, whether intentionally or not. Whether you want to explore this more or not is your choice. Nobody is asking you to engage in this conversation. If
58.
▲
by
andygrove
7y ago
Thanks eklavya but it's completely normal on HN to have trolls criticizing other people's work. Best thing is to just ignore them. They can move on to criticize other posts. Much easier than contributing.
59.
▲
by
andygrove
7y ago
I'm sensing that you are not familiar with Rust and the safety that it introduces compared to C/C++.
60.
▲
by
andygrove
7y ago
This is a personal open source project. I'm not sure I'm "marketing" it since I make zero dollars from this work. For benchmarks, check out the past 18 months of posts on my blog. Here is the most recent: https:/&#
More ›