Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pauldix
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
91.
▲
Future of InfluxDB OSS: More open, still permissive, complementary closed bits
(influxdata.com)
7 points
by
pauldix
6y ago
|
0 comments
92.
▲
InfluxData (InfluxDB) is hiring Rust and columnar database programmers (remote)
(boards.greenhouse.io)
1 points
by
pauldix
6y ago
93.
▲
by
pauldix
6y ago
There is absolutely no point benchmarking it at this stage. One of the many reasons we're not bothering to produce builds right now. We'll let you know when it's time to even have a look as I'm sure you'll want to c
94.
▲
by
pauldix
6y ago
Thanks :)
95.
▲
by
pauldix
6y ago
Not confrontational at all. In short, it's taken me years to be able to do this. Last year we finally got our engineering organization to a point where I no longer have any direct reports. Instead, everyone rolls up to our VP of engine
96.
▲
by
pauldix
6y ago
This blog post has more detail and screenshots. It'll probably be more interesting for this crowd: https://www.influxdata.com/blog/influxdb-2-0-open-source-is-...
97.
▲
by
pauldix
6y ago
For now the execution is in-memory only. Over time we want to be able to execute against Parquet files on disk. However, for large scale analytics on huge data sets where you're scanning all of it, we'll likely push you to EMR or
98.
▲
by
pauldix
6y ago
I'm not too familiar with their stuff, but I think in terms of approach, they're very different. This project doesn't aim to product materialized views, which I think is more of what naiad is for? I'm not sure about diff
99.
▲
by
pauldix
6y ago
Yes, absolutely. There are the Arrow libraries in Rust that we're already contributing to. DataFusion, for example is an SQL execution engine. We'll be publishing crates for parsing InfluxDB Line Protocol, Reading InfluxDB TSM fil
100.
▲
by
pauldix
6y ago
So if my users can't have DDL access, that means that they can't define the schemas for the analytics that they want to do? It only works if as the developer of an application I have a fixed schema that my users interact with?
101.
▲
by
pauldix
6y ago
This comparison is vs. InfluxDB. This thread is about a new project called InfluxDB IOx. It's under development and we're not producing builds, so any kind of operational comparison would be very premature. Its architecture is dra
102.
▲
by
pauldix
6y ago
My post is about InfluxDB IOx, which is the project this thread is about. You're correct about InfluxDB having HA and clustering under a closed source enterprise license. If you read the post, I even mention this as a shortcoming of th
103.
▲
by
pauldix
6y ago
It's under a community license, which has restrictions. The limitations on derivative works and value added products or services are the ones that will create the most problem for people trying to build a business on it: https:/&
104.
▲
by
pauldix
6y ago
Timescale is built on top of Postgres, which is a row oriented database. They've built a kind of columnar layer on top of it, which is quite interesting. Because it's Postgres you get their full SQL support. Meanwhile, InfluxDB IO
105.
▲
by
pauldix
6y ago
Right now it's just a project in development, so nothing yet. But it's supposed to be for time series data. This could be metrics, events, or any kind of semi-structured data that fits into tables where you want to ask questions a
106.
▲
by
pauldix
6y ago
Ah yes, of course, hi Nevi :). Thank you again for all your work on the Rust implementation. We're obviously big fans. Gandiva bindings is definitely something we should look into, but I'm guessing there's much lower hanging
107.
▲
by
pauldix
6y ago
We'll be producing builds early next year. Those won't be anything we're recommending for production. Our goal is to have an early alpha in our own cloud environment by the end of Q1. I stress alpha. But we'll also have
108.
▲
by
pauldix
6y ago
So far working with Rust has been great. We haven't had to build any special build tooling. The build times are greater than Go, but we're able to keep things manageable by breaking the project into separate Crates (all still with
109.
▲
by
pauldix
6y ago
It absolutely is and we've been contributing back. A even bigger amount of this is based on Wes McKinney's work on Arrow. Andy is great and he's been helpful as we've been working with DataFusion.
110.
▲
by
pauldix
6y ago
They're heavily into Arrow. A few years ago they contributed Gandiva, an LLVM expression compiler for super fast processing. https://arrow.apache.org/blog/2018/12/05/gandiva-donation/ It's
111.
▲
by
pauldix
6y ago
I think the best new projects are created by a small focused team. Adding too many people too early actually slows things down. But, of course, I'm biased. The thing about this getting to production next year is that we're doing i
112.
▲
by
pauldix
6y ago
Users of InfluxDB 1.x can upgrade to 2.0 today. It's an in-place upgrade and just requires a quick command to make it work. Further, InfluxDB 2.0 also has the API for InfluxDB 1.x. We've been putting out InfluxDB 1.x releases whil
113.
▲
by
pauldix
6y ago
Yeah, Parquet is awesome. One of the things we really want to do here is to push DataFusion (the Rust based SQL execution engine) to work on Parquet files, but to push down predicates and other things and operate on the data while it's
114.
▲
by
pauldix
6y ago
Execution is vectorized, but that's Arrow really. We'd like this to be useful for general OLAP workloads, but the focus for the next year is definitely going to be on our bread and butter time series stuff. That being said, Arrow
115.
▲
by
pauldix
6y ago
Yes, this will likely be the case, but it's not a specific goal. We're focusing our efforts right now on the InfluxDB time series use cases we see most. More general data warehousing work isn't our focus, but we expect to pic
116.
▲
by
pauldix
6y ago
Those two systems are designed to work with Prometheus style metrics, which are very specific. You have metrics, labels, float values, and millisecond epochs. I'm not totally sure how they index things, but I would guess that it's
117.
▲
by
pauldix
6y ago
InfluxDB creator and lead for InfluxDB IOx here. Happy to answer any questions people might have. In short, it's an in-memory columnar database with object storage for persistence. It's a bunch of other things as well, but that&#x
118.
▲
by
pauldix
6y ago
HN thread for that is here: https://news.ycombinator.com/item?id=25049253
119.
▲
by
pauldix
6y ago
InfluxDB creator here. I've actually been working on this one myself and am excited to answer any questions. The project is InfluxDB IOx (short for iron oxide, pronounced eye-ox). Among other things, it's an in-memory columnar dat
120.
▲
LakeFS – atomic, versioned data lake on object storage
(lakefs.io)
8 points
by
pauldix
6y ago
|
0 comments
More ›