3 ms·
In this case, the features we kept getting asked for by our customers necessitated a change in underlying database architecture. I talk about that quite a bit i
by pauldix 3y ago
In this case, the features we kept getting asked for by our customers necessitated a change in underlying database architecture. I talk about that quite a bit in the reddit thread.
I totally agree that a rewrite is risky. It's not something I'd choose to do again, but at the time we didn't really see any way around rewriting the bulk of the database (even if we kept it implemented in Go).
Using Rust and the Arrow ecosystem of projects (Parquet, DataFusion, Flight) meant that there were a ton of things we didn't have to do from scratch. One of our staff engineers, Andrew Lamb, has called it a toolkit for building databases. Thanks in part to his contributions, I think he's right.
- say_it_as_it_is 3y agoWhat is meant by separating compute from storage? This keeps being mentioned as if it were some new paradigm shift so I assume there's a non-obvious situation.
- pauldix 3y agoA traditional monolithic database assumes that you have locally attached storage. All of your ingest, indexing and query processing happens on the machine with that storage (i.e. your compute and storage live together). The cloud kind of complicated things with EBS and high IOPS network storage, but generally, those work kind of the same way. The volume is mounted on a single machine that uses it. When people talk about separating compute from storage, they mean pulling compute heavy tasks like query, ingest, indexing, and compaction apart and using a shared storage tier that many systems can talk to. Usually this is object storage paired with some sort of catalog (kept in either object storage or some other store or API). Snowflake popularized this approach in the data warehousing and OLAP space with great success. Their papers submitted to VLDB are great reads on the topoic.