3 ms·
Disclaimer: I work at Meta. I don't work on Velox, but my work intersects with Velox in multiple ways. The short answer is consistent semantics. We have a larg
by scott_s 4y ago
Disclaimer: I work at Meta. I don't work on Velox, but my work intersects with Velox in multiple ways.
The short answer is consistent semantics. We have a large data warehouse, and several different engines that query that data warehouse. We want, as much as possible, consistent semantics across all of our engines for our users. That is, as much as possible, we want the same query to produce the same results on different surfaces. The value of that is so that users can draft a query on one surface, and use that same query elsewhere with confidence that they will get the same results. If we can consolidate on the same execution engine inside of our query engines, we can achieve that.
Minor quibble on terminology: "Velox knows how to retrieve data from your database". I would instead say that Velox knows how to retrieve data from your storage. Velox is deeply integrated into the query engine, and the combination of the query engine and the storage is "the database." In large data warehouses, we've already separated storage from compute to achieve scalability.
If this is all too abstract, think of this way: your query engine (such as Presto) is like a full computer system, while Velox is like the processor. Processors, by themselves, are not useful. They need to be attached to a motherboard which has RAM and connections to hard drives, GPUs and other external devices. Your query engine is like the computer system that contains that motherboard and all the components connected to it. There's enormous value in having multiple computer systems with different capabilities, but using the same kind of processor: you get consistent behavior when the capabilities are the same. Velox is that processor, ready to be plugged into different query engines.