3 ms·
Insightful article / ad. Maybe the problem with the idea is that in practice not many users want to do joins at arbitrary depths levels, e.g. "sort all the chil
by geuszb 8y ago
Insightful article / ad. Maybe the problem with the idea is that in practice not many users want to do joins at arbitrary depths levels, e.g. "sort all the children of US presidents by height" is probably not a very common query needing a massively distributed architecture?
- yorwba 8y agoI suspect that although any given query that needs a join is rare, there is a long tail of many different such queries. Just having more than one predicate is enough, e.g. "weather in <city> during <event>" or any of the queries in footnote 2 of the article.
- mrjn 8y agoConsider [sort all comedy movies by rating]. Such queries happen on a daily basis on movie sites like IMDB or Rotten Tomatoes. The only way to avoid joins is when you specialize your data to a particular vertical. Therefore, your flat tables are then built to serve say movie data, and can avoid some joins. But, if you're building something spanning multiple verticals, like Knowledge Graph is, which houses movie data, celebrity data, music data, events, weather, flights, etc. Then building flat tables specific to each vertical's properties is almost impossible.
- geuszb 8y agoYes, you're right about the need for many flat tables specific to each vertical. However, if other aspects of these verticals need to be specific to the vertical anyway (for example, UI: weather for the week is presented differently than a list of movies and their ratings; or: data quality), then there's still a significant amount of eng work per vertical, so that data flattening step may not be the bottleneck in growing the number of verticals served...
- mrjn 8y agoThat was the exact issue Google found itself in. Every OneBox had their own backend, which meant many different teams were involved in running and maintaining them. Now, with the graph serving system in place, all they need to do is to slap a new UI for the vertical, while the backend remains the same. Of course, there's real effort involved in building a new UI for the vertical, but it's a lot smaller compared to building a whole stack for each vertical. Not to mention, just the movie vertical itself has many "types" of data. Movies, Actors, Directors, Producers, Cinematographers, etc. -- all of these have different properties. By the time one is done flattening all of these into relational tables, they've built a custom graph solution -- which is what happens repeatedly in companies.
- geuszb 8y agoThat makes sense... Though after reading the article, I thought Google wasn't using a system that supported arbitrary join depths?
- pacala 8y ago"Travel from NY to LA over the weekend" was probably not a very common demand needing a massive investment in air travel infrastructure and equipment circa 1900. User demand is bounded by the capability of the tools commonly available.