24 ms·
loved the article. I strongly believe that when doing geospatial indices, it is quite important to up front establish the nature and distribution of your data.
by mohaps 12y ago
loved the article. I strongly believe that when doing geospatial indices, it is quite important to up front establish the nature and distribution of your data.
Searching for a generic/theoretically all-encompassing solution is quite akin to looking for the perfect pub-sub system for all workloads. :)
My comment was related to lat/lon indexed data clustered around cities. If you know beforehand that you're going to deal with such a dataset (static or dynamic), one can pragmatically decide on some sort of by-convention partitioning (e.g. Partitioning by continent/region which will bring it down to in-memory indices not needing continuous disk access).
I love the non-euclidean bit in the piece you wrote. Anyone who has tried to do a k-nearest neighbor query east of New Zealand will appreciate that bit. :D
- jandrewrogers 12y agoIf you have a priori knowledge of the data distribution, you can construct an ideal partitioning scheme. In practice, this runs into some significant problems: - The average distribution of the data and the instantaneous distribution of the data can be very, very different. This means that some cells are overloaded while others are idle, and the whole system runs as slow as the overloaded cell. The canonical (and fairly benign) example is data following the sun. - Many data sources have inherently unpredictable data distributions. - Spatial joins across different data sources (say, weather and social media) require congruent partitioning or it won't scale. Unrelated data sources tend have unrelated data distributions, so this is a problem. Making your partitions match your data distribution is good practice for static data layers. With some caveats, this can be modified for spatial joins across static data layers as well. For dynamic data sources, you run into issues with data and load skew at scale.
- gcb0 12y agoThe sun is pretty predictable... everyone who built any online service know the access pattern is a wave with time of day. do you mean to say you've seen people trying to shard a geodb by latitude instead of longitude? ... that would be a very sloppy initial research.
- gcb0 12y agoreally? i got downvoted because i find it hard to imagine that people don't know that their systems will get more requests for the area that is day-time? or i guess the sun is not predictable? sigh...