3 ms·
(Author here, sorry I didn't catch this earlier) The 3.5 minutes is the preprocessing time (think of it as index-building time). After the index is built, the
by cscheid 11y ago
(Author here, sorry I didn't catch this earlier) The 3.5 minutes is the preprocessing time (think of it as index-building time).
After the index is built, the goal is to be able to generate outputs at a time bounded by (up to a poly-log loss) the size of your screen, and not of your data.
After the index is built you get sub-millisecond answers for the 2-D histograms over reasonably-sized screens, and that turns out to be pretty useful for the use cases we're going for.
- CyberDildonics 11y agoI was including the indexing time too, I think 10 million points in 4 seconds on one core was what I saw. These were single high dimensional points, I'm guessing that the data being dealt with here is different? The link in your post doesn't work but I did look at the paper and couldn't really see the uniqueness. I wasn't clear on what 'data cubes' was supposed to mean. When you say 'screen' you are talking about ranges or bounds of area right?
- cscheid 11y agoWe use "data cubes", as per Jim Gray's paper, http://web.stanford.edu/class/cs345d-01/rl/olap.pdf http://web.stanford.edu/class/cs345d-01/rl/olap.pdf. When I say "screen", I mean (e.g.) your laptop's screen resolution. If you're going to display a heatmap on a screen with p pixels, we (roughly) touch only O(p) memory cells on query time, independently of the dataset size. When you say "you don't see the uniqueness", I'm not sure what you're comparing it against, so I can't say anything more. You mentioned gkd-trees earlier: is that the comparison you mean? In that case, these are two completely different data structures. For example, (at least as described on the 2009 siggraph paper), you can't subset on a categorical dimension of the pixels. In the case of the data structure we created, we can report (for example) a heatmap of all geolocated tweets generated by an iPhone, or all geolocated tweets generated by a windows phone, or all geolocated tweets irrespective of device, without having to scan the 200M tweet database.