3 ms·
I like this write up a lot, particularly the examples and the metrics, but I wonder, what would be the relative performance of sampling using some of the middle
by techbio 6y ago
I like this write up a lot, particularly the examples and the metrics, but I wonder, what would be the relative performance of sampling using some of the middle digits of the UUID?
- wyohaines 6y agoThe key requirement is that the sample be distributed through the time series, and allow for arbitrary start and end points in that time series. It isn't enough for it to just be random. So while the UUID could be leveraged to provide a distribution space for random sampling, it doesn't help much with the requirement that the records be time distributed. In the article I used a 2 week span as an example, but in actual use this could be a 1 hour span or a 4 month span, too, with an arbitrary number of data points being collected, and I need a technique that works efficiently for all of those things. To the core of your thought, though, about using the UUID as a source of randomness for gathering a random sample, I think the challenge there is going to be indexing the UUID in a way that is amenable to that purpose.
- deleted 6y ago[deleted]