3 ms·
The key requirement is that the sample be distributed through the time series, and allow for arbitrary start and end points in that time series. It isn't enough
by wyohaines 6y ago
The key requirement is that the sample be distributed through the time series, and allow for arbitrary start and end points in that time series. It isn't enough for it to just be random.
So while the UUID could be leveraged to provide a distribution space for random sampling, it doesn't help much with the requirement that the records be time distributed.
In the article I used a 2 week span as an example, but in actual use this could be a 1 hour span or a 4 month span, too, with an arbitrary number of data points being collected, and I need a technique that works efficiently for all of those things.
To the core of your thought, though, about using the UUID as a source of randomness for gathering a random sample, I think the challenge there is going to be indexing the UUID in a way that is amenable to that purpose.
- deleted 6y ago[deleted]