4 ms·
Thanks. We did study this article a while back while building our data infrastructure. We wrote about our infrastructure here: engineering.viki.com/blog/2014/da
by huy 12y ago
Thanks. We did study this article a while back while building our data infrastructure. We wrote about our infrastructure here: engineering.viki.com/blog/2014/data-warehouse-and-analytics-infrastructure-at-viki/
In fact, to give you a little bit more background on the problem, if you look at our data infrastructure diagram (in the posted link), the thing we're trying to improve is the hydration system (where it takes in a record in real-time and try to inject more time-sensitive information into it).
E.g. When a user watches a video (thus a video_play event sent), we want to know if it's a free user or a paid user. Since the user could be a free user today and upgrade to paid tomorrow, the only way to correctly attribute the play event to free/paid bucket is to inject that status right right into the message when it's received.
Building the system this way (using the hydration service) makes our service very prone to error and indeterministic (since you only have a short window to hydrate the message, and you can't replay a hydration).
That's why we're looking at building a historical lookup service that remembers all the different changes of a data object over time, so that replaying a hydration becomes deterministic.
At the moment we're processing around 100M records a day and growing. Not a lot but still at some scale that puts us in the position to think about scalability and performance.