6 ms·
Two main differences - ability to time travel for training data generation and the ability to push compute to the write side of the view rather than the read si
by nikhilsimha 2y ago
Two main differences - ability to time travel for training data generation and the ability to push compute to the write side of the view rather than the read side for low latency feature serving.
- ShamelessC 2y ago> ability to time travel for training data generation What now?
- nikhilsimha 2y agoPardon the jargon. But it is a necessary addition to the vocabulary. To evaluate if a feature is valuable, you could attach the value of the feature to past inferences and retrain a new model to check for improvement in performance. But this “attach”-ing needs the feature value to be as of the time of the past inference.
- csmpltn 2y ago> "ability to time travel for training" Nah, this is nothing new. We've solved this for ages with "snapshots" or "archives", or fancy indexing strategies, or just a freaking "timestamp" column in your tables.
- nikhilsimha 2y agoSnapshots can’t travel back with milliseconds precision or even minute level precision. They are just full dumps at regular fixed intervals in time.
- _se 2y agoDatabases have had many forms of time travel for 30+ years now.
- threeseed 2y agoNot at the latency needed for feature serving and most databases struggle with column limits. But please enlighten us on which databases to use so Airbnb (and the rest of us) can stop wasting time.
- refset 2y agoShameless plug, but XTDB v2 is being built for low-latency bitemporal queries over columnar storage and might be applicable: https://docs.xtdb.com/quickstart/query-the-past.html https://docs.xtdb.com/quickstart/query-the-past.html We've not been developing v2 with ML feature serving in mind so far, but I would love to speak with anyone interested in this use case and figure out where the gaps are.
- mulmen 2y agoSnapshots don’t have to be at regular intervals and can be at whatever resolution you choose. You could snapshot as the first step of training then keep that snapshot for the life of the resulting model. Or you could use some other time travel methodology. Snapshots are only one of many options.
- nikhilsimha 2y agoThese are reconstruction of features / columns that don’t exist yet.
- jyhu 2y agoHave you guys considered Rockset? What you mentioned are some classic real-time aggregation use cases and Rockset seems to support that well: https://docs.rockset.com/documentation/docs/ingestion-rollups https://docs.rockset.com/documentation/docs/ingestion-rollup...