4 ms·
Interesting - I was vaguely aware of the existence of bookmarks. I'd be interested to know about what you're using them for - they definitely _sound_ useful. I
by RobinL 6y ago
Interesting - I was vaguely aware of the existence of bookmarks. I'd be interested to know about what you're using them for - they definitely _sound_ useful. I guess it probably depends what sort of workloads you're doing. At the moment we use Airflow to manage DAGs/retries etc. I like it as a user, but from what I understand from our ops people it's a pain to manage.
- orf 6y agoThe use case is pretty simple. You’ve got a bucket that you want to load data from and shuffle it away somewhere else (redshift, s3, whatever). This could be populated by a Firehose, another system, etc etc. Bookmarks just store the greatest “created time” for the files you’re loading from s3. So when you trigger a job it will only load files created since the last successful run. It does some funky stuff to handle s3’s eventual consistency with LIST operations. Super simple incremental loading. This also works when loading data from a relational database, by storing the greatest primary key value.
- RobinL 6y agoThanks, that's really useful