5 ms·
Ask HN: What tools do you use to monitor your “data stack”?
I don't mean Hadoop, etc. I mean "data engineering", using tools like Spark, Kafka, Airflow, BigQuery, Redshift. And then on the visualization side Tableau, Looker, etc. and the impact of poorly written SQL queries on your infrastructure.
- mattbillenstein 9y agoWith Airflow you can have it send an email on task failure which you can then hookup to a bunch of things -- slack, pagerduty, etc. We've also used rollbar for logging, so I've hacked the airflow script to setup a logger that reports warn and above to rollbar so we'll get alerts that way as well if something is deadlocked or whatever.
- scapecast 9y agoMatt, thanks - can you elaborate on your hack of the airflow script? What type of events do you set alerts on?
- mattbillenstein 9y agoTake a look at: https://github.com/rollbar/pyrollbar/blob/master/rollbar/logger.py https://github.com/rollbar/pyrollbar/blob/master/rollbar/log... for usage -- I just stick a call to a function in bin/airflow before the cli is run: https://gist.github.com/anonymous/7d2a50d863ff2210ac69563c76db0d3d https://gist.github.com/anonymous/7d2a50d863ff2210ac69563c76... This was to mainly catch "DAG is deadlocked" type of errors that seem to only come out of the logs.
- mattbillenstein 9y agoI'll add, my current data stack is mostly python jobs (Airflow BashOperator) moving data in and out of BigQuery using their json api. Visualization is using Metabase talking to BQ. GCS for data archival in .json.gz format.