10 ms·
How Airbnb Built “Wall” to prevent data bugs
- SOLAR_FIELDS 4y agoIve been in the end stage of this (worked on data validation for a good chunk of my career) and these are my thoughts on the article: Determining blocking vs non blocking is a big issue - deciding which checks should be stoppers and which shouldn’t is often a matter of extensive debate. In my experience, only a few data checks are absolute show stoppers under any circumstance and a lot of things need to spawn tickets that should be routed to the correct team and followed up on. Some type of tracking system is necessary for this. Defining the logic of checks themselves in YAML is a trap. We went down this DSL route first and it basically just completely falls apart once you want to add moderately complex logic to your check. AirBnB will almost certainly discover this eventually. YAML does work well for the specification of how the check should behave though (eg metadata of the data check). The solution we were eventually able to scale up with was coupling specifications in a human readable but parseable file with code in a single unit known as the check. These could then be grouped according to various pipeline use cases. A model that plugs into an Airflow DAG as AirBnB has designed seems like a good approach. Often when it was time to incorporate checks into the pipeline we had heterogenous strategies to invoke our checks engines. Having a standardized approach helps drive adoption across the organization- oftentimes I’ve found that people are reluctant to run non critical checks if it’s a significant time and effort cost and will only run critical ones to try and push data quality accountability either upstream or downstream. If it’s really easy to turn on and incorporate that’s one less excuse that can be used to not run the checks.
- hitekker 4y agoYou seem to know what you're talking about. Ignorant question: do you think Dagster would work better as an orchestration/validation tool than AirBnB's Wall?
- SOLAR_FIELDS 4y agoI don’t know much about Dagster but it does not look like they have a validation tool equivalent to Wall, which requires Airflow. So you would not get validation with Dagster unless you brought it yourself.
- apahwa 4y agothe logic in Wall isn't defined in YAML, the logic is defined as SQL or code and then configured via YAML.
- thrav 4y agoYep. Sounds just like dbt’s approach.
- alexott 4y agoFor blocking checks - I personally use notion of errors and warnings, with errors definitely going to quarantine and propagated to good data, and warnings going to both good data and quarantine. It’s a trade off between not blocking all data and having a visibility of what is potentially bad. Another approach is to send everything into quarantine, but then giving users an instrument for rescuing their data, and further tuning checks to avoid this happening.
- maartenatsoda 4y agoRegarding YAML/DSL being a trap. Did you consider a fallback option to allow users to just plug in SQL? See this for example: https://docs.soda.io/soda-cl/user-defined.html https://docs.soda.io/soda-cl/user-defined.html
- testbjjl 4y agoMaybe Jim Buckmaster and Craig Neumark are taking notes.
- metadat 4y ago(For those ignorant like me, these are both Craigslist guys. Not yet clear what their relation to TFA is.)
- hideo 4y agoYeah I didn't get it either. Perhaps a reference to the time AirBNB allegedly farmed craigslist to grow listings? https://www.businessinsider.com/airbnb-harvested-craigslist-to-grow-its-listings-says-competitor-2011-5 https://www.businessinsider.com/airbnb-harvested-craigslist-...
- a2tech 4y ago
- daniel-cussen 4y agoScams could be exploiting data bugs. Literally all scams depend on verification failure.
- iratewizard 4y agoI'm the first person to criticize the dumpster fire that is Airbnb. I've hosted with them for 5 years and they've made awful decisions time and time again. Scams aren't as bad on Airbnb compared to VRBO and TripAdvisor, though.
- geoffjentry 4y agoIs this available for others to use or internal only? I think the answer is the latter as a google search didn't turn anything up and I didn't see anything in the article. But if I'm wrong I'd love to kick the tires a bit.
- d_burfoot 4y ago> Hive SQL, Spark SQL, Scala Spark, PySpark and Presto are widely used as different execution engines This makes me think they're doing something very very wrong. AirBNB does not have data on the scale that would require these tools. They have 5.6 million listings, 150 million users, and 1 billion total person-stays. These numbers can easily be processed with Postgres or SQLite on single machines. Spark and Hive are for companies like Google and Facebook. https://www.thezebra.com/resources/home/airbnb-statistics/#infographic https://www.thezebra.com/resources/home/airbnb-statistics/#i...
- rebelos 4y agoHave you ever worked in data engineering? They're using these systems for event data, data generated through transformations (multiplicative effect on base size), data used for ML, etc.
- Dylan16807 4y agoHow many events would there be per stay, in your estimation? And how many are actually important?
- tuckerman 4y agoThese events aren't just being generated per stay. A company like Airbnb will have events about logins, searches, site interactions, etc. You'll also be transforming the raw data and storing it again as higher level, materialized tables. Disclaimer: Worked at Airbnb (not on a data engineering or data infra team)
- Dylan16807 4y agoSo all unimportant data? I mean sure you can squeeze insights out of that but if a third of it disappeared overnight it wouldn't be a big deal. And even then anything short of obsessive mouse tracking won't be that much data. This isn't doing much to prove that the stuff in the article matters. Maybe it does but it's not self-evident and the criticism upthread makes sense. (Please note that I am not ignorantly saying the job is easy. I'm mostly wondering if it affects revenue and satisfaction by more than a tiny sliver to do the hard job with all these different big data engines as opposed to doing a much simpler job.)
- quadrifoliate 4y agoI'm a little bit annoyed at reading about details that seem closely connected to internal code (e.g. CheckConfigModel classes) without being able to see the source. I am not sure what others find so compelling about this blog post. Granted it's from Airbnb which probably has one of the more interesting data sets, but honestly it looks to me like an internal blog post that's been reposted to Medium without considering the viewpoint of an external user. I understand if they don't want to open source the framework; but then most of the blog post should be about design principles, maybe a bit about the process itself — not implementation details that seem directed towards an internal audience.
- jm1271 4y agoThanks for this post! Naive question: why not "just use Great Expectations"? At first blush GE seems like it has a lot of what you need out of the box: checks definable in YAML, extensibility, and connectors to many major data sources. Was there something you all found lacking there which made "roll your own" the right approach here?
- jabagonuts 4y agoAs a software engineer new to the data space, I am baffled by why people recommended great_expectations. It has a lot of questionable dependencies that inflate image sizes and lead to conflicts at scale. It is also a very ambitious project that fails to deliver on many fronts, including documentation and basic data quality checks. The complexity in writing your own checks is way too high. There’s a lot of very abstract concepts you have to understand before you can write a single line of code. If you think I’m wrong, stop now and go look at some of their code examples. You’re better of using python’s built-in unittest to run a query and then make assertions on the result as a task in your DAG
- charlysl 4y agoFrom related https://medium.com/airbnb-engineering/data-quality-at-airbnb-e582465f3ef7 https://medium.com/airbnb-engineering/data-quality-at-airbnb... > The new role requires Data Engineers to be strong across several domains, including data modeling, pipeline development, and software engineering. > comprehensive guidelines for data modeling, operations, and technical standards for pipeline implementation > Tables must be normalized (within reason) and rely on as few dependencies as possible. Minerva does the heavy lifting to join across data models. > When we began the Data Quality initiative, most critical data at Airbnb was composed via SQL and executed via Hive. This approach was unpopular among engineers, as SQL lacked the benefits of functional programming languages (e.g. code reuse, modularity, type safety, etc) > made the shift to Spark, and aligned on the Scala API as our primary interface. Meanwhile, we ramped investment into a common Spark wrapper to simplify reads/write patterns and integration testing. > needed to improve was our data pipeline testing. This slowed iteration speed and made it difficult for outsiders to safely modify code. We required that pipelines be built with thorough integration tests > tooling for executing data quality checks and anomaly detection, and required their use in new pipelines. Anomaly detection in particular has been highly successful in preventing quality issues in our new pipelines. > important datasets are required to have an SLA for landing times, and pipelines are required to be configured with Pager Duty > a Spec document that provides layman’s descriptions for metrics and dimensions, table schemas, pipeline diagrams, and describes non-obvious business logic and other assumptions > a data engineer then builds the datasets and pipelines based on the agreed upon specification