3 ms·
Just to name a few off the top of my head -- - Data type mismatches between systems - Differences in handling ambiguous or bad data (e.g., null characters) -
by carlineng 4y ago
Just to name a few off the top of my head --
- Data type mismatches between systems
- Differences in handling ambiguous or bad data (e.g., null characters)
- Handling backfills
- Handling table schema changes
- Writing merge queries to handle deletes/updates in a cost-effective way
- Scrubbing the binlog of PII or other information that shouldn't make its way into the data warehouse
- Determining which tables to replicate, and which to leave behind
- Ability to replay the log from a point-in-time in case of an outage or other incident
And I'm sure there are a lot more I'm not thinking of. None of these are terribly difficult in isolation, but there's a long tail of issues like these that need to be solved.