3 ms·
I wonder if the underlying technology is from their acquisition of Alooma [1] several years back. High volume change data capture from an OLTP database to a dat
by carlineng 4y ago
I wonder if the underlying technology is from their acquisition of Alooma [1] several years back. High volume change data capture from an OLTP database to a data warehouse is a problem that seems simple on its face, but has a lot of nasty edge cases. Alooma did a decent job of handling the problem at moderate scale, so I'm cautiously optimistic about this announcement.
[1]: https://techcrunch.com/2019/02/19/google-acquires-cloud-migration-platform-alooma/ https://techcrunch.com/2019/02/19/google-acquires-cloud-migr...
- base3 4y agoDefinitely alooma. It even uses alooma's semantics for naming variables and staging tables.
- dominotw 4y ago> problem that seems simple on its face, but has a lot of nasty edge cases. I highly recommend teams not rolling out their own solutions because of all these edge cases. My boss refused to pay a vendor solution because it was a 'easy problem' according to him. Denied whole team promotion because we spent lot of time solving 'easy problem' , outright asked us 'whats so hard about it that it deserves promotion' . Worst person i've ever worked for.
- anonymousDan 4y agoWell what are the hard parts then?
- mywittyname 4y agoKeeping track of new record without hammering the database. Keeping track of updates without a dedicated column. Ditto for deletes. Schema drift. Managing datatype differences. I.e., Postgres supports complex datatypes that other DBs do not. Stored procedures. Basically any features of a DB that are beyond the bare minimum necessary to call something a relation database is an edge case. I rolled a solution for Postgres -> BQ by hand and it was much more involved than I expected. I had to set up a read-only replica and completely copying over the databases each run, which is only possible because they are so tiny.
- carlineng 4y agoJust to name a few off the top of my head -- - Data type mismatches between systems - Differences in handling ambiguous or bad data (e.g., null characters) - Handling backfills - Handling table schema changes - Writing merge queries to handle deletes/updates in a cost-effective way - Scrubbing the binlog of PII or other information that shouldn't make its way into the data warehouse - Determining which tables to replicate, and which to leave behind - Ability to replay the log from a point-in-time in case of an outage or other incident And I'm sure there are a lot more I'm not thinking of. None of these are terribly difficult in isolation, but there's a long tail of issues like these that need to be solved.