4 ms·
Isn't it just a paradox to store infinite data, to use it later for very specific things without having to define it first? It sounds very common sense to not
by cateye 7y ago
Isn't it just a paradox to store infinite data, to use it later for very specific things without having to define it first?
It sounds very common sense to not to "limit the potential of intelligence by enforcing schema on Write" while in reality, the same problem just shifts (or gets hidden) in the next steps.
For example: there are 10 data sources with each 100TB of data. I aggregate these to my new shiny data lake with a fast adapter. Just suck it all without any worries about Schema. So, now I have 1PB of semi unstructured data.
How do I find the fields X and Y when these are all named differently in 10 sources? Can I even find it without having business domain experts for each data source? How do I keep things in sync when the structure of my data sources change (frequently)?
It seems like there is an underlying social/political problem that technology can't really fix.
Reminds me the quote: "There are only two hard things in Computer Science: cache invalidation and naming things."
- bradleyjg 7y ago> Reminds me the quote: "There are only two hard things in Computer Science: cache invalidation and naming things." and off by one errors!
- derefr 7y agoYou're not necessarily ingesting semi-unstructured data. Common Data Lake file formats (Avro, Parquet, ORC) are in fact highly structured, and even self-describing in their schema, with format-features like schema evolution allowing sibling datasets produced at different times to have "different" schemas which nevertheless have a single defined schema as the output when the datasets are unioned together. The idea, though, is that, if your Data Warehouse wants the data in the form of e.g. a daily-aggregate accounting ledger, then your data sources might be of various time granularities and might be denormalized in different ways (one source with separate Invoices with Transactions foreign-keyed to an Invoice; another with just Transactions with root-level metadata like timestamp directly on them; etc.) All of the transformations between the source formats and the destination format here are, in some sense, "transparent"—a sufficiently-advanced DBMS query planner could generate an OLAP expression to turn one into the other without understanding the problem domain. It's precisely because of this that, in many cases, it's cheaper to not worry about these kinds of transformations until you need to compute on the data. It's just a bunch of trivial stuff, that you can easily normalize in the computation step, but where fixing it on ingest would have been a whole expensive cluster operation to rewrite terabytes of data, and would require the OpEx of a whole additional set of always-online Hadoop cluster-nodes to fix marginal data as it comes in. Even though you're just going to be touching it all again anyway when you run it through the compute step.