3 ms·
> It is also terrible design - having data in two places means you now have to implement access control, logging, auditing, data access and so on twice. I've u
by atwebb 5y ago
> It is also terrible design - having data in two places means you now have to implement access control, logging, auditing, data access and so on twice.
I've understood and implemented differently. With Spectrum (or Polybase for SQL Server / Synapse), you can extended into the data lake. Copy over aggregate/curated data or something you need to special use cases on. Leave the structured, columnar data within the cheap storage. You pay per scan but it is cheap (at least to a point).
Also, Databricks took the Lakehouse moniker and sprinted with it. AWS was late to the game from what I saw (at least for marketing terminology adoption).
- glogla 5y agoYou can do "lakehouse" just with Redshift, but in AWS pictures, you'll see Glue Jobs, Glue Elastic Views, Sagemaker, Aurora ... it's a huge mess. With ra3 redshift, you pay storage cost of s3 for internal data as well, so unless you use the s3 with something else, I don't see much point in using spectrum. Still, something like Snowflake works much better. They actually seem to have vision and not just "us too!" like AWS.
- atwebb 5y agoYou pay the storage cost in S3 as well, depending on tier Snowflake will not necessarily be a cost savings from compute either. Redshift could really use some elasticity beyond a factor of 2 and some warm resume features.
- glogla 5y agoJust to clarify, I meant Redshift ra3 storage costs as much as s3, so you're not saving much by keeping things in s3 instead of in Redshift. Although I don't know how Redshift compression compares to something like gzipped parquet. Maybe the data ends up taking more space and thus money. Agreed on that elasticity.