4 ms·
So, the general architecture described here is solid, and I support it, but I take issue with the "Lakebase" naming thing. Disaggregated storage and disaggrega
by gavinray 5mo ago
So, the general architecture described here is solid, and I support it, but I take issue with the "Lakebase" naming thing.
Disaggregated storage and disaggregated compute have been an open trend in DBMS development for the last half-decade. This is an obvious move with modern computing paradigms, and the academic literature has a standard name for it.
This feels like "JAMStack" from Netlify happening all over again.
I tweeted about this in 2022, as a general trend, and also from the RocksDB meetup emphasizing disaggregated storage:
- https://x.com/GavinRayDev/status/1607769112234823680 https://x.com/GavinRayDev/status/1607769112234823680
- https://x.com/GavinRayDev/status/1600666127025156096 https://x.com/GavinRayDev/status/1600666127025156096
- jeremyjh 5mo agoI don't think it should be surprising that vendors are not going to lead with "disaggregated storage". I don't see that taking off either. This isn't a paper in a journal. Aurora doesn't call it that either. But yes, it is not a new idea.
- gavinray 5mo agoAvoiding industry-standard names and trying to introduce your own convention comes off as hubristic and grift-ey to me. "Basic literacy" -> "Prompt Engineering" "P2P networking" -> "Web3" "Service-Oriented Architecture" -> "Microservices" Maybe I'm old-man-yelling-at-cloud.
- jeremyjh 5mo agoIts more like old-man-yelling-at-billboard. Its just marketing. Its like complaining about the font they chose.
- nikita 5mo agoLakebase is referring to the fact that in addition to disaggregated storage s3 is authoritative storage for older data. Since data is on s3 (or lake) you can perform direct to s3 type operations like data loading, reading this data by engines that are not Postgres and more
- gavinray 5mo ago> in addition to disaggregated storage s3 is authoritative storage for older data Suppose a person retrives cold data from another Object Storage protocol rather than S3. This is no longer a "Lakebase", so we have to come up with a different name to avoid confusion. But if you say "Disaggregated Storage on S3" then you have the flexibility to change that to "Disaggregated Storage on FOOBAR" to avoid confusion.
- deleted 5mo ago[deleted]
- majormajor 5mo ago> Suppose a person retrives cold data from another Object Storage protocol rather than S3. This is no longer a "Lakebase", so we have to come up with a different name to avoid confusion. I've never seen "lake" or adjacent terminology refer to S3 specifically like that vs other object storage. A data lake on Ceph would still be a data lake. (My quibble would be that "lake" often refers to inconsistent or unstructured, and itself has always been a bit handwavy compared to "warehouse," whereas this is very structured data on object storage.)
- andyferris 5mo agoYes. Maybe I’m wrong, but AFAICT this is block (page) storage backed by S3, tuned for Postgres with some paxos-linked storage/caching servers sitting in front? Sounds good, but I’m not sure “lake” or “warehouse” is a word I’d choose… much closer to Litestream-with-reads, or the somewhat-famous “I ran out of RAM so I downloaded some more” blog article.