14 ms·
The broad set of computer science problems faced at cloud database companies
- xnx 3y agoWhen I read that Google installed their own atomic clocks in each datacenter for Spanner, I knew they were doing some real computer science (and probably general relativity?) work: https://www.theverge.com/2012/11/26/3692392/google-spanner-atomic-clocks-GPS https://www.theverge.com/2012/11/26/3692392/google-spanner-a...
- ignorantguy 3y agoyeah synchrozing clocks across distributed systems is really hard without expensive hardware
- audioheavy 3y ago...but for distributed databases specifically, you can use a different algorithm like Calvin[0] or Fauna[1] that do not require external atomic clocks… but the CS point and the wealth of info in research papers (in distributed systems stuff) are solid ...but there is a lot of noise in those software papers, too - you are often disappointed by fine print, or have good curators/thought-leaders [2] - we all should share names ;) enjoying the discussion though - very timely if you ask me. -L, author of [1] below. [0] - The original Calvin paper - https://cs.yale.edu/homes/thomson/publications/calvin-sigmod12.pdf https://cs.yale.edu/homes/thomson/publications/calvin-sigmod... [1] - How Fauna implements a variation of Calvin - https://fauna.com/blog/inside-faunas-distributed-transaction-engine-dte#a-flexible-layered-system https://fauna.com/blog/inside-faunas-distributed-transaction... [2] - A great article about Calvin by Mohammad Roohitavaf - https://www.mydistributed.systems/2020/08/calvin.html?m=1#:~:text=Calvin%20is%20a%20transaction%20scheduling,that%20unlike%202PL%20is%20deterministic https://www.mydistributed.systems/2020/08/calvin.html?m=1#:~....
- kmod 3y agoI hear this from tech people, but hft people are happily humming along with highly-synchronized clocks (mifid ii requires clocks to be synchronized to 100us). I wouldn't say it's "easy" but apparently if you need it then you do it and it's not that bad.
- joshuamorton 3y ago> (mifid ii requires clocks to be synchronized to 100us) That only applies if the clocks are within 1ms of each other, so around 100 miles (or equivalently: within a single cloud region), and only came in to force in 2014. The bound that Spanner-likes keep is ~3ms for datacenters across continents, and that was in 2012.
- liquidgecka 3y agoWe had hardware for time sync at Google 15 years ago actually, thats not a new thing. Actually.. we had time sync hardware (via GPS) in the early Twitter datacenters as well until it became clear that was impossible to support without getting roof access. =)
- pokstad 3y agoAccurate clocks in data centers predates Google. Telecom needed accurate clocks for their time sliced fiber infrastructure. Same for cellular infrastructure.
- totetsu 3y agoI wonder if you have to account for localized gravity differences for that kind of thing.
- LAC-Tech 3y agoChatGPT tells me that the time difference between the top of mount everest and sea level is in the nano second range - so completely dwarfed by network latency and maybe doesn't matter? Pure speculation on my behalf though.
- defrost 3y agoIf you were to supply that answer in a group concerned with time and gravity they might ask you if that was the right question to ask wrt "greatest relativistic time seperation between two points on 'the earth's surface'" ChatGPT likely won't help - but you could look into the fact the the earth isn't round, being an oblate spheroid with 20km less radius at the poles than the equator. Of course the fact the ideal WGS84 ellipsoid, the official mean global sea level, and the geoid (gravitational surface of equipotential) don't all align must surely come into play here - and that bloody great "gravitational hole" somewhere south of Ceylon. https://www.e-education.psu.edu/geog862/node/1820 https://www.e-education.psu.edu/geog862/node/1820 https://en.wikipedia.org/wiki/Gravity_of_Earth https://en.wikipedia.org/wiki/Gravity_of_Earth https://en.wikipedia.org/wiki/World_Geodetic_System https://en.wikipedia.org/wiki/World_Geodetic_System
- LAC-Tech 3y agoI've been working hard to up skill on the consistency and distributed systems sides of things. General recommendations: - Designing Data Intensive Applicatons. Great overview of... basically everything, and every chapter has dozens of references. Can't recommend it enough. - Read papers. I've had lots of a-ha moments going to wikipedia and looking up the oldest paper on a topic (wtf was in the water in Massachusetts in the 70s..). Yes they're challenging, no they're not impossible if you have a compsci undergrad equivalent level of knowledge. - Try and build toy systems. Built out some small and trivial implementations of CRDTs here https://lewiscampbell.tech/sync.html https://lewiscampbell.tech/sync.html, mainly be reading the papers. They're subtle but they're not rocket science - mere mortals can do this if they apply themselves! - Follow cool people in the field. Tigerbeetle stands out to me despite sitting at the opposite end of the consistency/availability corner where I've made my nest. They really are poring over applied dist sys papers and implementing it. I joke that Joran is a dangerous man to listen to because his talks can send you down rabbit-holes and you begin to think maybe he isn't insane for writing his own storage layer.. - Did I mention read papers? Seriously, the research of the smartest people on planet earth are on the internet, available for your consumption, for free. Take a moment to reflect in how incredible that is. Anyone anywhere on planet earth can git gud if they apply themselves.
- autotune 3y ago>- Did I mention read papers? Seriously, the research of the smartest people on planet earth are on the internet, available for your consumption, for free. Take a moment to reflect in how incredible that is. Anyone anywhere on planet earth can git gud if they apply themselves. There is a flood of papers out there with unrepeatable processes. Where can you find quality papers to read?
- bilalq 3y agoThis is a great point. It's true that there's a wealth of good information out there. But there's so much bad information that we now struggle with a signal vs noise problem. If you don't have enough context and knowledge yet to make the distinction, it's very easy to go down a wild goose chase. Having access to an expert in the field who can mentor and direct you is invaluable.
- deleted 3y ago[deleted]
- avrionov 3y agoFrom the article: > Another example is figuring out the right tradeoffs between using local SSD disks and block-storage services (AWS EBS and others). Local disks on AWS are not appropriate for long term storage, because when an instance reboot the data will be lost. AWS also doesn't offer huge amounts of local storage.
- AdamProut 3y agoyeah, that is part of the trade off. Using an ephemeral SSD (for a database) means the database needs to have another means of making the data durable (replication, storing data in S3, etc.). There are AWS instance types (I3en) with large and very fast SSDs (many times higher IOPS then EBS).
- anonymousDan 3y agoUnless you manage the replication across different local disks yourself
- betaby 3y ago> For instance, blob storages such as S3 have enabled cloud database providers to offer flexible, unlimited storage (SingleStoreDB even coined the term “bottomless storage” for this). Can someone please elaborate that? What does it mean in conjunction of S3 and DB. I know how traditional DBs work (PostgreSQL and MySQL). I know how S3 work (opensource implementation like minio). But S3 is not a random access file on block storage which is a prerequirement for PostgreSQL and MySQL. How is that solved for S3 based DBs? Can someone point out to the doc, or even better an opensource implementation.
- monero-xmr 3y agoS3 is a key-value store with random access and various pricing / limits. If you can make the disk IO part of your DB map to S3 API there is nothing stopping you other than network latency, and ensuring the way you use S3 is cost effective.
- richieartoul 3y agoYou usually have to redesign the storage from the ground up around S3. Some databases do transparent data tiering with S3 which can work ok too depending on the use case.
- deleted 3y ago[deleted]
- AdamProut 3y agoIts a popular design for SQL Data warehouses. I think almost all of them (snowflake, redshift, etc.) store cold data in S3 and hot data on local disk[1][2]. It works well if the data is stored as immutable files (i.e., A log structure merge tree) or is not index at all (classical columnstores). S3 doesn't provide an efficient way to update a file. [1] https://dl.acm.org/doi/10.1145/2882903.2903741 https://dl.acm.org/doi/10.1145/2882903.2903741 (snowflake SIGMOD paper) [2] https://dl.acm.org/doi/10.1145/3514221.3526055 https://dl.acm.org/doi/10.1145/3514221.3526055 (singlestore SIGMOD paper)
- zX41ZdbW 3y agoClickHouse has two modes of operation on S3: 1. S3 as main storage with a write-through cache. 2. S3 as a cold tier in tiered storage. It works well because the data is organized by a set of immutable parts called MergeTree. These data parts are atomically created, merged, and deleted, but never modified. S3 does not work well with random access... but neither it's needed.
- TuringNYC 3y agoI worked at a specialty database software vendor for almost 4yrs, albiet I worked on ML connectors. I recall some of the hardest challenges as figuring out each cloud vendor's poorly documented and rapidly changing/breaking marketplace launch mechanisms (usually built atop k8s using their own flavor (eks, aks, gke, etc)).
- richieartoul 3y agoI’ve built pretty much my entire career around this problem and it still feels evergreen. If you want a meaningful and rich career as a software engineer, distributed storage is a great area to focus on.
- LAC-Tech 3y agoCould I send an email? I'm sunsetting my CRUD career and going down this route, would appreciate some perspective.
- potsandpans 3y agoI'm a principal engineer in a faang-adjacent ecomm domain. I've been doing this for 11 years and I'm absolutely sick of it. there are no novel problems, just quarterly goals that compete with and dominate engineering timelines. how did you get started, and what would you recommend for pivoting into this space?
- ibgeek 3y agoI was hoping that the blog post would actually spell out examples of problems. Is it just me or have there been a lot of shorter blog posts on HN lately that are really no more than an introduction section rather than an actual full article?
- tayo42 3y agoI'm kind of confused by companies like the one in the post. What is the selling point of the these hosted DB companies running in AWS, when aws and the rest of the providers them selves provide pretty good, probably much better, DB services? Is there that much money to make between running DBs on EC2 compared to the existing offerings they have? Amazon, google, MS, these companies print money, have built up massive engineering cultures to run reliable storage. I just dont see what the value is with trusting data with some VC funded group over proven engineering work. I worked on one of these in house storage systems, all we did was look at how the cloud providers did things already for inspiration. Might as well just use those. IDK maybe someone can convince me of the value?
- audioheavy 3y agoSometimes the value is not in having a zillion configuration options with these providers but, instead, having a more accessible service that doesn't require a PhD. And some of the people in those VC-funded groups were alumni of those providers too. :) -L
- qaq 3y agoCause they provide solution AWS does not ? SingleStore is hybrid OLAP/OLTP AWS does not have one of those. Neither does it have a proper horizontally scalable new SQL DBMS like CockroachDB. Snowflake is way better than what you can get from AWS for OLAP. and so on
- cmrdporcupine 3y agoThe risk is not that AWS offers a better service now (many of these companies are actually offering something better). But that they copy your idea or service and do it later. I'm given to understand Snowflake runs its own cloud platform, at least in part.
- deleted 3y ago[deleted]
- pradeepchhetri 3y agoIf you are interested in performance aspect of databases, I would recommend watching this great talk [0] from Alexey, ClickHouse Developer, where he talks about various aspects like designing systems by first realising the hardware capabilities and understanding the problem landscape. [0]: https://www.youtube.com/watch?v=ZOZQCQEtrz8 https://www.youtube.com/watch?v=ZOZQCQEtrz8
- samsquire 3y agoThere is so many interesting problems to solve. I just want there was available libraries or solutions that solved a lot of them for the least cost, so that I may build on some good foundations. RocksDB is an example of that. I am playing around with SIMD, multithreaded queues and barriers. (Not on the same problem) I haven't read the DDIA book. I used Michaeln Nielsen's consistent hashing code for distributing SQL database rows between shards. I have an eventually consistent protocol that is not linearizable. I am currently investigating how to schedule system events such as TCP ready for reading EPOLLIN or ready for writing EPOLLOUT efficiently rather than data events. I want super flexible scheduling styles of control flow. Im looking at barriers right now. I am thinking how to respond to events with low latency and across threads. I'm playing with some coroutines in assembly by Marce Coll and looking at algebraic effects
- jollyllama 3y agoVery similar to the problems traditionally faced by engineered storage solution engineers one to two decades ago. There's a mix of those engineers and newer folks from academia or cloud in general leading the solutions for the cloud.