26 ms·
How Discord Stores Billions of Messages (2017)
- jjice 5y agoI love Discord's tech blog. There are a few corporate tech blogs that are just fantastic. Fly.io is another one that has great writing and interesting topics.
- dragonfax 5y agoKKV databases (Cassandra and DynamoDB are good examples) have a common problem with hotspots or "hot partitions". The most common mistake is to use a timestamp of any kind in the range (cluster) column. Then, whatever partition represents "today" or "this hour" ends up being the hot partition. The article mentions hot partitions becomming a problem with max partition size, but they're also a problem with scalability. Say, if your writing a very high throughput of logs into the table (contrived example), then your bottlenecked by the rate at which you can write to one partition. Adding the bucket id (say, the current day or hour), is a common solution, and solves the max partition size issue, but not the scalability issue of hot partitions.
- geenat 5y agoCockroach DB recently addressed hotspots on sequence/timestamp workloads with: https://www.cockroachlabs.com/blog/hash-sharded-indexes-unlock-linear-scaling-for-sequential-workloads/ https://www.cockroachlabs.com/blog/hash-sharded-indexes-unlo... Does what it says on the tin for the primary key. That said, hotspots are 100% the reason why Cockroach encourages UUID primary keys. The disadvantage to UUID is you want sequential data, you then need a secondary index which you'll have to bucket anyway.
- welder 5y agoTwo horrible choices of databases... first they used MongoDB then they migrated to Cassandra. I've used tons of databases [1] in production and those two are the worst. [1] I've used RethinkDB, Postgres, MongoDB, MySQL, Cassandra, CockroachDB, TimescaleDB, SSDB, and others
- kin 5y agoCassandra has been successfully deployed at many companies. Would you care to provide some insight into your experience and why you consider it one of the worst?
- welder 5y agoHere's a bug I ran into a few years ago: https://datastax-oss.atlassian.net/browse/PYTHON-891 https://datastax-oss.atlassian.net/browse/PYTHON-891 With all the issues I encountered using in prod, it gave the impression of an overly complicated key/value store.
- winrid 5y agoWhat would you have chosen?
- welder 5y agoAn in-house data service built with RocksDB/LevelDB. [1] [1] https://rocksdb.org/ https://rocksdb.org/
- anaganisk 5y agoIm not questioning your intellect, but each database has its own use case for usage. You might be expecting wrong things from Mongo or cassandra. Using gazillion databases for wrong use cases doesnt mean nothing. From your child comment, you cooked up an in house solution, it may be best suited for you. But for others it would be horrible too. If Cassandra is suited for Facebook, I think the problem is with you making a choice to use it for something its not suited for rather than the database itself.
- welder 5y agoYes, you're right.
- lmilcin 5y agoWell... we have 3 node MongoDB cluster and are processing up to a million trades... per second. And a trade is way more complex than a chat message. Has tens to hundreds of fields, may require enriching with data from multiple external services and then requires to be stored, be searchable with unknown, arbitrary bitemporal queries and may need multiple downstream systems to be notified depending on a lot of factors when it is modified. All this happens on the aforementioned MongoDB cluster and just two server nodes. And the two server nodes are really only for redundancy, a single node easily fits the load. What I want to say is: -- processing a hundred million simple transactions per day is nothing difficult on modern hardware. -- modern servers have stupendous potential to process transactions which is 99.99% wasted by "modern" application stacks, -- if you are willing to spend a little bit of learning effort, it is easily possible to run millions of non trivial transactions per second on a single server, -- most databases (even as bad as MongoDB is) have a potential to handle much more load than people think they can. You just need to kind of understand how it works and what its strengths are and play into rather than against them. And if you think we are running Rust on bare metal and some super large servers -- you would be wrong. It is a normal Java reactive application running on OpenJDK on an 8 core server with couple hundred GB of memory. And the last time I needed to look at the profiler was about a year ago.
- kevinsundar 5y agoThats not really what this article is about. Their problem wasn't throughput. What's the size of all the data in your MongoDB instance? And what's the latency in your reads? In the big data world the "complexity" of the data doesn't really mean much. It's just bytes.
- lmilcin 5y ago> What's the size of all the data in your MongoDB instance? 3x12TB > In the big data world the "complexity" of the data doesn't really mean much. Oh how wrong you are. It is much easier to deal with data when the only thing you need to do is to just move it from A to B. Like "find who should see this message, make sure they see it". It is much different when you have large, rich domain model that runs tens of thousands of business rules on incoming data and each entity can have very different processing depending on its state and the event that came. I am writing whole applications just to data-mine our processing flow just to be able to understand a little bit of what is happening there. At that traffic you can't even log anything for each of the transactions. You have to work indirectly through various metrics, etc.
- greenbeans1991 5y agoCould somebody help me understand the reasoning behind `timestamp = snowflake_id >> 22` Thanks :)
- karmanyaahm 5y agoIIRC a snowflake ID's first few bits (or characters?) are by definition a unix timestamp.
- greenbeans1991 5y agoAh it gets rid of the non-time related bits. Below is a description of the formatting * id is composed of: * time - 41 bits (millisecond precision w/ a custom epoch gives us 69 years) * configured machine id - 10 bits - gives us up to 1024 machines * sequence number - 12 bits - rolls over every 4096 per machine (with protection to avoid rollover in the same ms) thanks :)
- Notanothertoo 5y agoWe use scylla for our IoT stream, bucket per day, with a date index for second resolution data. The current day is a hot spot of, but we throw that in redis. It's running one of the largest re insurance providers IoT deployments.
- PeterCorless 5y agoThat sounds like one of my favorite use cases, Meshify / MunichRE! :) A little history lesson is in order: https://www.scylladb.com/2019/02/01/meshify-and-scylla-an-industrial-strength-iot-solution/ https://www.scylladb.com/2019/02/01/meshify-and-scylla-an-in...
- dancemethis 5y agoWell, very good for them and their "business partners". That data and metadata all for them to snoop into... Can't wait for OpenFeint 2.0 to have its scandal lawsuit too.
- tsjq 5y ago"Billions of messages" that's a prime target for acquisition by big-tech / big-data companies
- siculars 5y agoPartition Key selection aside, if you want a better Cassandra, look at ScyllaDB. Much, much better engine, imho. /disclaimer/ I used to work at Scylla.
- tbarbugli 5y agoIf you can, use Scylla over Cassandra. The performance difference is tremendous in my experience and replacing can be trivial (easier if you start with Scylla on day 0)
- anonymoushn 5y agoI am excited for the upcoming launch of full text search on discord.
- tester756 5y agoDiscord had like $300M invested and they created unparalleled piece of software that ate whole market, damn. One of the most impressive softwares that I've seen and use after years of using ventrilo/mumble/teamspeak.
- PHGamer 5y agoits just slack for gaming. the ui is ripped off as suck. it is better than ventrilo but its not like they are that much better, just they realized a good concept and took it.
- Aeolun 5y agoI don’t know about your definitions, but 300M is a fuckton of money to me.
- dvt 5y ago> ventrilo/mumble/teamspeak To be fair, Mumble is FOSS, and Ventrilo and Teamspeak have literally not iterated since 2005. Discord is pretty mediocre software (remember when they accidentally allowed iframe XSS RCE attacks? A very amateurish mistake), but the incumbents were an absolute dumpster fire.
- the_duke 5y ago> and Ventrilo and Teamspeak have literally not iterated since 2005 True, but to be fair: the next iteration of Teamspeak will be based on the Matrix protocol, which is quite a big iteration. See https://news.ycombinator.com/item?id=25743874 https://news.ycombinator.com/item?id=25743874 .
- Scaless 5y agoFor Mumble in particular, the devs had their heads in the clouds for so long that it is no surprise that it is no longer relevant. If you had a mic that had issues in any way (buzzing, volume, balance), "The Wizard" and "AGC" were supposed to fix it for you. Do not fret little one, for you do not need nor want to manually fiddle with settings, The Wizard will make everything right [1]! The pivotal feature that was the reason so many people I know stopped using it is the ability to change the volume of an individual person [2]. It has been a requested feature since the beginning of time, yet it took until 2016 to implement in dev branch and didn't actually make it into a release version until 2020! Too little, too late. [1] https://web.archive.org/web/20200223143654/https://wiki.mumble.info/wiki/FAQ/English#Can_I_change_to_volume_of_a_specific_user.3F https://web.archive.org/web/20200223143654/https://wiki.mumb... [2] https://github.com/mumble-voip/mumble/issues/1156 https://github.com/mumble-voip/mumble/issues/1156
- paulryanrogers 5y agoTLDR MongoDB then Cassandra
- typon 5y agoDid they make the move to ScyllaDB, as mentioned in their "Future work" section?
- misframer 5y agoTheir jobs pages (e.g. [0]) mention "ScyllaDB/Cassandra". [0] Senior Site Reliability Engineer: https://discord.com/jobs/4004051002 https://discord.com/jobs/4004051002
- mikesun 5y agoAlso: https://discord.com/jobs/5411664002 https://discord.com/jobs/5411664002 https://discord.com/jobs/5426301002 https://discord.com/jobs/5426301002
- Sikul 5y agoWe've moved quite a few datasets from Cassandra to Scylla, but not messages. I think we're planning to make a blog post about our experience with Scylla at some point.
- PeterCorless 5y agoDefinitely lemme know when you are going to do that! (peter@scylladb.com here).
- yeswecatan 5y agoStill a good read if you're curious about Cassandra
- gregoriol 5y ago2017 is pretty antique though now: the scaling and the ecosystem change fast
- jamesdwilson 5y agoWith privacy concerns, companies should be shamed for storing billions of messages.
- losteric 5y agoThat is a fair concern. However, as a customer - searching my message history is a desirable feature. I would rather see meaningful individual and corporate accountability for privacy breaches. The threat of jail and/or 100MM's in fines should motivate better data handling.
- jamesdwilson 5y agoWe've seen this happen over and over again: - a company amasses a large trove of sensitive information - it is exposed to adversaries or political enemies - the information is used against the people
- colesantiago 5y ago> With privacy concerns, companies should be shamed for storing billions of messages. should we shame ycombinator for storing the messages, accounts and comments on hacker news then? I am still unable to delete my account here even though the CCPA and the GDPR exists. But here we are.
- jamesdwilson 5y agoyes. whataboutism.
- colesantiago 5y agoSo I shouldn't be able to delete account information about me on HN or Discord? Care to explain this?
- jamesdwilson 5y agono, sorry if unclear. i meant the opposite, you should be able to.
- NotAnOtter 5y agoThis is like a masterclass in how to answer system design questions. Maybe a bit verbose. They cover requirements, how to answer those requirements, relevant tech for the problem, implementation, and techniques for maintenance
- deleted 5y ago[deleted]
- bcrosby95 5y agoAre you saying that one person, without consulting other engineers, or having to do any research, made this decision in less than a day? Because I have some bad news for you. The sentiment you express here is why interviewing is so shit these days.
- golergka 5y agoWell, that's true for any task that you're asked to complete at an interview: you're doing the fast, draft version that is obviously not comparable in quality to what you would do in a real setting. But still, even this draft could be illuminating.
- whymauri 5y agoAre you replying to the right comment?
- imdsm 5y agoThey must be. Totally out of context response.
- NotAnOtter 5y agoNo? But I am saying this is like an extremely articulate, over the top, excessively detailed answer for a system design question. I'm not saying this is what people should aim for, just that it's a good example of the types of things you should discuss during a system design interview. I mean that's what it IS. It's a system design.
- vortico 5y agoI'm curious of the 2021 measure of total disk space that Discord consumes. Servers that I'm in share images every few minutes, which must add up pretty quick.
- SilverRed 5y agoI'd imagine that the images are not part of the main database and that they are in some kind of s3 like file storage system.
- latchkey 5y agoUnique images or just copied from elsewhere?
- objektif 5y agoHow does that matter? Do they keep track of all images on the internet?
- jjice 5y agoI don't know much about image de-deuplication, but maybe they can get some sort of fingerprint/hash for an image, see if they already have it, and then serve that already existing image. I'd imagine a hash like SHA256 would be tricky because if that image was compressed an additional time at all throughout it's internet journey, then we'd get a different resulting hash, but maybe there is an effective way to fingerprint images. I have a utility on my machine (czkawka maybe?) that does really good image de-duplication with what seemed like a common algorithm (based on a quick look at the source). No idea though, just spit balling.
- PeterCorless 5y agoYes. There are ways to group images that seem to be the same. TinEye and Google image search do that. So you'd have a collection of related hashes that equal "Bob's prom photo where he looks like a goofer."
- klaussilveira 5y agoDid they ever make it to ScyllaDB?
- jhgg 5y agoYea. Everything but messages is on Scylla.
- res0nat0r 5y agoIt's unfortunate Discord is still requiring relocation to SFO, the product is amazing and it looks like some awesome engineering behind the scenes that would be fun to work on!
- Sikul 5y agoThis isn't true anymore. We switched to allowing and supporting permanent remote work (writing this from Seattle).
- res0nat0r 5y agoI applied the other month for a job that mentioned SFO or remote, then halfway through the signup it stated that they were allowing folks to work remote until COVID was better, then wanted folks to be onsite, and then prompted for a yes/no if I was willing to move to SFO at a later date. Didn't get a chance to talk to anyone and expect it was because of this, so is a bit disappointing.
- misframer 5y agoLooks like they migrated (at least partially) to Scylla: "Discord Chooses Scylla as Its Core Storage Layer" (2020) https://www.scylladb.com/press-release/discord-chooses-scylla-core-storage-layer/ https://www.scylladb.com/press-release/discord-chooses-scyll...
- PeterCorless 5y agoYes. They started experimenting with Scylla earlier, and made the switch in 2020. Here's more on their logic — they like "opinionated systems": https://www.scylladb.com/2019/03/20/discord-on-the-joy-of-opinionated-systems/ https://www.scylladb.com/2019/03/20/discord-on-the-joy-of-op...
- tomnipotent 5y agoScylla is really impressive. The only complaint I've heard is about the shard-per-cpu approach, causing some issues when data has bad distribution.
- umvi 5y agoDiscord is so good. I just can't imagine it can stay this good forever. My fear is that eventually it will be bought out and aggressively monetized.
- dancemethis 5y agoIt's not good. It's user hostile software.
- knownjorbist 5y agoOnly out-of-touch tech elitists on HN think this.
- throwaddzuzxd 5y agoAs someone who loves Discord, I'm curious why you'd think that, can you elaborate how it's user hostile?
- alpb 5y agoIf you're paying for Discord every month, it's actually fairly expensive. A lot of the good features unlock once people start boosting servers with Nitros and those aren't cheap either. So I'd assume they aren't bleeding cash left and right on infra costs. They might actually breaking even on the infra costs at least.
- helen___keller 5y agoFor years I've had a little bet going with friends about who ends up buying them to subsidize all this. My money was on amazon, because it could work so well with twitch + amazon prime.
- kerblang 5y agoI've used cassandra quite a bit and even I had to go back and figure out what this primary key means: ((channel_id, bucket), message_id) The primary key consists of partition key + clustering columns, so this says that channel_id & bucket are the partition key, and message_id is the one and only clustering column (you can have more). They also cite the most common cassandra mistake, which is not understanding that your partition key has to limit partition size to less than 300MB, and no surprise: They had to craft the "bucket" column as a function of message date-time because that's usually the only way to prevent a partition from eventually growing too large. Anyhow, this is incredibly important if you don't want to suffer a catastrophic failure months/years after you thought everything was good to go. They didn't mention this part: Oh, I have to include all partition key columns in every query's "where" clause, so... I have to run as many queries as are needed for the time period of data I want to see, and stitch the results together... ugh... Yeah it's a little messy.
- habibur 5y agoReading the article I was right now visiting Cassandra site to figure out what the catch is. Surely there should be a catch. Well, here it is. The partitioning in manual upto the SQL level.
- PretzelPirate 5y agoThe bigger catch is that when your partition grows too big and your nodes are hit by the OOMKiller, you have very few options other than create a new table and replay data, or use a cli tool to manually partition your data while the node is offline. Using Cassandra tends to mean pushing costs to your developers instead of spending more money on storage resources, and your devs will almost certainly spend a ton of time fixing downed nodes. Apple supplied some of the biggest contributors to Cassandra who were optimizing things like how to read data in a partition without fully reading the partition into memory to avoid the terrible GC cost. They put in a ton of engineering effort that probably could have been better spent elsewhere if they’d used a different database.
- 5y ago
- jgilias 5y agoCan anyone share experiences with using Discord as a communications tool in a workplace? We're currently on Google Chat because it comes with the package that we pay for anyway, but it's pretty lame. So from time to time we consider jumping to Slack. But then, why not Discord?
- kroltan 5y agoWe used Discord for a while for our team at work for over a year. Stopped using because company policy changed (we had no centralized chat program, some teams were on Skype, some on WhatsApp, then Slack was instituted). Context: 10-person, mostly technical, project-oriented, game development team. --- It was a joy to use, we created channels left right and center and knew everyone needed would be in them thanks to the centralized "role-based" permission system. (We would create project-specific channels and an accompanying role, or client-specific roles for the few high-throughput clients that had lots of small projects) At the time it did not have threading, which was one of the biggest pain points on the text-chat front. --- The voice chat is very good, and having dedicated voice channels means you can emulate meeting rooms or desks and have people join as desired/needed. You could be working and idle at "kroltan's desk" voice channel, but even if you weren't, joining one is trivial (a single click, can be done independently by many people) compared to Slack (find the call button somewhere different each time because they redesign the UI every week, then wait for your peer to join the call). Screen sharing is 720p on the free plan, so for meetings, it was hard to read documents, requiring zooming and whatnot. At the time there was also no setting to optimize for framerate or definition, so even 720p felt closer to 480p. Nowadays you can lower the framerate and also select the desired optimization, so you can ask Discord to optimize the stream for quality which is much better for documents, even in 720p. --- The client is also much more responsive than other Electron-based chat programs, especially with big workspaces with close to a hundred channels (yes, for a 10-person team, we sure type a lot), search is basically instant and has very useful filters, mentioning roles is great and the notification settings are fine-grained enough to please everyone.
- onkoe 5y agoIt's everything you could want, but this a lot of asterisks. A lot of things are limited (check other comment), but most importantly, their policies force you to hand over any text you write on their platform. But frankly, I'm not sure if there's anything near a perfect solution. A lot of companies use a big clump of services including both self-hosted and "rented". I hope to one day see a better, more comprehensive solution in line with the future of online work.
- ChrisArchitect 5y agoplenty of discussion when this was news: https://news.ycombinator.com/item?id=13439725 https://news.ycombinator.com/item?id=13439725
- dang 5y agoThanks! Macroexpanded: How Discord Stores Billions of Messages Using Cassandra - https://news.ycombinator.com/item?id=13439725 https://news.ycombinator.com/item?id=13439725 - Jan 2017 (155 comments)
- ryanianian 5y agoTFA states: > we knew we were not going to use MongoDB sharding because it is complicated to use and not known for stability But then goes on to describe using Cassandra and overcoming sharding and stability issues. I.e., changing the key, changing TTL knobs, adding anti-entropy sweepers, and considering switching to a different cassandra impl entirely. Are these issues significantly harder to solve in MongoDB than Cassandra?
- PeterCorless 5y agoAt the time MongoDB's sharding story wasn't great. They've gotten better since, but still have a primary-replica set model that has a single point of failure/failover. Cassandra (and Scylla) are leaderless, peer-to-peer clustering. Any node can go offline and the cluster keeps humming. Cassandra shards per node. Scylla goes beyond that and shards per core. Cassandra and Scylla also use hinted handoffs so if a node is unavailable temporarily (up to a few hours) you can store "hints" for it when it comes back online. Handy for short admin windows.
- trashcan 5y agoMongoDB has the equivalent of hinted handoffs. Changes are streamed to secondary nodes via the oplog, and the secondary just resumes where it was once it is back online. There is a limit to how long it can be offline (based on the size of the oplog), but that is the same limitation as hinted handoffs.
- PeterCorless 5y agoThanks! Good to know.
- calmoo 5y agoA MongoDB shard isn't necessarily a single-point-of-failure since a shard is usually deployed as a replica set. If a shard's primary node goes down, a secondary node in the replica set is elected as a primary and takes reads + writes. Similar to what you mentioned for Scylla - a node can go offline on a shard in a MongoDB cluster and it keeps humming.
- Ansil849 5y agoNothing about any sort of encryption.
- PeterCorless 5y agoScylla, Discord's replacement for Cassandra, supports both encryption in transit (server-to-server within the cluster; client-to-server) and encyption at rest for stored data. More on the latter here: https://docs.scylladb.com/operating-scylla/security/encryption-at-rest/ https://docs.scylladb.com/operating-scylla/security/encrypti...
- imdsm 5y agoThat's an interesting point too. They talk about not being a blob store, not wanting the serialisation cycle to hamper performance but makes you wonder how exactly they're storing the data. I'd guess it's not encrypted at all. ETA: Going back to the original thread, the whole question of encryption seems to be dodged and that usually means the answer isn't the one people are looking for: https://news.ycombinator.com/item?id=13440921 https://news.ycombinator.com/item?id=13440921
- mikesun 5y agoWe're actively hiring for our storage infrastructure team (SF or remote) ! If this sounds interesting, check us out! https://discord.com/jobs/5411664002 https://discord.com/jobs/5411664002 https://discord.com/jobs/5426301002 https://discord.com/jobs/5426301002
- c7DJTLrn 5y agoThey wouldn't need to store so many if they actually let people delete their messages on account deletion. Instead, they ban many people who attempt to do so via automated scripts.
- coldblues 5y agoOne of the big reasons I refuse to use Discord. Deleting your messages is a right every user should have. Whether it be individually or in bulk. The way it's done now just makes it more susceptible for users to be open to malicious attacks. Whether someone archives your content before you delete it, that's not of importance, that can happen on any internet medium.
- anigbrowl 5y agoI don't think they delete anything really. I've retrieved stuff from servers that were ostensibly deleted over a year ago.
- c7DJTLrn 5y agoWell, data is their greatest asset I'd assume. They don't want that precious user data going anywhere.
- simonw 5y agoDeletion of data at scale is a really difficult technical problem, unfortunately. I'm not saying they shouldn't do that though - especially given regulations like GDPR. Designing systems for deletion is important! But it's also really hard, especially if you didn't design for it from the start. There's also no way the tiny fraction of users who want to delete their data would make up a significant enough proportion of the messages that it would impact their scaling strategy.
- dave_sullivan 5y agoMaybe I don't know enough about databases, but could you really not have done this with postgres?
- simonw 5y agoI'm very much in the "use PostgreSQL unless you have absolutely proven to yourself that it won't work for your project" camp but in this case it really does look like moving to Cassandra was a good choice. NoSQL scalable stores like Cassandra basically only work well if you have a very strong model of the queries that you will need to make. In this case, that's exactly what they had: they knew what their read/write patterns looked like and they knew that they would be growing at hundreds of millions of rows per month, so easy horizontal scalability was a hard requirement. The biggest weakness of classical relational databases like PostgreSQL come when you have a super high volumes of inserts (as opposed to updates) which will continue to grow your database over time, and you need to keep all of that data accessible for real-time queries. They might have been able to achieve something like this using a PostgreSQL extension such as Citus, but it really does look like what they are doing fits Cassandra's sweet spot.
- chillacy 5y agoAfaik you could up until the data exceeds your capacity to fit in one machine, at which point you have to figure out how to split your data up in a way which lets you preserve all the strengths of sql (strong consistency). At that point you run into a lot of complexity with managing your shards.
- Quarrelsome 5y agodo they ever explain what their "anti-entropy" processes are for and what they do?
- PeterCorless 5y agoThere are a few: 1. Hinted Handoffs - if a node has a transient failure, the other nodes store up messages, like your buddy might take notes in class if you had to go to the bathroom. They'd pass you those notes when you got back. "Here's what you missed." When the node comes back online it processes all new operations and works through its backlog of hinted handoffs to get caught up. Because of the backlog it creates, hinted handoffs are only stacked up for a few hours. If the node never comes back up, or comes back after that window... 2. Repairs - in an eventually-consistent database you might miss an update or two over time. Or maybe you're a replacement node that has to fill in for a failed node. The replacement will get streamed data from the other replicas to get it started, or you might restore sstables from a backup, but then you should run a repair job to make sure all your replicas are properly in sync. (That's my understanding. Let me know if that sounds correct from the hands-on experts.)
- andrewstuart 5y agoI'd first reach for Postgres to do this. Anyone have any idea how Postgres would stack up in a similar challenge?
- aeyes 5y agoI see several issues: - No out of the box horizontal sharding, according to the post they had 4TB (compressed) data in the cluster in 2017. Looking at their growth I think it is safe to assume that today they would have >50TB which can't be done on a single node. You could use Citus but this is not exactly vanilla Postgres anymore. For such a simple data model wasting time implementing your own sharding solution and (more importantly) shard migration makes no sense. - Discord is storing text data, in Postgres this will be stored in TOAST tables which has some drawbacks. - Their workload is mostly inserts, almost no updates. Vacuum only operates on complete tables so you would wast I/O and CPU processing data which you don't even touch. You can partition tables but it's a manual process and you have to make compromises. In 2017, Postgres partitioning still had many performance drawbacks. - No out of the box redundancy. - Once your data doesn't fit in memory, Postgres performance becomes unpredictable. Personally I would have chosen ElasticSearch for this project.
- emptysea 5y agoWhy would the text data be stored in TOAST? My understanding was PG only uses TOAST when the data is too large to fit in the row, and since PG compresses data before inserting wouldn't user messages be fine?
- daveidol 5y agoI think Discord messages can be unbounded in size
- emptysea 5y agoInteresting, trying it out on Discord the default max msg size is 2000 chars and with Discord Nitro the max is raised to 4000 chars. Testing with Postgres, a 2000 char random sequence doesn't result in TOASTing, but a 4000 random sequence does get TOASTed And for kicks, 4000 chars that aren't random compress well enough that they don't end up in TOAST.
- secondcoming 5y agoThey should check out Scylla if they want even faster queries and - depending on their workload - fewer instances. Edit: It seems they have moved to Scylla
- kaladin-jasnah 5y agoThey already did—https://news.ycombinator.com/item?id=28293097 https://news.ycombinator.com/item?id=28293097, main article: https://www.scylladb.com/press-release/discord-chooses-scylla-core-storage-layer/ https://www.scylladb.com/press-release/discord-chooses-scyll...
- josephd79 5y agooh great, here come the mongodb haters. I do use discord for a few groups, too bad they will not allow 3rd party clients because a discord terminal app similar to irssi would be awesome
- PeterCorless 5y agoMongoDB is great for developers. Very facile to get started. However, it tends to fall over when it hits scale — which could be in total data set size (like, >TB scale), transaction scale (>100k ops) or in low latencies (submillisecond to single-digit millisecond). In any of those domains, if you are trying to solve your problem with MongoDB you are in for a world of hurt. That's generally when people start looking at other options. Whether an in-memory system for pure speed, or a horizontally scalable system for raw size or throughput.
- legerdemain 5y agoWe took a big bet on Cassandra, and then on an opinionated wrapper around Cassandra at $PASTJOB. The use case was a text search engine for syslog-type stuff. The product we built using Cassandra was widely known as our buggiest and least maintainable, and it died a merciful death after several years of being inflicted on customers. We didn't have a good handle on the exact perf implications of different values of read/write replication. Writing product code to handle a range of eventual consistency scenarios is challenging. The memory consumption and duration of compactions and column/node repair jobs is hard to model and accommodate. It's hard to tell what the cluster is doing at any given moment. Our experience with support plans from Datastax was also pretty dismal. Maybe the situation has changed since 2016. In my experience with several employers since then, it seems like every enterprise architect fell in love with Cassandra around 2014-2015 and then had a long, painful, protracted breakup.
- throwdbaaway 5y ago> Maybe the situation has changed since 2016. In my experience with several employers since then, it seems like every enterprise architect fell in love with Cassandra around 2014-2015 and then had a long, painful, protracted breakup. I think 2012-2014 was peak marketing from DataStax. There would be some new major feature with every new blog post, and it would mostly never work as expected. Between 2017 and now, things have settled down.
- trashcan 5y agoI've used Cassandra at two companies, and had the exact same experience as you at the first company. At a much bigger company that had some very, very highly paid Cassandra DBAs it was actually a relatively smooth experience.
- jitans 5y agoI would have used CockroachDB, it has all the requirements listed and you don't need to know in advance the queries you will perform when deciding the database schema.
- rnotaro 5y agoOn the day that this blog post was written, CockroachDB was only at beta-20170112 and didn't even had a production release yet. v1.0 was released on May 10, 2017 [1], so I doubt it was even on their mind when they started working on the project. [1] https://www.cockroachlabs.com/docs/releases/index.html https://www.cockroachlabs.com/docs/releases/index.html
- PeterCorless 5y agoAmazing. The whole industry has come quite a ways since 2017!
- PeterCorless 5y agoThis would not work at scale for a company like Discord, with its volume of traffic. Cockroach, being consistency-oriented would quickly become transaction-bound. You want a database like a Cassandra or Scylla that is more performance/availability oriented. Otherwise you are going to see a lot of lag and latency in the Discord chat. Cockroach is very, very good for a distributed SQL database. But it's still performance-limited in its very nature. More here on the difference between NoSQL/NewSQL performance, using Scylla (a CQL-workalike) as a point of comparison: https://www.scylladb.com/2021/01/21/cockroachdb-vs-scylla-benchmark/ https://www.scylladb.com/2021/01/21/cockroachdb-vs-scylla-be...