4 ms·
"Instead, they keep a Thing Table and a Data Table. Everything in Reddit is a Thing: users, links, comments, subreddits, awards, etc. Things keep common attribu
by ndemoor 14y ago
"Instead, they keep a Thing Table and a Data Table. Everything in Reddit is a Thing: users, links, comments, subreddits, awards, etc. Things keep common attribute like up/down votes, a type, and creation date. The Data table has three columns: thing id, key, value."
I hope they introduced some NoSQL sweetness by now.
- encoderer 14y agoExactly. They turned their RDBMS into a NoSQL database that's still slowed-down by all the relational machinery. Though I'm sure this was an informed choice at the time, I seriously hope nobody finds this advice actionable anymore. Use Cassandra. (Or your NoSQL DB of choice)
- rfurmani 14y agoThey in an ad hoc way turned postgres into a key-value store, but in reality they have Cassandra running as a "permacache" in front of it and, in practice, everything really just hits Cassandra. It looks like postgres could at some point be phased out. Source: ive hacked at the codebase to produce arxaliv
- jebblue 14y agoThe only problem with NoSQL is what happens when one day _you_need_to_relate_data?
- encoderer 14y agoIt's an all-of-the-above strategy. I remember when the Bigtable paper was released. It was very early in my career and I remember it sounding so alien to me. Sure, i had Memcached in my stack, but no SQL? It seemed like something they had to trade off to be able to build the kind of services they offer. I felt the same way after reading Dynamo. Sure, I thought a lot about data design. I thought about usage patterns to inform how we denormalize. And I grew into using, eg, Gearman, to pre-compute dozens of tables every night. I evolved, a bit. But a few years ago, a little before this OP was written, I had a great experience with some Facebook engineers and had an a-ha moment that has made me a much better software engineer. Basically, I realized that I needed to let my data be itself. If I have inherintly relational data, then it should be in a relational database. But I've built EAVs, queues, heaps, lists, all of these on top of MySQL and Postgres. Let that data be itself. We have more options now than ever before. K/V stores, Column stores, etc. I use a lot of Cassandra. A lot of Redis. Some Mongo. And a put a lot more in flat files than I ever thought I would. I know a lot of people that are smarter than me left the womb knowing these things. But for me it was transformational and has made me much happier. I realized how much energy I wasted fighting my own tools.
- notimetorelax 14y agoDoes it mean that in a single project you may use two different storage solutions? SQL DB for relational data and NoSQL for queues, heaps, etc. ? I'm asking because for me it is a kind of a paradigm shift, I have used plain files and SQL DB in the same project before, but never two different databases.
- blaines 14y agoYes. Always use the right tool for the job. Its okay to have some duplication (an object stored in postgres and also part of a mongo doc or cached in redis). Don't forget storage is cheap.
- ndemoor 14y agoIn the stack I am working on we have a variety of databases all serving a different type of data storage: - memcache: for caching of data that doesn't persist - redis: caching of data that needs a to be persisted short term but not on the longer term (eg. sessions) - MySQL: for user-like data (account details, addresses, projects, ...) - DynamoDB: for millions of data points that only needs to be queried in 1 dimension, so are not related or compared to one another. eg. give me all values from this table containing a given datatype, between 2 dates - MongoDB: for millions of datapoints that need to be queried on deeper levels - etc.
- notimetorelax 14y agoGreat, thank you for the detailed overview. If you don't mind me asking, do you ever end up having problems with consistency? (I imagine it could happen if some data was written to 2 databases and the second one rejected the transaction.)
- ndemoor 14y agoWell, the downside of using multiple DB stores is that the logic of keeping everything consistent is in the hands of the developer. So you have to make sure that everything is written correctly. For instance, if you write to MySQL and Mongo, but your Mongo is down, you'll either have to queue the data item somewhere for a write once the system is back up, or you have a migration system in place that gets everything from MySQL since the downtime and writes it back to Mongo. Depending on the type of data we have a few easing factors: for some data stores it is not that big of a deal if it doesn't get written to it's 2nd layer (eg. cache) as we can rewrite it the next time it is requested in layer 1.
- zorked 14y agoIf you need to relate you process the data offline with something like Pig, Hive or one of many other data analysis tools built on top of Hadoop.
- masklinn 14y ago> I hope they introduced some NoSQL sweetness by now. Postgres has good support for being used as a key/value store. http://blog.reddit.com/2012/01/january-2012-state-of-servers.html http://blog.reddit.com/2012/01/january-2012-state-of-servers... indicates they're still using Postgres for the main backend, and Cassandra next to it as both cache (it replaced memcached back in 2010) and to store data for some features (moderation log and flairs)