5 ms·
It strikes me that so much of the components they use (e.g. under "Storage") are in-house built (several dbs, blob store, caches, etc). Is that because at that
by rollulus 10y ago
It strikes me that so much of the components they use (e.g. under "Storage") are in-house built (several dbs, blob store, caches, etc). Is that because at that time equivalent solutions didn't exist? Is that because Twitter suffers from NIH?
- daveFNbuck 10y agoProbably a bit of both. I know the team I was on when I worked there was fairly proud of duplicating effort implementing things that other teams had already done. But Twitter does tend to build really good stuff.
- fizx 10y agoI built custom storage for Twitter back in the 2010-12 period. There wasn't much off the shelf in those days that worked out-of-the-box at scale besides Cassandra, and Twitter had a well-documented attempt at using Cassandra for primary workloads that failed due to the amount of variance in IO and latency for high-volume workloads. Most of the time, we were building custom distribution layers on top of open source storage (e.g. memcached/mysql/redis/etc). I think blobstore was the first thing twitter put in production that was mostly custom, followed a year or two later by manhattan. I'm not sure if there's even now good open-source competitors for those projects, largely because any reasonable smaller company uses s3 or dynamo. There's plenty of open-source things twitter created, or nurtured out of the existing ecosystem, from mesos to memcached to some of the hadoop/scalding/parquet stuff.
- seanp2k2 10y agoThe flip side of that imo is that if you don't have Twitter-scale needs for the specific things they've optimized their infrastructure for, you probably don't need their solutions :)
- TheAceOfHearts 10y agoI haven't used it, but one tool I've read good things about Minio [0]. I don't know if it's able to match twitter-scale, though. https://minio.io/ https://minio.io/
- throwawasiudy 10y agoSounds like it to me. Storing a crap ton of 300 byte messages is pretty common. Thousands of companies have been doing it for a decade. Anyone that does analytics or email probably stores far more of these than twitter. Blob store is perhaps the only forgivable custom solution. Besides the eventually expensive S3 you pretty much have to roll your own large binary storage at that scale. For regular DB sized workloads 1-16k bytes they had hundreds of options to choose from. Same with caching. They're both relatively solved problems.
- thinkingfish 10y agoIn general, Twitter infrastructure didn't start by building solutions in house. Like most early stage companies, it started by using whatever it can find in the OSS repertoire. But a lot of problems that are non-existing at a smaller scale or at least tolerable become more acute when you scale up. This is when the developers face a decision: to improve the existing solution, or develop a new one? There are many factors that impact the final decision. Maturity of the technology? Community support/inclusion? How big is the necessary change? What scale can existing solution handle by design? How does in-house talent compare to those maintaining the original? What's the focus/need of the in-house use cases compared to the broader objective of the OSS solution? It should not be a surprise that the answer differs from project to project. But if the decision is to use the existing solutions, it probably won't raise much eyebrow/question. It is very common for larger-scale operations to write up their own solutions. Because a different scale can change the nature of the problem fundamentally. Google, Facebook and whoever has production fleet above the 10K range tend to build a lot of in-house solutions, often after similar considerations as mentioned above.