28 ms·
Production Twitter on one machine? 100Gbps NICs and NVMe are fast
- morphle 4y agoWhy not a single FPGA with 100Gbps ethernet or pcie with NVM attached? Around $5K for the hardware and $5K for the traffic per month. The software would be a bit trickier to write, but you now get 100x performance for the same price
- tluyben2 4y agoThat would be quite a nice project for fun.
- morphle 4y agoI have seen a number of papers of people doing such a fun project as their PhD thesis. I'll try to find some more [1] examples with links on Google Scholar. Try searching for p4 programmable dataplane content address". [1] Flexible Content-based Publish/Subscribe over Programmable Data Planes https://pure.rug.nl/ws/files/206093441/Flexible_Content_based_Publish_Subscribe_over_Programmable_Data_Planes.pdf https://pure.rug.nl/ws/files/206093441/Flexible_Content_base...
- mpoteat 4y agoLet's spend multi-million dollars a year on a team of highly specialized FPGA engineers writing assembly and HDL so that we can save 5k a month. Feature velocity will be 100x slower as well, but at least our application is efficient. I think that this may make sense for some applications, but I also think that if you can utilize software abstractions to improve developer efficiency, it reduces risk in the long run.
- morphle 4y agoThose millions of dollars have already been spent. For example the P4 [1] language (a HDL language) and the Tofino 3 chip. It started out as FPGA (NetFPGA) to do programmable packet routing at linespeed. You now have the P4 language to define your packet routers with and generate your code. Ten years later we have 25 Tbps software defined packet routers. [1] https://en.wikipedia.org/wiki/P4_(programming_language) https://en.wikipedia.org/wiki/P4_(programming_language) [2] https://www.intel.com/content/www/us/en/products/network-io/programmable-ethernet-switch/tofino-3-brief.html https://www.intel.com/content/www/us/en/products/network-io/...
- nimish 4y ago>Let's spend multi-million dollars a year on a team of highly specialized FPGA engineers I agree with you specialists are expensive but even a team of software engineers runs into the multi million dollar territory. Why not spend the same amount AND cut down resource use? Hyperscalers have shifted to custom hardware already.
- tylerhou 4y agoA bit trickier is a huge understatement.
- morphle 4y agoThat depends. If you would use our asynchronous runtime reconfigurable array called Morphle Logic [1] instead of FPGA (Field Programmable Gate Array), you could program this hardware Twitter in a 1000 hours and have it run at 50 Tbps. You would not need to know much about hardware but you would get a thousandfold speedup of your software. [1] https://github.com/fiberhood/MorphleLogic/blob/main/README_MORPHLE_LOGIC.md https://github.com/fiberhood/MorphleLogic/blob/main/README_M...
- twp 4y agoThis post solves all of the easy problems (i.e. make simple stuff go fast) and none of the hard problems (i.e. build a system that still works when other stuff breaks). This post is perfect world thinking. We don't live in a perfect world.
- throwmeup123 4y agoThe title is highly misleading for some theoretical "exploration".
- dang 4y agoOk, we've put a question mark up there to make it more explorationy.
- trishume 4y agoAs the author, this sounds good to me! I'll probably even change the actual title to match. I originally was going to make it a question mark and the only reason I didn't is https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headlines https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headline... when I think the answer is probably "could probably be somewhat done" rather than "no".
- dang 4y agoWell this may be the first time that's ever happened :) Betteridge antiexamples are always welcome. I once tried to joke that Mr. Betteridge had "retired" and promptly got corrected about his employment status (https://news.ycombinator.com/item?id=10393754 https://news.ycombinator.com/item?id=10393754).
- kierank 4y agoThis is as realistic as the moon rocket in my back garden.
- Cyph0n 4y agoMore like “barebones, in-memory, English-only Twitter clone on one machine”. Edit: Still a nice writeup!
- Dylan16807 4y agoWhat strikes you as English-only about it?
- Cyph0n 4y agoMy bad - not English, but ASCII. The assumed max tweet size is in bytes rather than (UTF) characters.
- trishume 4y agoI specifically assumed a max tweet size based on the maximum number of UTF-8 bytes a tweet can contain (560), with a link to an analysis of that, and discussion of how you could optimize for the common case of tweets that contain way fewer UTF-8 bytes than that. Everything in my post assumes unicode.
- Tepix 4y agoDid you consider URLs? They don't seem to count against tue size and can be very large indeed (like 4k)
- modeless 4y agoURLs are shortened and the shortened size counts against the tweet size. The URL shortener could be a totally separate service that the core service never interacts with at all. Though I think in real twitter URLs may be partially expanded before tweets are sent to clients, so if you wanted to maintain that then the core service would need to interact with the URL shortener.
- Cyph0n 4y agoThanks for clarifying. I missed the max vs. average analysis because I was focused on the text. Still, as noted in the Rust code comment, the sample implementation doesn’t handle longer tweets.
- eatonphil 4y agoThis is a great exercise in napkin math, even with constraints you've set for yourself that don't fully approximate Twitter (yet). Thanks!
- agilob 4y agoThe title reminded me about this https://www.phoronix.com/news/Netflix-NUMA-FreeBSD-Optimized https://www.phoronix.com/news/Netflix-NUMA-FreeBSD-Optimized (2019) and this 2 years later: https://papers.freebsd.org/2021/eurobsdcon/gallatin-netflix-freebsd-400gbps/ https://papers.freebsd.org/2021/eurobsdcon/gallatin-netflix-... (2021)
- syoc 4y agoLatest version: http://nabstreamingsummit.com/wp-content/uploads/2022/05/2022-Streaming-Summit-Netflix.pdf http://nabstreamingsummit.com/wp-content/uploads/2022/05/202... (2022)
- jeffbee 4y agoI like this kind of exercise. One thing I am not seeing is analytics, logs and so forth that as I understand it are significant portions of Twitter's production cost story.
- tluyben2 4y agoAnyone have a complete list of functional blocks that form Twitter? Beyond the obvious and what we see?
- Marazan 4y agoYou need the blocks for the obvious for what we see because it is not necessarily obvious to everyone. Over the last couple of months I've seen comments that summarise Twitter as a read-only service that doesn't have any real time posting requirements and similarly other comments that treat it as a write-only service with no real time read / fast search requirements. Without _all_ the blocks even the simple surface level Twitter will have complexity people miss.
- lazyasciiart 4y agoIf it's this cheap to run you don't need analytics because you don't need to monetize it, and if it's this simple you don't need logs because it'll all work correctly the first time!
- kevingadd 4y ago"You don't need to monetize it" who's going to fund your Twitter-as-a-charity? What happens when the free money goes away? Businesses have to pay the bills eventually one way or another, you need to plan for that in advance
- Nextgrid 4y agoIf you're a small team running this and can actually deliver it with a single machine (or two), just charging a few bucks a month for verification should net enough money to run it and provide a decent living (and out of 300M users, there will be people who would pay).
- varunkmohan 4y agoGood analysis. Obviously, this doesn't handle cases like redundancy and doesn't handle some of other critical workloads the company has. However, it does show how much real compute bloat these companies actually have - https://twitter.com/petrillic/status/1593686223717269504 https://twitter.com/petrillic/status/1593686223717269504 where they use 24 million vcpus and spend 300 million a month on cloud.
- judge2020 4y agoOn the other hand, Twitter does (or did) handle over 450 million monthly active users (based on stats websites), with a target for 315 monetizable daily active users (based on their earnings calls pre-privatization). Handling that amount of concurrency and beaming millions of tweets a day to home feeds and notifications is going to be logistically hard.
- WJW 4y agoIs that 315 million monetizable DAUs? That sounds like a lot if the total is only 450 MAU. OTOH, 315k DAU seems like it wouldn't be enough to pay the bills.
- judge2020 4y agoThere were some quarters with profit, some without; the past few years were mostly without IIRC. They were targeting 315 mDAUs for Q4 2023, but in the final earnings it was only 238 mDAUs. Actual MAU stats weren't public iirc but some random stats sites seemed to say 450m global MAUs, which likely includes people with no ad preferences or who only view NSFW content (which can't be shown next to (most?) ads). https://www.forbes.com/sites/johnkoetsier/2022/11/14/twitter-would-need-64-million-subscribers-to-replace-existing-revenue-and-cover-existing-losses/?sh=45d00edda95f https://www.forbes.com/sites/johnkoetsier/2022/11/14/twitter...
- saagarjha 4y agoThis includes users who aren’t shown ads, like people who are logged out or use third party clients.
- britneybitch 4y ago> colo cost + total server cost/(3 year) => $18,471/year Meanwhile the company I just left was spending more than this for dozens of kubernetes clusters on AWS before signing a single customer. Sometimes I wonder what I'm still doing in this industry.
- tluyben2 4y agoVery often. I work with companies spending 10x as much on bad (bizarrely complex; indeed kubernetes, lambda, gateway, rds etc) setup and bad code on aws. Almost no traffic (b2b). Makes no sense at all.
- Existenceblinks 4y agoTechies in tech industry are basically eating the rich .. A lot of buzzword to suck investment money in.
- paulryanrogers 4y agoOr rather cloud providers are eating the rich. Techies are carrying the plates.
- Existenceblinks 4y agoIt's a big meal. Those old-industry tycoons (oil/real estate/finance/alcohol/export/etc.) from gold rush era have too much money. There are a few places for money to land. Cloud seems to be the first wave, social media, messaging, big data, ai, smart everything. $200k-400k + $80k-$150k ranges are like hyenas army!
- traceroute66 4y ago> spending more than this for dozens of kubernetes clusters on AWS before signing a single customer Yup. Cloud is 21st century Nickel & Diming. Sure it sounds cheap, everything is priced in small sounding cents per unit. But then it very quickly becomes a compounding vicious circle ... a dozen different cloud tools, each charged at cents per unit, those units often being measured in increments of hours....next thing you know is your cloud bill has as many zeros on the end of it as the number of cloud services you are using. ;-) And that's before we start talking about the data egress costs. With colo you can start off with two 1/4 rack spaces at two different sites for resilience. You can get quite a lot of bang for your buck in a 1/4 rack with today's kit.
- jasonhansel 4y agoIf you really wanted to run Twitter on one machine at any cost, wouldn't an IBM mainframe be much more practical? You can even run Linux on them now. The specs he cites would actually be fairly small for a mainframe, which can reach up to 40TB of memory. I'm not saying this is a good idea, but it seems better than what the OP proposes.
- trishume 4y agoMy friend mentioned this just before I published and I think that probably is the fastest largest thing you can get which would in some sense count as one machine. I haven't looked into it, but I wouldn't be surprised if they could get around the trickiest constraint, which is how many hard drives you can plug in to a non-mainframe machine for historical image storage. Definitely more expensive than just networking a few standard machines though. I also bet that mainframes have software solutions to a lot of the multi-tenancy and fault tolerance challenges with running systems on one machine that I mention.
- toast0 4y ago> I wouldn't be surprised if they could get around the trickiest constraint, which is how many hard drives you can plug in to a non-mainframe machine for historical image storage. Some commodity machines use external SAS to connect to more disk boxes. IMHO, there's not a real reason to keep images and tweets on the same server if you're going to need an external disk box anyway. Rather than getting a 4u server with a lot of disks and a 4u additional disk box, you may as well get 4u servers with a lot of disks each, use one for tweets and the other for images. Anyway, images are fairly easy to scale horizontally, there's not much simplicity gained by having them all in one host, like there is for tweets.
- trishume 4y agoYah like I say in the post, the exactly one machine thing is just for fun and as an illustration of how far vertical scaling can go, practically I'd definitely scale storage with many sharded smaller storage servers.
- sethev 4y agoJohn Carmack tweeted something that made me noodle on this too: >It is amusing to consider how much of the world you could serve something like Twitter to from a single beefy server if it really was just shuffling tweet sized buffers to network offload cards. Smart clients instead of web pages could make a very large difference. [1] Very interesting to see the idea worked out in more detail. [1] https://twitter.com/id_aa_carmack/status/1350672098029694998 https://twitter.com/id_aa_carmack/status/1350672098029694998
- varjag 4y agoIsn't that what an OPA sorta kinda does.
- threeseed 4y ago> just shuffling tweet sized buffers to network offload cards Except that's not what it is doing at all. It assembles all the Tweets internally, applies an ML model to produce a finalised response to the user.
- seritools 4y ago> if it really was [which it isn't]
- threeseed 4y agoBut many people including the OP think it is. It’s like me running a web crawler on my phone and saying I can replace Google.
- Dylan16807 4y agoIf you were able to index the same amount of content you'd have a damn good alternative. And that's the level of this experiment.
- threeseed 4y ago
- andrewstuart 4y agoMost projects I encounter these days instantly reach for kubernetes, containers and microservices or cloud functions. I find it much more appealing to just make the whole thing run on one fast machine. When you suggest this tend to people say "but scaling!", without understanding how much capacity there is in vertical. The thing most appealing about single server configs is the simplicity. The more simple a system easy, likely the more reliable and easy to understand. The software thing most people are building these days can easily run lock stock and barrel on one machine. I wrote a prototype for an in-memory message queue in Rust and ran it on the fastest EC2 instance I could and it was able to process nearly 8 million messages a second. You could be forgiven for believing the only way to write software is is a giant kablooie of containers, microservices, cloud functions and kubernetes, because that's what the cloud vendors want you to do, and it's also because it seems to be the primary approach discussed. Every layer of such stuff add complexity, development, devops, maintenance, support, deployment, testing and (un)reliability. Single server systems can be dramatically mnore simple because you can trim is as close as possible down to just the code and the storage.
- judge2020 4y agoKubernetes and containers are a means to service architecture; It enabled scalability but does not require it. You should still be containerizing your applications to ensure a consistent environment, even if you only throw it in a docker-compose file on your production server.
- andrewstuart 4y agoI don't even use containers - I aim primarily for simplicity and so far I have found I am able to build entire sophisticated systems without a single container. Containers I find make things much more complex.
- John23832 4y agoThat just means you don't know how to architect a machine with containers (and if you're effective in what you do, that's ok). But it's a pretty objective notation that manually scaled single machines don't scale as well as automation.
- BeefWellington 4y agoI'm going to preface this criticism by saying that I think exercises like this are fun in an architectural/prototyping code-golf kinda way. However, I think the author critically under-guesses the sizes of things (even just for storage) by a reasonably substantial amount. e.g.: Quote tweets do not go against the size limit of the tweet field at Twitter. Likely they are embedding a tweet reference in some manner or other in place of the text of the quoted tweet itself but regardless a tweet takes up more than 280 unicode characters. Also, nowhere in the article are hashtags mentioned. For a system like this to work you need some indexing of hashtags so you aren't doing a full scan of the entire tweet text of every tweet anytime someone decides to search for #YOLO. The system as proposed is missing a highly critical feature of the platform it purports to emulate. I have no insider knowledge but I suspect that index is maybe the second largest thing on disk on the entire platform, apart from the tweets themselves.
- chippiewill 4y ago> I have no insider knowledge but I suspect that index is maybe the second largest thing on disk on the entire platform, apart from the tweets themselves. I'd probably go as far to say that the indexes _generally_ at twitter could be larger than the tweets
- trishume 4y agoQuote tweets I'd do as a reference and they'd basically have the cost of loading 2 tweets instead of one, so increasing the delivery rate by the fraction of tweets that are quote tweets. Hashtags are a search feature and basically need the same posting lists as for search, but if you only support hashtags the posting lists are smaller. I already have an estimate saying probably search wouldn't fit. But I think hashtag-only search might fit, mainly because my impression is people doing hashtag searches are a small fraction of traffic nowadays so the main cost is disk, not sure though. I did run the post by 5 ex-Twitter engineers and none of them said any of my estimates were super wrong, mainly just brought up additional features and things I didn't discuss (which I edited into the post before publishing). Still possible that they just didn't divulge or didn't know some number they knew that I estimated very wrong.
- drewg123 4y agoHow much bandwidth does Twitter use for images and videos? Less than 1.4Tb/s globally? If so, we could probably fit that onto a second machine. We can currently serve over 700Gb/s from a dual-socket Milan based server[1]. I'm still waiting for hardware, but assuming there are no new bottlenecks, that should directly scale up to 1.4Tb/s with Genoa and ConnectX-7, given the IO pathways are all at least twice the bandwidth of the previous generation. There are storage size issues (like how big is their long tail; quite large I'd imagine), but its a fun thing to think about. [1] https://people.freebsd.org/~gallatin/talks/euro2022.pdf https://people.freebsd.org/~gallatin/talks/euro2022.pdf
- cortesoft 4y agoIt is way more than 1.4TBs a second globally.
- xyzzy123 4y agoI wonder how much is api traffic and how much is assets & images.
- koolba 4y agoI wonder how much of that is crypto spam bots replying to each other.
- skywhopper 4y agoCan you detect and block those spam bots with less effort than it would take to process them?
- koolba 4y agoThat’s a not so simple calculation as comparing the raw request process time isn’t the complete picture. Spam bot content must be persistently stored, indexed, archived, etc so the long term cost is much more than the one off POST to create the entity. You’d also have to quantify the improved user experience from seeing less spam v.s. inflated ad revenue for garbage views / content.
- Tepix 4y agoThe new EPYC servers can be filled with 6TB of RAM and 96 cores per socket. Fun times.
- wonnage 4y agoThis doesn’t seem to support fetching a specific tweet by id?
- sgt 4y agoNo need, for those you can display a "Twitter is experiencing issues" or something similar. That will encourage the user to return to the main app and enjoying infinite scroll.
- habibur 4y agoHe will be up for surprise. HTTP with connection: keep-open can serve 100k req/sec. But that's for one client being served repeatedly over 1 connection. And this is the inflated number that's published in webserver benchmark tests. For more practical down to earth test, you need to measure performance w/o keep-alive. Request per second will drop to 12k / sec then. And that's for HTTP without encryption or ssl handshake. Use HTTPS and watch it fall down to only 400 req / sec under load test [ without connection: keep-alive ]. That's what I observer.
- JanisErdmanis 4y agoI guess that the biggest chunk of a slowdown from TLS comes due to group operations alone. So wouldn't it be practical to configure TLS for session resumption and limit the number of handshakes per second it could do?
- trishume 4y agoI agree most HTTP server benchmarks are highly misleading in that way, and mention in my post how disappointed I am at the lack of good benchmarks. I also agree that typical HTTP servers would fall over at much lower new connection loads. I'm talking about a hypothetical HTTPS server that used optimized kernel-bypass networking. Here's a kernel-bypass HTTP server benchmarked doing 50k new connections per core second while re-using nginx code: https://github.com/F-Stack/f-stack https://github.com/F-Stack/f-stack. But I don't know of anyone who's done something similar with HTTPS support.
- sayrer 4y agoUserspace networking is pretty common. The chair of the IETF even wrote one: https://github.com/NTAP/quant https://github.com/NTAP/quant "Quant uses the warpcore zero-copy userspace UDP/IP stack, which in addition to running on on top of the standard Socket API has support for the netmap fast packet I/O framework, as well as the Particle and RIOT IoT stacks. Quant hence supports traditional POSIX platforms (Linux, MacOS, FreeBSD, etc.) as well as embedded systems."
- lossolo 4y ago
- keewee7 4y agoIn the coming years we will probably see a lot of complicated microservice architectures be replaced by well-designed and optimized Rust (and modern C++) monoliths that use simple replication to scale horizontally.
- pixl97 4y agoReplication and simple never belong in the same sentence. DNS which is one of the simplest replication systems I know of has its own complex failure modes.
- qaq 4y agoIt can be outsourced complexity as in just using horizontally scalable database like Spanner, CockroachDB etc.
- pixl97 4y agoAll complexity must be paid for and typically the more complex the problem the higher the cost to manage.
- TideAd 4y agoCockroachDB is nice to use but every database has complexity you have to deal with. Here's one I ran into recently: if a range has only 1 of 3 replicas online then it will stop accepting traffic for that range until it has 3 replicas again. (for the folks at home, "range" is a technical term for 512 bit slice of the data - CRDB replicates at the range level) So, in some code I wrote, I had account for not only 1) the whole DB being unavailable but also 2) just one replica being unavailable (they're different failure modes that say different things about the health of the system). It's a good behavior! Good for durability. But I had to do some work to deal with it, spend an hour coming up with a solution, etc. There are databases that work at Twitter scale but no there are no silver bullets among those that do. You need full time engineers to manage the complexity and keep it online, or else it could cost the company shitloads of money - I've seen websites of similar scale where a two-hour outage cost them $20 million.
- PragmaticPulp 4y agoVery cool exercise. I enjoyed reading it. I see a lot of comments here assuming that this proves something about Twitter being inefficient. Before you jump to conclusions, take a look at the author’s code: https://github.com/trishume/twitterperf https://github.com/trishume/twitterperf Notably absent are things like serving HTTP, not to even mention HTTPS. This was a fun exercise in algorithms, I/O, and benchmarking. It wasn’t actually imitating anything that resembles actual Twitter or even a usable website.
- trishume 4y agoWhich I think I'm perfectly clear about in the blog post. The post is mostly about napkin math systems analysis, which does cover HTTP and HTTPS. I'm now somewhat confident I could implement this if I tried, but it would take many years, the prototype and math is to check whether there's anything that would stop me if I tried and be a fun blog post about what systems are capable of. I've worked on a team building a system to handle millions of messages per second per machine, and spending weeks doing math and building performance prototypes like this is exactly what we did before we built it for real.
- PragmaticPulp 4y ago> Which I think I'm perfectly clear about in the blog post. Of course. I was commenting here to counter all of the comments declaring that this proves Twitter doesn’t need all of their servers, etc. It’s a fun article, but the comments here interpreting it as proving something about Twitter engineers being bad are kind of depressing.
- PerilousD 4y ago[flagged]
- kureikain 4y agoNot to the extreme of fitting everything into one machine but I have explorer the idea of separate stateless workload into its own machine. However, the stateless workload can still operate in a read-only manner if the stateful component failed. I run an email forwarding service[1], and one of challenge is how can I ensure the email forwarding still work even if my primary database failed. And I come up with a design that the app boot up, and load entire routing data from my postgres into its memory data structure, and persisted to local storage. So if postgres datbase failed, as long as I have an instance of those app(which I can run as many as I can), the system continue to work for existing customer. The app use listen/notify to load new data from postgres into its memory. Not exactly the same concept as the artcile, but the idea is that we try to design the system in a way where it can operate fully on a single machine. Another cool thing is that it easiser to test this, instead of loading data from Postgres, it can load from config files, so essentially the core biz logic is isolated into a single machine. --- https://mailwip.com https://mailwip.com
- samsquire 4y agoI recommend this table of latency figures for any software engineer: https://gist.github.com/jboner/2841832 https://gist.github.com/jboner/2841832 Essentially IO is expensive except within a datacenter but even in a data center, you can do a lot of loop iterations in a hot loop in the time it takes to ask a server for something. There is a whitepaper which talks about the raw throughput and performance of single core systems outperforming scalable systems. These should be required reading of those developing distributed systems. http://www.frankmcsherry.org/assets/COST.pdf http://www.frankmcsherry.org/assets/COST.pdf A summary: http://dsrg.pdos.csail.mit.edu/2016/06/26/scalability-cost/ http://dsrg.pdos.csail.mit.edu/2016/06/26/scalability-cost/
- SilverBirch 4y agoI think one of the under-estimated interesting points of twitter as a business is that this is the core. Yes, Twitter is 140 characters, it's got "300m users" which is probably 5m real heavy users. So yes, you could do a lot of "140 characters, a few tweets per person, few million users" on very little hardware. But that's why Twitters a shit business! How much RAM did your advertising network need? Becuase that is what makes twitter a business! How are you building your advertiser profiles? Where are you accounting for fast roll out of a Snapchat/Instagram/BeReal/Tiktok equivalent? Oh look, your 140 characters just turned into a few hundreds megs of video that you're going to transcode 16 different ways for Qos. Ruh Roh! How are your 1,000 engineers going to push their code to production on one machine? Almost always the answer to "do more work" or "buy more machines" is "buy more machines". All I'm saying is I'd change it to "Toy twitter on one machine" not Production.
- reacharavindh 4y agoThe author claimed early on, and very clearly that this was a fun exercise of thought and engineering rather than saying “Look this is how Twitter should be run”. After all this is Hacker News. Such exercises, and engaging other hackers to pick something out of there is how we progress(and get our tickles). So, may be instead think about how one could tackle the advertising/indexing needs in a similar fashion(could it be done in just another server? 5 more servers?)..
- SilverBirch 4y agoYeah I completely get that, but I think a lot of hacker news tends to think of companies as the sum of their engineering resources, rather than what they really are. Which is weird, because it's hosted by YC which is meant to be the polar opposite of that. The point of my comment was that the decision making that lead to Twitter's design actually make sense when you understand the business model behind them. It's not good enough to just run a webpage.
- Nextgrid 4y agoIf it actually takes you a single machine to run this, you don't really need an advertising network to fund it. Out of 5M users (let alone the theoretical 300M) there will be enough people who'd be happy to pay for verification or an exclusive badge on their profile. > How are your 1,000 engineers going to push their code to production on one machine? That might actually be the reason why Twitter barely keeps afloat. 1k engineers for a product that's already built and hasn't fundamentally changed nor evolved in years makes me wonder what business value those engineers actually provide.
- thriftwy 4y agoI remember Stack Overflow running on a single Windows Server box and mocking fellow LAMP developers with their propensity towards having dozens of VMs to same effect. That was some time ago, though.
- Nextgrid 4y agoI believe that's still the case: https://stackexchange.com/performance https://stackexchange.com/performance 9 web servers to serve the entire network. I wish more developers were aware of just how performant modern (non-cloud) hardware is.
- sgt 4y agoSO only has two proxy servers, in which one is a failover? I bet a lot of people wouldn't believe that if you casually mentioned it.
- justapassenger 4y agoSaying this is production Twitter is like saying that rsync is a Dropbox.
- betaby 4y agorsync is way better than Dropbox for large volumes / high speed file transfers . So, yes.
- pengaru 4y agoThis post reminds me of an experience I had in ~2005 while @ Hostway Chicago. Unsolicited story time: Prior to my joining the company Hostway had transitioned from handling all email in a dispersed fashion across shared hosting Linux boxes with sendmail et al, to a centralized "cluster" having disparate horizontally-scaled slices of edge-SMTP servers, delivery servers, POP3 servers, IMAP servers, and spam scanners. That seemed to be their scaling plan anyways. In the middle of this cluster sat a refrigerator sized EMC fileserver for storing the Maildirs. I forget the exact model, but it was quite expensive and exotic for the time, especially for an otherwise run of the mill commodity-PC based hosting company. It was a big shiny expensive black box, and everyone involved seemed to assume it would Just Work and they could keep adding more edge-SMTP/POP/IMAP or delivery servers if those respective services became resource constrained. At some point a pile of additional customers were migrated into this cluster, through an acquisition if memory serves, and things started getting slow/unstable. So they go add more machines to the cluster, and the situation just gets worse. Eventually it got to where every Monday was known as Monday Morning Mail Madness, because all weekend nobody would read their mail. Then come Monday, there's this big accumulation of new unread messages that now needs to be downloaded and either archived or deleted. The more servers they added the more NFS clients they added, and this just increased the ops/sec experienced at the EMC. Instead of improving things they were basically DDoSing their overpriced NFS server by trying to shove more iops down its throat at once. Furthermore, by executing delivery and POP3+IMAP services on separate machines, they were preventing any sharing of buffer caches across these embarrassingly cache-friendly when colocated services. When the delivery servers wrote emails through to the EMC, the emails were also hanging around locally in RAM, and these machines had several gigabytes of RAM - only to never be read from. Then when customers would check their mail, the POP3/IMAP servers always needed to hit the EMC to access new messages, data that was probably sitting uselessly in a delivery server's RAM somewhere. None of this was under my team's purview at the time, but when the castle is burning down every Monday, it becomes an all hands on deck situation. When I ran the rough numbers of what was actually being performed in terms of the amount of real data being delivered and retrieved, it was a trivial amount for a moderately beefy PC to handle at the time. So it seemed like the obvious thing to do was simply colocate the primary services accessing the EMC so they could actually profit from the buffer cache, and shut off most of the cluster. At the time this was POP3 and delivery (smtpd), luckily IMAP hadn't taken off yet. The main barrier to doing this all with one machine was the amount of RAM required, because all the services were built upon classical UNIX style multi-process implementations (courier-pop and courier-smtp IIRC). So in essence the main reason most of this cluster existed was just to have enough RAM for running multiprocess POP and SMTP sessions. What followed was a kamikaze-style developed-in-production conversion of courier-pop and courier-smtp to use pthreads instead of processes by yours truly. After a week or so of sleepless nights we had all the cluster's POP3 and delivery running on a single box with a hot spare. Within a month or so IIRC we had powered down most of the cluster, leaving just spam scanning and edge-SMTP stuff for horizontal scaling, since those didn't touch the EMC. Eventually even the EMC was powered down, in favor of drbd+nfs on more commodity linux boxes w/coraid. According to my old notes it was a Dell 2850 w/8GB RAM we ended up with for the POP3+delivery server and identical hot spare, replacing racks of comparable machines just having less RAM. >300,000 email accounts.
- ricardobeat 4y ago> super high performance tiering RAM+NVMe buffer managers which can access the RAM-cached pages almost as fast as a normal memory access are mostly only detailed and benchmarked in academic papers Isn't this exactly what modern key value stores like RocksDB, LMDB etc are built for?
- z3t4 4y agoA Twitter clone could probably run in a teenagers closet, but not after it has been iterated by 10000 monkeys.
- fleddr 4y ago"Through intense digging I found a researcher who left a notebook public including tweet counts from many years of Twitter’s 10% sampled “Decahose” API and discovered the surprising fact that tweet rate today is around the same as or lower than 2013! Tweet rate peaked in 2014 and then declined before reaching new peaks in the pandemic. Elon recently tweeted the same 500M/day number which matches the Decahose notebook and 2013 blog post, so this seems to be true! Twitter’s active users grew the whole time so I think this reflects a shift from a “posting about your life to your friends” platform to an algorithmic content-consumption platform." I know it's not the core premise of the article, but this is very interesting. I believe that 90% of tweets per day are retweets, which supports the author's conclusion that Twitter is largely about reading and amplifying others. That would leave 50 million "original" tweets per day, which you should probably separate as main tweets and reply tweets. Then there's bots and hardcore tweeters tweeting many times per day, and you'll end up with a very sobering number of actual unique tweeters writing original tweets. I'd say that number would be somewhere in the single digit millions of people. Most of these tweets get zero engagement. It's easy to verify this yourself. Just open up a bunch of rando profiles in a thread and you'll notice a pattern. A symmetrical amount of followers and following typically in the range of 20-200. Individual tweets get no likes, no retweets, no replies, nothing. Literally tweeting into the void. If you'd take away the zero engagement tweets, you'll arrive at what Twitter really is. A cultural network. Not a social network. Not a network of participation. A network of cultural influencers consisting of journalists, politicians, celebrities, companies and a few witty ones that got lucky. That's all it is: some tens of thousands of people tweeting and the rest leeching and responding to it. You could argue that is true for every social network, but I just think it's nowhere this extreme. Twitter is also the only "social" network that failed to (exponentially) grow in a period that you might as well consider the golden age of social networks. A spectacular failure. Musk bought garbage for top dollar. The interesting dynamic is that many Twitter top dogs have an inflated status that cannot be replicated elsewhere. They're kind of stuck. They achieved their status with hot take dunks on others, but that tactic doesn't really work on any other social network.
- yodsanklai 4y ago> Musk bought garbage for top dollar Totally out of topic here, but could be he just wants the ability to amplify his own ideas. Also, why measure Twitter value (arbitrarily?) by number of unique tweets, rather than by read tweets?
- jiggawatts 4y agoSomething I've found a lot modern IT architects seem to ignore is "write amplification" or the equivalent effect for reads. If you have a 1 KB piece of data that you need to send to a customer, ideally that should require less than 1 KB of actual NIC traffic thanks to HTTP compression. If processing that 1 KB takes more than 1 KB of total NIC traffic within and out of your data centre, the you have some level of amplification. Now, for writes, this is often unavoidable because redundancy is pretty much mandatory for availability. Whenever there's a transaction, an amplification factor of 2-3x is assumed for replication, mirroring, or whatever. For reads, good indexing and data structures within a few large boxes (like in the article) can reduce the amplification to just 2-3x as well. The request will likely need to go through a load balancer of some sort, which amplifies it, but that's it. So if you need to process, say, 10 Gbps of egress traffic, you need a total of something like 30 Gbps at least, but 50 Gbps for availability and handling of peaks. What happens in places like Twitter is that they go crazy with the microservices. Every service, every load balancer, every firewall, proxy, envoy, NAT, firewall, and gateway adds to the multiplication factor. Typical Kubernetes or similar setups will have a minimum NIC data amplification of 10x on top of the 2-3x required for replication. Now multiply that by the crazy inefficient JSON-based protocols, the GraphQL, an the other insanity layered on to "modern" development practices. This is how you end up serving 10 Gbps of egress traffic with terabits of internal communications. This is how Twitter apparently "needs" 24 million vCPUs to host text chat. Oh, sorry... text chat with the occasional postage-stamp-sized, potato quality static JPG image.
- threeseed 4y agoYour point makes sense if you have no idea how Twitter works. It needs to assemble tweets internally, sort them with an ML model, add in relevant ads and present a single response to the user because end-user latency matters. And each of these systems eg. ads has their own features, complexities, development lifecycle and scaling requirements. And of course deploying them continuously without downtime. That is how you end up with disparate services and a lot of them for redundancy reasons. I know you think you’re smarter than everyone at Twitter. But those who really know what they are doing have a lot more respect for the engineers who built this insanity. There are always good intentions.
- sitkack 4y agoI am both embarrassed and disappointed with the negativity this post has attracted. A litanny of "gotchas", where someone attempts to best the OP. What about x, y and z? It can't possibly scale. Twitter is so much more than this, etc. The OP isn't making the assertion that Twitter should replace their current system with a single large machine. The whole thread paints a picture of HN like it is full of a bunch of half-educated, uncreative negative brats. To the people that encourage a fun discussion, thank you! Great things are not built by people who only see how something cannot possibly work.
- mcenedella 4y agoAmen!
- Xeoncross 4y agoI must admit, as a bystander I'm torn here. I love that Tristan put out this post and made it so detailed with plenty of assumptions to cover. I also like to hear about possible issues and assumptions which the crowd calls out. Even naysayers can be helpful. I want both, but I don't want to crowd to go to far and kill the desire to produce this kind of content.
- racingmars 4y ago> I want both, but I don't want to crowd to go to far and kill the desire to produce this kind of content. I think it's easy to have both. It's all about the tone of the responses. For example, instead of "your assumptions are wrong, this would collapse because X" or "this is dumb because real Twitter does Y which yours doesn't handle," I think responses could be framed as: "Wow, neat thought experiment! If I were to approach this same problem, I might make an allowance of more than 280 bytes of storage per tweet to allow for additional metadata that is probably needed to make everything work together; I wonder if that can be accommodated with an even beefier big computer?" Or "What a great writeup of building a simplified Twitter! After the features you've accounted for, the next most important feature of Twitter for me personally is Y. What kinds of things would we have to do to stretch your idea to handle that? [or, I bet with the addition of X we could make that happen in this setup too!]" I think many criticisms could be turned into constructive positive additions to the original article versus attacks against the idea of the article.
- knubie 4y ago> I’m not sure how real Twitter works but I think based on Elon’s whiteboard photo and some tweets I’ve seen by Twitter (ex-)employees it seems to be mostly the first approach using fast custom caches/databases and maybe parallelization to make the merge retrievals fast enough. I think Twitter does (or at some point did) use a combination of the first and second approach. The vast majority of tweets used the first approach, but tweets from accounts with a certain threshold of followers used the second approach.
- mgaunard 4y agoAll web and cloud technologies are inherently inefficient, and most programmers don't know networking or even how hardware works sufficiently well to optimize for high througput and low-latency. There was an article just yesterday about how Jane Street had developed an internal exchange way faster than any actual exchange by building it from the ground up, thinking about how the hardware works and how agents can interact with it. Modern software like Slack or Twitter are just reinventing what IRC or BBS did in the past, and those were much leaner, more reliable and snappier than their modern counterparts, even if they didn't run at the same scale. It wouldn't be surprising at all that you could build something equivalent to Twitter on just one beefy machine, maybe two for redundancy.
- therealdrag0 4y agoIt’s “easy” to optimize for speed when you build from the ground up with no real customers or feature requirements that you can’t just conveniently ignore.
- deleted 4y ago[deleted]
- drowsspa 4y agoI'm always reminded of that tweet from @SwiftOnSecurity https://mobile.twitter.com/SwiftOnSecurity/status/1485822027923603456 https://mobile.twitter.com/SwiftOnSecurity/status/1485822027... > Once you understand your computer has 16 cores running at 3GHz and yet doesn't boot up in .2 nanoseconds you understand everything they have taken from you. With their infinite VC money at their disposal, and with their programmers having 100 GHz machines with thousands of cores, 128 TB of RAM and FTL internet connections, tech companies don't really have any incentive to actually reduce bloat. Edit: it's still quite sad. I feel like we had languages with a way better future, and more promising programming architectures, back in the 80s.
- kaba0 4y agoBooting has fundamental hardware limits, though. Will that harddrive or whatever really reply all in 0.2 ns?
- viraptor 4y agoThis is in no way a criticism of the analysis. But what I think is a hidden cost of an idea like this (that hasn't been pointed out) is the ability to extend the features. With a tightly integrated system like that you may want to add a frobnicator as a test - now that whole system would need to change to accommodate that, because all the timeline processing happens more or less in memory. Making things external / network based adds overhead, but makes plugging in / removing an extra feature much easier. If you count the cost of work required for making changes, then burning money on the "unnecessary" horizontal scaling may not be a bad idea. Wanna add new ads analytics? Just plug into this common firehouse/summary endpoint without worrying about the internals. Wanna test a new implementation of some component? Run both in parallel, plugging into same inputs. Etc.
- betimsl 4y agoI'm just curious as to what kind of motherboard this personal computer is going to have? I'm asking this because of the limit on PCIe bandwidth. 100gbit NIC? How?
- tpmx 4y agoServing Netflix Video Traffic at 800Gb/s and Beyond [pdf] https://news.ycombinator.com/item?id=32519881 https://news.ycombinator.com/item?id=32519881
- betaby 4y agoI have some servers in production with https://www.nvidia.com/en-us/networking/ethernet/connectx-6/ https://www.nvidia.com/en-us/networking/ethernet/connectx-6/ There is a an AUX card to load PCI lanes properly, like on this picture https://lenovopress.lenovo.com/assets/images/LP1195/Mellanox%20ConnectX-6%20HDR%20QSFP56%201-port.jpg https://lenovopress.lenovo.com/assets/images/LP1195/Mellanox...
- siliconc0w 4y agoEnjoyed the write up, would be curious to see the twitter spend broken down by functionality given all the extra stuff they do. I imagine it's a non-linear relationship where the company has to burn more and more cash with every new feature (and esp things like Advertising which you need once your spend surpasses what a simple subscription can offer), more scale adds more complexity, bureaucracy and overhead (management, hr & recruiting, legal&accounting, etc). While it's likely there is waste (some of which is inevitable, see 'overhead' above) a super bareboens twitter can maybe run within one beefy machine but a 'real' twitter ends up needing millions + lots of people.
- lightlyused 4y agoIt is all fun running on one machine until a capacitor leaks or something else goes south.
- kissgyorgy 4y agoI feel like people writing posts like this never worked in a big team at a big company on a big project. It is so obviously impossible to do this and Twitter has so many more features users will never even see, but sure, re-implement it in a couple hundred lines of Rust and Twitter will be saved...
- LastTrain 4y agoEvery intern proposes something cute like this in my org, every year.
- threeseed 4y agoHN unfortunately has a lot of people like George Hotz. They are knowledgeable to a certain level but they simply aren't great engineers who almost always are humble, cautious, thoughtful and respectful of the intentions behind what other engineers build. Anyone who thinks they can jump in and replace any tech stack without an extensive deep dive of the business requirements, design decisions, cost constraints, resource limitations etc that drove the choices deserves the pain and unemployment that inevitably follows.
- TacticalCoder 4y agoTFA, to me, touches about something I've wondered about a very long time ago: what are the implications of CPU and storage growing at much faster rates than human population? Back in the 486 days you wouldn't be keeping, in RAM, data about every single human on earth (let's take "every single human on earth" as the maximum number of humans we'll offer our services to with on our hypothetical server). Nowadays keeping in RAM, say, the GPS coordinates of every single human on earth (if we had a mean to fetch the data) is doable. On my desktop. In RAM. I still don't know what the implications are. But I can keep the coordinates of every single humans on earth in my desktop's RAM. Let that sink in. P.S: no need to nitpick if it's actually doable on my desktop today. That's not the point. If it's not doable today, it'll be doable tomorrow.
- moeny 4y ago's/(human|desktop)/smartphone/g'
- untech 4y agoUsing two 32 bit numbers for coordinates, each record would take 8 bytes, which is 64 gigabytes for 8 billion population. Don’t think many smartphones have this RAM today.
- cellis 4y agoDon't you need 3 numbers? Unless you believe in a flat earth ;). Also you need some slack space for metadata. Let's call it 100GB all in.
- toast0 4y agoLat long and assuming at ground gets you most of the way there.
- Dylan16807 4y ago"Add it all up, and the US has around 340 billion square feet of building stock[3]. This is about 12,200 square miles, or 0.00032% of US land area" Judging by that, you need a negligible increase in the number of locations you can represent to handle everywhere stable and off the ground someone could be. Much less than one bit per person. If you want to deal with people currently in airplanes then you could give them an extra couple bytes. It's less than a million people so it won't affect your total storage at all.
- gravypod 4y agoAs a side note: > I did all my calculations for this project using Calca (which is great although buggy, laggy and unmaintained. I might switch to Soulver) and I’ll be including all calculations as snippets from my calculation notebook. I've always wanted an {open source, stable, unit-aware} version of something like this which could be run locally or in the browser (with persistence on a server). I have yet to find one. This would be a massive help to anyone who does systems design.
- swellguy 4y agoYou have violated the number one rule in Silicon Valley: If it doesn't take at least "N" "engineers" to "solve" a problem who report directly to moi, then how am I relevant? So I agree this is entirely possible, but no one would build this with any funding.
- VLM 4y agoInteresting optimization idea: If 95% of the users are bots, and your ML algorithms are smart enough to figure who's bot and who's bio, you could save a lot of traffic by not publishing tweets to bots as no one is going to read them anyway. Of course if that traffic included advertisements, you'd also lose 95% of your ad revenue. The ultimate extension of this "run it all on one machine" meme would be to run the bots on the single machine along with the service.
- Hasu 4y ago> Of course if that traffic included advertisements, you'd also lose 95% of your ad revenue. Not so, serving an ad to a bot gains you no revenue, because ad networks charge for clicks, not impressions. If a significant percentage of your ad clicks are from bots, you're running a defective advertising platform and won't have customers for long regardless.
- toteno 4y agoI agree with you overall. Just a small nitpick: most ad networks optimize for price of impression, so at the end of the day they charge for impressions (just not always directly). If your ad has low click rate and average price then it just won't be shown, because it's more profitable for an ad network to show ad with better click rate or with better price (i.e. with better price for impression) .
- mattbillenstein 4y agoInteresting abstract system design type problem - I think it becomes difficult if you have to shard the data though because all the assumptions about the hot set being in RAM break all the performance guarantees now I think... Which is I think basically what Twitter's existing backend has to deal with.
- 0837560358 4y ago[flagged]
- jacobsenscott 4y agoNice. While there may be some impracticalities to actually doing this for twitter, 99% of the software out there could run on a fraction of a single commodity server. People complain about the carbon burn of crypto, and they are right, but I bet it is dwarfed by the carbon burn of all the shitty over-provisioned and over-architected CRUD apps running interpreted languages. Unfortunately with universities teaching python of all things we'll have (or maybe already do have) a whole generation of developers that actually have no idea how powerful a modern computer is. I suppose there's a chance AI will get to the point where we can feed it a ruby/python/js/whatever code base and it can emit the functionally equivalent machine code as a single binary (even a microservices mess).
- deleted 4y ago[deleted]
- etaioinshrdlu 4y agoI've been experimenting with GPT3 as a "compiler" and it's always amazing when it works. Extrapolating from here, I think this is right on the nose -- we feed in any high level language and get reasonably efficient assembly out the other end, furthermore, the assembly is more human-like than a compiler would give you. There's some big problems with this approach today, namely, it's not always right, and it may sometimes be half right (miss edge cases). But think of where this AI technology is headed -- it stands to reason it will eventually work pretty much perfect. And then I think we'll see another very strong trend -- large AI models replacing other forms of software. Why write a compiler when GPT3 can compile C to asm? Why write an interpreter when GPT3 can "compile" python to C? The AI model is hilariously less-efficient than traditional software, but it will be far far cheaper and faster to create than the traditional equivalent. What other types of software will be replaced by AI models?
- jbotdev 4y agoI think people overestimate how much CPU time a typical CRUD app spends on actual business logic, even with an interpreted language like Ruby or Python. What I’ve seen is the bottleneck is largely memory, such that you can pack a ton of these apps on a machine with a few cores and a lot of RAM. The stuff that actually is CPU-bound often ends up being written in an appropriate language, or uses C extensions (e.g. ML and data science libraries for Python).
- bitbckt 4y agoRegarding tweet distribution, I was one of the folks who built the first scalable solution to this problem at Twitter (called Haplocheirus). We used the Yahoo “Feeding Frenzy” design, pushing tweets through a Redis-backed caching layer. Feel free to continue using that (historically-correct) answer in interviews. :P
- deleted 4y ago[deleted]
- snotrockets 4y agoI ask candidates I interview to design a certain service. Most ask about scale, to which I like to direct the question back at them: it's going to be huge. As big as Twitter. How big would that be, do you think? Most then suggest scale that would make the service run comfortable from a not-too powerful machine, and then go to design data-center spanning distributed service.
- jameshart 4y agoGetting everything onto one machine works great until... it no longer fits on one machine. You add another feature and it requires a little bit more RAM, and another feature that needs a little bit more, and.. eventually it doesn't all fit. Now you have to go distributed. And your entire system architecture and all your development approaches are built around the assumptions of locality and cache line optimization and all of a sudden none of that matters any more. Or you accept that there's a hard ceiling on what your system will ever be able to do. This is like building a video game that pushes a specific generation of console hardware to its limit - fantastic! You got it to do realtime shadows and 100 simultaneous NPCs on screen! But when the level designer asks if they can have water in one level you have to say 'no', there's no room to add screenspace reflections, the console can't handle that as well. And that's just a compromise you have to make, and ship the game with the best set of features you can cram into that specific hardware. You certainly could build server applications that way. But it feels like there's something fundamental to how service businesses operate that pushes away from that kind of hyperoptimized model.
- eismcc 4y agoYou’d end up synchronizing feature releases to Moores law. Which while it sounds untenable there are large corporations that continue to use monolithic approach and vertical scaling.
- nyanpasu64 4y agoVideo games have had 3D water for decades before screen-space reflections, and many look serviceable decades later (Super Mario Sunshine looks great at 480p though dated at higher resolutions).
- jameshart 4y agoCurses. That undermines my argument completely. It wasn’t merely a random example of how you have to compromise on features to fit in the available hardware budget, but actually the premise upon which my entire argument rested.
- joshspankit 4y ago
- mizzao 4y agoWhile you might not want to do this with actual Twitter, any sort of high-performance computing workload can run substantially faster on a single optimized machine than on a distributed computing environment. I learned this the hard way when I was running a medium-sized MapReduce job in grad school that was over 100x faster when run as a local direct computation with some numerical optimizations.
- robluxus 4y agoObligatory "you don't need Hadoop": https://news.ycombinator.com/item?id=14401399 https://news.ycombinator.com/item?id=14401399
- truth_seeker 4y agoHow to subdue the cost of abstraction. Very well explained !
- spullara 4y agoWhen I was working there I implemented my patent during a hack week (given a set of follows return the list of matching tweet ids, very similar to his prototype): https://patents.google.com/patent/US20120136905A1/en https://patents.google.com/patent/US20120136905A1/en (licensed under Innovators Patent Agreement, https://github.com/twitter/innovators-patent-agreement https://github.com/twitter/innovators-patent-agreement) I could have definitely served all the chronological timeline requests on a normal server with lower latency that the 1.1 home timeline API. There are a bunch of numbers in the calculations that he is doing that are off but not by an order of magnitude. The big issue is that since I left back then Twitter has added ML ads, ML timeline and other features that make current Twitter much harder to fit on a machine than 2013 Twitter.
- steviesands 4y agoA few thoughts. The first is, are we asking the wrong questions? Should it be, "If I spend 10m on hardware for predicting ads (storage/compute) that generates 25m in revenue, should I buy the hardware?". Sure, we can "minify" twitter, and it's a wonderful thought experiment, but it seems devoid of the context of revenue generation. The second is, it's interesting to understand social media industry wide infra cost per user. If you look at FB, Snap, etc. they are within all within an order of magnitude in cost per DAU (DAU / Cost of revenue) of each other. This can be verified via 10-ks which show Twitter at $1.4B vs. SNAP 1.7B Cost of Revenue. The major difference between the platforms is revenue per user, with FB being the notable exception. Also would you summarize the patent/architecture? The link is a bit opaque/hard to read. Note: Cost of Revenue does also include TAC and revenue sharing (IIRC) and not just Infra costs but in theory they would also be at similar levels. eg. SNAPs 10-k https://d18rn0p25nwr6d.cloudfront.net/CIK-0001564408/da8288aa-d492-4fd1-b6ce-62dce8206e9f.pdf https://d18rn0p25nwr6d.cloudfront.net/CIK-0001564408/da8288a...
- spullara 4y agoThe basic idea of the system was to scan a reverse chronologically ordered list of "user id, tweet id", filtering out any tweet whose user wasn't in the follow set (or sets in the case of scan sharing) until you retrieved enough tweets for the timeline request. There are a bunch of variants in the patent, but that is the basic idea. At the time, I estimated that Twitter was spending 80% of its CPU time in the DC doing thrift/json/html serialization/deserialization and mused about merging all the separate services into a single process. Lot's of opportunity for optimization.
- castratikron 4y agoDoes 4chan fit on one machine?
- henning 4y agoNo rate limiting, API data, quote tweets, view count, threads, likes, mentions, notifications, ads, video, images, account blocking (permanent or TTL), account muting (permanent or TTL), word filtering (permanent or TTL), moderation/reporting, user profile storage, or the fact that tweets that display show more than just the tweet itself. No mention that tweet activity all occurs concurrently and therefore the loading script is not at all a realistic estimate of real activity. But sure, go ahead and take this as evidence that 10 people could build Twitter as I'm sure that's what will happen to this post. If that's true, why haven't they already done so? It should only take a couple weeks and one beefy machine, right?
- systemvoltage 4y agoIt's worth noting Stackoverflow's production arch: https://stackexchange.com/performance https://stackexchange.com/performance
- deleted 4y ago[deleted]
- summerlight 4y agoI think many people in this thread are making the mistake of ignoring evolutionary factors in system engineering. If a system doesn't need to adopt or change, lots of things can be much more efficient, easier and simpler, likely the order of 10x~100x. But you gotta appreciate that we're all paid because we need to swap wheels on running trains (or even engines in flying airplanes). A large fraction of demand for redundancy, introspection, abstraction and generalization comes from this. Why do we want to apply ML at the cost of a significant fleet cost increase? Because it can make the overall system consistently perform against external changes via generalization, thus the system can evolve more cheaply. Why do we want to implement a complex logging layer although it doesn't bring direct gains on system performance? Because you need to inspect the system to understand its behavior and find out where it needs to change. The list can go on and I can give you hundreds of reasons why we need all these apparently unnecessary complexities and overheads can be important for systems' longevity. I don't deny the existence of accidental complexities (probably Twitter can become 2~3x simpler and cheaper given sufficient eng resource and time), but in many cases you probably won't be able to confidently say if some overheads are accidental or essential since system engineering is essentially a highly predictive/speculative activity. To make this happen, you gotta have a precise understanding of how the system "currently works" to make a good bet rather than re-imagination of the system with your own wish list of how the system "should work". There's a certain value on the latter option, but it's usually more constructive to build an alternative rather than complaining about the existing system. This post is great since the author actually tried to build something to prove its possibility, this knowledge could turn out to be valuable for other Twitter alternatives later on.
- ilyt 4y ago> A large fraction of demand for redundancy, introspection, abstraction and generalization comes from this. Sure, you need to invest into it but those are things you can reuse for every app and feature you build. And those are not the reason why those systems are so complex, those are just ways to keep complex systems running and manageable. In most they also do not stand in the way of making system better but help in it. They need to exist because the architecture of system grew organically from smaller system over and over again and big restructurization was deemed not worth it. It's "just have a bunch more hardware and engineers" vs "we're not delivering features and we might not get rewrite right". And every time you throw money at the problem the problem becomes a bigger problem and potential benefits from "getting it right" are also getting bigger. But nobody wants to be herald that tells management "we 're going to spend 6-12 months" for somethinkg that have few years of pay-off
- irq 4y agoExcellent article! I wish the font size on mobile was bigger.
- KaiserPro 4y ago> A friend points out that IBM Z mainframes have a bunch of the resiliency software and hardware infrastructure I mention, Sure its expensive, and you have to deal with IBM, who are either domain experts or mouth breathers. Sure it'll cost you $2m but! the opex of running a team of 20 engineers is pretty huge. Especially as most of the hard bits of redundant multi-machine scaling are solved for you by the mainframe. Redundancy comes for free(well not free, because you are paying for it in hardware/software) Plus, IBM redbooks are the golden standard of documentation. Just look at this: https://www.redbooks.ibm.com/redbooks/pdfs/sg248254.pdf https://www.redbooks.ibm.com/redbooks/pdfs/sg248254.pdf its the redbook for GPFS (scalable multi-machine filesystem, think ZFS but with a bunch more hooks.) Once you've read that, you'll know enough to look after a cluster of storage.
- jideel 4y ago[dead]
- Halan 4y agoLet’s hope Elon doesn’t read this
- firstSpeaker 4y agoThis is one of the most interesting part of the whole post for me: Through intense digging I found a researcher who left a notebook public including tweet counts from many years of Twitter’s 10% sampled “Decahose” API and discovered the surprising fact that tweet rate today is around the same as or lower than 2013! Tweet rate peaked in 2014 and then declined before reaching new peaks in the pandemic. Elon recently tweeted the same 500M/day number which matches the Decahose notebook and 2013 blog post, so this seems to be true! Twitter’s active users grew the whole time so I think this reflects a shift from a “posting about your life to your friends” platform to an algorithmic content-consumption platform. So, the number of writes has been the same for a good long while.
- fortran77 4y agoI've thought about this problem, too, and blocklists seem like a hard problem to implement efficiently. I have a few thousand users blocked, and several hundred keywords, phrases and emoji. How are these processed efficiently?
- aetimmes 4y ago(Disclaimer: ex-Twitter SRE) > There’s a bunch of other basic features of Twitter like user timelines, DMs, likes and replies to a tweet, which I’m not investigating because I’m guessing they won’t be the bottlenecks. Each of these can, in fact, become their own bottlenecks. Likes in particular are tricky because they change the nature of the tweet struct (at least in the manner OP has implemented it) from WORM to write-many, read-many, and once you do that, locking (even with futexes or fast atomics) becomes the constraining performance factor. Even with atomic increment instructions and a multi-threaded process model, many concurrent requests for the same piece of mutable data will begin to resemble serial accesses - and while your threads are waiting for their turn to increment the like counter by 1, traffic is piling up behind them in your network queues, which causes your throughput to plummet and your latency to skyrocket. OP also overly focuses on throughput in his benchmarks, IMO. I'd be interested to see the p50/p99 latency of the requests graphed against throughput - as you approach the throughput limit of an RPC system, average and tail latency begin to increase sharply. Clients are going to have timeout thresholds, and if you can't serve the vast majority of traffic in under that threshold consistently (while accounting for the traffic patterns of viral tweets I mentioned above) then you're going to create your own thundering herd - except you won't have other machines to offload the traffic to.
- MichaelZuo 4y agoWhat do you think about his interesting comment on the possibility of a mainframe? "I also didn’t try to investigate configuring an IBM mainframe, which stands a chance of being the one type of “machine” where you might be able to attach enough storage to fit historical images." It seems theoretically possible it could accomodate the entirety of Twitter in 'one machine'.
- aetimmes 4y agoIt depends on what you (or OP) mean by "one machine". There was a HPC cluster at Princeton when I worked there (which, looking at their website, has since been retired) that was assembled by SGI and outfitted with a customized Linux unikernel that presented itself as a single OS image, despite being comprised several disparate racks of individual 2-4u servers. You might be able to metaphorically duct-tape enough machines together with a similar technique to be able to run the author's pared-down scope within a single OS image. With respect to the IBM z-series specifically - if the goal of the exercise is to save money on hardware costs, I'm imagining purchasing an IBM mainframe is in direct opposition to that goal. :) I'm not familiar enough with its capabilities to say one way or the other.
- mcqueenjordan 4y agoFun thought experiment! I can't help but be reminded of the Good Will Hunting quote, though: SEAN: So if I asked you about art you’d probably give me the skinny on every art book ever written. Michelangelo? You know a lot about him. Life’s work, political aspirations, him and the pope, sexual orientation, the whole works, right? But I bet you can’t tell me what it smells like in the Sistine Chapel. You’ve never actually stood there and looked up at that beautiful ceiling. Seen that.
- surume 4y agoOps: "One of our instances went down" Everyone else: "Gaaaahhh"
- sammy2255 4y agoA bit out of touch to think that the bandwidth alliance will let you push 500TB a month through them for free