5 ms·
Actually, when done right this dramatically simplifies a backend architecture. Even a low-scale application commonly uses multiple databases (e.g. Postgres plus
by nathanmarz 3y ago
Actually, when done right this dramatically simplifies a backend architecture. Even a low-scale application commonly uses multiple databases (e.g. Postgres plus ElasticSearch) and queues+workers for background work.
Our Twitter-scale Mastodon implementation is a direct demonstration of this. It's literally 100x less code than Twitter had to write to build the equivalent feature-set at scale, and it's more than 40% less code than Mastodon's official implementation. This isn't because of being able to design things better with the same tooling the second time around – it's because it's built using fundamentally better abstractions.
- vidarh 3y agoThe thing is, the abstractions you're offering is sugar on top of a well-known, well-understood old architectural pattern. The Twitter / Mastodon comparison is unconvincing exactly because both Twitter and Mastodon seemingly totally ignored the well-known, well-trodden ground on this, for no good reason. If you can convince more developers to apply what is effectively canonical store -> workers -> partitioned denormalized materialized views as a pattern where it makes sense, then great. But you can do that with just the tools people already have available. Heck, you can do that with just multiple postgres servers (as the depot, and for the "p-stores" and for the indexing functions), and then you don't need to ditch the languages people are familiar with for both specifying the materialization and the queries. Part of the reason we use a "hodgepodge of narrow tooling" however, tends to be that it allows us to pick and choose languages, and APIs, depending on what developers are familiar with and what suits us, and it allows people to pick and mix. Convincing people to give up that in favour of a fairly arcane-looking API restricted to the JVM is going to be a tough sell a lot of places.
- graemep 3y agoIf you complimented Mastodon why are you comparing with Twitter? Why not compare with Mastodon or other ActivityPub implementations? Does your implementation match the features of Mastodon?
- nathanmarz 3y agoMastodon is extremely similar to the Twitter consumer product. The differences are minor. We compare with Twitter because unlike Mastodon Twitter runs at large scale. Yes, we implemented the entirety of Mastodon from scratch.
- saberience 3y agoEvent-Sourcing is just a bad idea, period. It adds more complexity, more cognitive load, and results in more bugs and more downtime.
- nathanmarz 3y agoA good idea with a flawed implementation is still a good idea
- jddj 3y ago> Fundamentally, [the things you deal with] are materialized views over an event sourced log. This is also true of most of the main SQL database engines, no?
- yowlingcat 3y agoTwitter is a simple product that only has value because of the eyeballs that look at it, not the depth of the product. I fundamentally disagree with the premise of your blog post but not the premise of your company, so why don't I make an ask for you to write a much more practically useful application. Design a basic shopping cart using your system and compare and contrast it with a well designed relational equivalent. A system that allows products to be purchased and fulfilled is a far closer match to what the majority of companies are using software for compared to writing a Twitter clone. Here's my take -- at a certain point in scale and volume using a database, it actually does make sense to rewrite all of the following from scratch: - query planning - indexing (btree vs GIN, etc) and primary/foreign/unique key structure - persistence layer - locking - enforcing ACID constraints - schema migrations/DDL - security, accounts and permissioning - encryption primitives But more crucially, my belief and experience is that most companies making lots of money from software products will NEVER remotely reach that scale -- and prematurely optimizing for that is not just the wrong decision but borderline professional malpractice. You can get very far with Postgres and JSONB if you really need it, and you'll spend more time focusing on your business logic than reinventing the wheel. I'd like to be proven wrong. But I get sinking feeling that I'm not wrong, and while your product is potentially valuable for a very specific use case, you're doing your own company a disservice by distorting reality so strongly to both yourselves and your prospective customers. I'll round out this comment by linking another comment in this thread that goes very well into the perils of event sourcing when the juice isn't worth the squeeze: https://news.ycombinator.com/item?id=38930591 https://news.ycombinator.com/item?id=38930591
- nathanmarz 3y agoDifferent applications will be more relatable to different developers. And that is the path we are going down, of steadily building up more and more examples of applying Rama towards different use cases. Some developers will get their light-bulb moment on a compare/contrast to a shopping cart, others vs. a time-series analytics app, etc. It's completely different for different developers, so building up that library of examples will take time. At the moment, we're focused on our private beta users who are technically savvy enough to be able to understand Rama through the first principles on which it's based. We started with a Twitter demonstration because: a) its implementation at scale is extremely difficult, b) I used to work there and am intimately familiar with what they went through on the technical end, and c) the product is composed of tons of use cases which work completely differently from each other – social graph, timelines, personalized follow suggestions, trends, search, etc. A single platform able to implement such diverse use cases with optimal performance, at scale, and in a comparatively tiny amount of code is simply unprecedented.
- groovecoder 3y agoLook I don't have any reason to praise Twitter, but ... This "Twitter-scale mastadon implementation" is when my red flags went up. It's meant to demonstrate a simpler and more performant architecture, but it actually demonstrates "things you should never do" #1: rewrite the code from scratch. https://www.joelonsoftware.com/2000/04/06/things-you-should-never-do-part-i/ https://www.joelonsoftware.com/2000/04/06/things-you-should-... "The idea that new code is better than old is patently absurd. Old code has been used. It has been tested. Lots of bugs have been found, and they’ve been fixed." The "1M lines of code" and "~200 person-years" of Twitter being trashed on in this article is the outcome of Twitter doing the most important thing that software should do: deliver value to people. Millions of people (real people, not 100M bots) suffered thru YEARS of the fail-whale because Twitter's software gave them value. This software has only delivered some artificial numbers in a completely made-up false comparison. Okay it's built on "fundamentally better abstractions", but until it's running for people in the real world, that's all it is: abstract. Please don't tout this as a demonstration of how to re-create all of Twitter with simpler and more performant back-end architecture.
- fl0ki 3y agoThis is a well-known essay that everyone should read, and yet nobody should ever cite it as a commandment for what other engineers can or cannot do. Joel was talking about commercial desktop software in the extremely competitive landscape of the 90s, he wasn't talking about world-scale internet service infrastructure. The architecture that delivered a set of X features to the initial N users isn't always going to be enough for X+Y features to the eventual 1000*N users that you promised your investors. Companies like Google are quite public about how much they rewrite internal software, and that's just what the public hears about. A particular service might have been written with all of the care and optimization that a world-class principal engineer can manage, serve perfectly for a particular workload for several years, and yet still need to be entirely replaced to keep up with a new workload for the following several years. You wouldn't tell another engineer that they shouldn't rewrite a single function because software should never be rewritten, so it also doesn't make sense to tell them not to rewrite an entire project either. It's their call based on the requirements and constraints that they know much more about. Nobody should be rushing out to rewrite Linux or LLVM from scratch, and yet we wouldn't even have Linux or LLVM if their developers didn't find reasons to create them even while other projects existed. In hindsight it's clear those projects needed to be created, but at the time people would have said you should never rewrite a kernel or compiler suite.