9 ms·
Scaling Rails and Postgres to millions of users at Microsoft
- cdiamand 2y agoI ran into some scaling challenges with Postgres a few years ago and had to dive into the docs. While I was mostly living out of the "High Availability, Load Balancing, and Replication" chapter, I couldn't help but poke around and found the docs to be excellent in general. Highly recommend checking them out. https://www.postgresql.org/docs/16/index.html https://www.postgresql.org/docs/16/index.html
- jbverschoor 2y agoLike many of the BSDs
- aerzen 2y agoDid Postgres used to be a BSD? Are they known for good documentation?
- password4321 2y agoBSD? No, that's operating system(s) Good documentation? Yes
- andrewf 2y agoBSD was the Unix distribution; BSD and Postgres/Ingres development did overlap at UC Berkeley.
- danpalmer 2y agoThey are excellent! Another great example is the Django project, which I always point to for how to write and structure great technical documentation. Working with Django/Postgres is such a nice combo and the standards of documentation and community are a huge part of that.
- irjustin 2y agoInterestingly I have had almost the exact opposite experience being very frustrated with the Django docs. To be fair, it could be because I'm frustrated with Django's design decisions having come from Rails. When learning Django a few years ago, I still carry a deep loathing against polymorphism (generic relations[0]), and model validations (full clean[1]), You know what - it's design decisions... [0] https://docs.djangoproject.com/en/5.1/ref/contrib/contenttypes/ https://docs.djangoproject.com/en/5.1/ref/contrib/contenttyp... [1] https://docs.djangoproject.com/en/5.1/ref/models/instances/#validating-objects https://docs.djangoproject.com/en/5.1/ref/models/instances/#...
- rtpg 2y agogeneric relations are hard to get right, really if you can avoid using them you're going to avoid a lot of trickiness. When you need them... it's nice to have them "just there", implemented correctly (at least as correctly as they can be in an entirely generic way). Model validations is a whole thing... I think that Django offering a built-in auto-generated admin leads to a whole slew of differing decisions that end up coming back to be really tricky to handle.
- globular-toast 2y agoWould love to hear more about what you don't like with model validations (full clean).
- irjustin 2y agoSorry on the slow reply. But yea, I can complain at length. - Model validations aren't run automatically. Need to call full_clean manually. - EXCEPT when you're in a form! Forms have their own clean, which IS run automatically because is_valid() is run. - This also happens to run the model's full_clean. - DRF has its own version of create which is separate and also does not run full_clean. - Validation errors in DRF's Serializers are a separate class of errors from model validations and thus model Val Errors are not handled automatically. - Can't monkey patch models.Model.save to run full_clean automatically for because it breaks some models like User AND now it would run twice for Forms+Model[0]. Because of some very old web-forum style design decisions, model validations aren't unified thus the fragmentation makes you need to know whether you're calling .save()/.create() manually, are in a form, or in DRF. And it's been requested to change this behavior but it breaks backwards compat[0]. It's frustrating because in Rails this is a solved problem. Model validations ALWAYS run (and only once) because... I'm validating the model. Model validations == data validations which means it should be true for all areas regardless of caller, except in exceptions, then I should be required to be explicit when skipping (i.e. Rails) where as in Django I need to be explicit in running it - sometimes... depends where I am. [0] https://stackoverflow.com/questions/4441539/why-doesnt-djangos-model-save-call-full-clean https://stackoverflow.com/questions/4441539/why-doesnt-djang...
- pajeets 2y agoPostgres can be scaled vertically like Stackoverflow did. With cache on edge for popular reads if you absolutely must (but you most likely dont). No need to microservice or sync read replicas even (unless you are making a game). No load balancers. Just up the RAM and CPU up to TB levels for heavy real world apps (99% of you wont ever run into this issue) Seriously its so create scalable backend services with postgrest, rpc, triggers, v8, even queues now all in Postgres. You dont even need cloud. Even a mildly RAM'd VPS will do for most apps. got rid of redis, kubernetes, rabbitmq, bunch of SaaS tools. I just do everything on Postgres and scale vertically. One server. No serverless. No microservice or load handlers. It's sooo easy.
- seabrookmx 2y ago> One server What happens if this server dies?
- wongarsu 2y agoThen your service is offline until you fix it. For many services a completely acceptable thing to happen once in a blue moon Most would probably get two servers with a simple failover strategy. But on the other hand servers rarely die. At the scale of a datacenter it happens often, but if you have like six of them, buy server grade stuff and replace them every 3-5 years chances you won't experience any hardware issues
- pajeets 2y agoif you cant risk this rarity then get a failover server with equal specs maybe add another for good measure....if the biz insurance needs extreme HA then absolutely have multiple failover my point is you arent doing extreme orchestration or routing throw a cloudflare ddos protection too
- mr_toad 2y agoStack overflow absolutely had load balancers, and 9 web servers, and Redis caches. They also use 4 SQL servers, so not entirely vertical either. And they were only serving 500 requests a second on average (peak was probably higher).
- djaouen 2y agoI don't understand why you wouldn't just use Elixir/Phoenix if you need to scale?
- seabrookmx 2y agoI don't understand why you wouldn't use <compiled language that's faster than the BEAM> if you need to scale? /s
- djaouen 2y agoI mean, you could, but you'd be missing out on the Rails-esque nature of Elixir/Phoenix.
- foundart 2y agoPerhaps because you need to scale quickly and already have a large Rails app that would take a long time to recreate in another language and framework.
- SkyPuncher 2y agoIt’s hard to compete with Rails productivity
- giovannibonetti 2y agoWhat a small world. Earlier today I got tagged in a PR [1] where Andrew became the maintainer of a Ruby gem related to database migrations. Good to know he is involved in multiple projects in this space. [1] https://github.com/lfittl/activerecord-clean-db-structure/issues/33 https://github.com/lfittl/activerecord-clean-db-structure/is...
- andatki 2y agoHi there! That's funny! This interview and those gem updates were unrelated. However both are part of the sweet spot for me of education, advocacy, and technical solutions for PostgreSQL and Ruby on Rails apps. I hope you’re able to check out the podcast episode and enjoy it. Thanks for weighing in within the gem comments, and for commenting here on this connection. :)
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- benwilber0 2y agoPostgres can scale to millions of users, but Rails definitely can't. Unless you're prepared to spend a ton of money.
- petcat 2y agoFor real. Show me a company that has scaled RoR or Django to 1 million concurrent users without blowing $250,000/month on their AWS bill. I've worked at unicorn companies trying to do exactly that. Their baseline was 800 instances of the Rails app...lol. I'm not going to name-names (you've heard of them) ... but this is a company that had to invent an entirely new and novel deployment process in order to get new code onto the massive beast of Rails servers within a finite amount of time.
- ainiriand 2y agoWe use 5 ec2 instances to serve around 32 million requests per day on PHP, all under 100ms. It is not the language.
- nov21b 2y agoThe language/runtime certainly has an impact. But indeed, in reality there is no way to compare these scaling claims. For all we know people are talking about serving a http-level cache without even hitting the runtime.
- ainiriand 2y agoEach and every request reach the DB and/or Redis. MyISAM is deprecated, but is crazy fast if you mainly read.
- charlie0 2y agoFramework or custom app?
- ainiriand 2y ago
- datadeft 2y agoScaling a non-scalabe by default framework that should have been few services written in a performance first language at a billion+ USD company. I am not sure why are we boliling the oceans for the sake of a language like Ruby and a framework like Rails. I love those to death but Amazons approach is much better (or it used to be): you can't make a service for 10.000+ users in anything else than: C++, Java (probably Rust as well nowadays). For millions of users the CPU cost difference probably justifies the rewrite cost.
- ainiriand 2y agoYou really do not know what you are talking about, it is not about the language, like it was repeated in this forum many many times already. We serve an application in PHP to thousands of users per second in less than 100ms constantly.
- hamandcheese 2y agoSometimes it is the language. Or at least the ecosystem and libraries available. My go-to example is graphql-ruby, which really chokes serializing complex object graphs (or did, it's been a while now since I've had to use it). It is pretty easy to consume 100s of ms purely on compute to serialize a complex graphql response.
- ainiriand 2y agoI would make a case that that's not the language's fault. You need to assess how critical is speed in your requirements and adapt your solutions.
- Lio 2y agoI have mixed feelings about this. It's saying that python is too slow for data science ignoring that python can outsource that work to Pandas or NumPy. For GraphQL on Rails you can avoid graphql-ruby and use Agoo[1] instead so that that work is outsourced to C. So in practice it's not a problem. 1. https://github.com/ohler55/agoo https://github.com/ohler55/agoo
- teleforce 2y agoPlease check this excellent book by former Microsoft and Groupon engineer on scaling Rails and Postgres: [1] High Performance PostgreSQL for Rails Reliable, Scalable, Maintainable Database Applications by Andrew Atkinson: https://pragprog.com/titles/aapsql/high-performance-postgresql-for-rails/ https://pragprog.com/titles/aapsql/high-performance-postgres...
- jojobas 2y agoWhat's Rails and Postgres? Do they mean ASP.NET and MS SQL Server?
- andatki 2y agoRails and Postgres (and AWS) was the pre-acquisition stack, and development continued with that stack during this time period (2020-2021). https://en.wikipedia.org/wiki/Flip_(software) https://en.wikipedia.org/wiki/Flip_(software) Microsoft acquired companies with web and mobile platforms with varied backgrounds at a high rate. I got the sense that the tech stack—at least when it was based on open source—was evaluated for ongoing maintenance and evolution on a case by case basis. There was a cloud migration to Azure and encouragement to adopt Surface laptops and VS Code, but the leadership advocated for continuing development in the stack as feature development was ongoing, and the team was small. Besides hosted commercial versions, I was happy to see Microsoft supporting community/open source PostgreSQL so much and they continue to do so. https://en.wikipedia.org/wiki/List_of_mergers_and_acquisitions_by_Microsoft https://en.wikipedia.org/wiki/List_of_mergers_and_acquisitio... https://techcommunity.microsoft.com/t5/azure-database-for-postgresql/what-s-new-with-postgres-at-microsoft-2024-edition/ba-p/4140085 https://techcommunity.microsoft.com/t5/azure-database-for-po...
- neonsunset 2y agoPostgreSQL has been the most popular choice for greenfield .NET projects for a while too. There really isn't any vendor lock-in as most of the ecosystem is built with swappable components.
- rubyfan 2y ago15 years ago I worked on a couple of really high profile rails sites. We had millions of users with Rails and a single mysql instance (+memcached and nginx). Back then ruby was a bit slower than it is today but I’m certain some of the challenges you face at that scale are things people still do today… 1. try to make most things static-ish reads and cache generic stuff, e.g. most things became non-user specific HTML that got cached as SSI via nginx or memcached 2. move dynamic content to services to load after static-ish main content, e.g. comments, likes, etc. would be loaded via JSON after the page load 3. Move write operations to microservices, i.e. creating new content and changes to DB become mostly deferrable background operations I guess the strategy was to do as much serving of content without dipping into ruby layer except for write or infrequent reads that would update cache.
- neonsunset 2y agoThis desperately needs the Walmart treatment of JET.com’s teams past acquisition :)
- cies 2y agoMy experience scaling up Rails (mostly in size of codebase NOT in size of traffic) really made me love typesafe languages. IDE smartness (auto complete, refactoring), compile error instead of runtime, clear APIs... Kotlin is a pretty nice "Type-safe Ruby" to me nowadays.
- Alifatisk 2y agoI had a similar experience, working in a large Ruby codebase made me realise how important type-hints is, sometimes I had to investige what types where expected and required because the editor where unable to tell me. I hope RBS / Sorbet solves this.