10 ms·
How We Built Fly Postgres
- asguy 4y ago> we’re good at consul Thank god someone is. I’ve lost more of my life to consul partition failures that’s any other part of the nutech stack.
- madeofpalk 4y agoThey've had their fair share of pains with it, but they still seem pretty happy with it https://fly.io/blog/a-foolish-consistency/ https://fly.io/blog/a-foolish-consistency/
- tptacek 4y agoHappy is a word. There are lots of words. I like words! You could be creative about what word could take the place of "happy" in that sentence, and probably still be correct. Get weird with it! Maybe "sanguine" would work. "Engaged". The reality is: we've got a fair bit of experience with Consul at this point, we respect the hell out of it for the problems it was designed to solve, and we're unlikely to stretch it any further than we've already stretched it. Distributed lock service for Postgres clusters? Sure. Source of truth for all our app state? We've built our own thingy ("corrosion", a Rust distribute state system) to phase Consul out with. I'll get Jerome to say things about it.
- subarctic 4y ago>Sure. Source of truth for all our app state? We've built our own thingy ("corrosion", a Rust distribute state system) to phase Consul out with. I'll get Jerome to say things about it. Looking forward to the blog post!
- madeofpalk 4y agoSorry yes. 'Invested' was more along what I was thinking :)
- WFHRenaissance 4y agoWhat else would you include in the nutech stack? HashiCorp products in general?
- atonse 4y agoYup our experience with brittleness of consul and nomad soured me on the HashiCorp stack even though I was very excited about it. We went from buying enterprise licenses to throwing it all out without ever pushing to production in the span of a year. This stuff, for all the hype about raft and things, is brittle enough and requires enough special attention. Which is why I’d gladly rather have it be fly.io’s problem.
- tptacek 4y agoFor what it's worth, if I was deploying a bounded set of applications or services in a small number of geographies and couldn't use Fly.io for some reason, I wouldn't hesitate to use Nomad. Nomad is pretty great; it's like Flask to K8s's Django.
- atonse 4y agoI did love the simplicity of nomad. And in general, nomad worked pretty well for us but our consul cluster kept mysteriously failing. I think that caused our nomad cluster to fail because it was backed by Consul. The one complaint I did have about nomad (same as consul) was that the recovery process was manual where you had to manually generate a peers.json. I was shocked when I saw that. Truly one of the "finding how the sausage is made" moments even though I've managed linux servers for two decades – I always assumed it would use zeroconf/bonjour/multicast DNS (remember cloud auto-join?) or something similarly elegant to auto discover other nodes in the network and just reconnect and rebuild a cluster. I mean what's the point of all this stuff if it can't be used to recover a cluster and just Do The Right Thing™? The shiny new experience is stellar (like sales, or setting up a new cluster), but the flip side (when things go wrong) is a mess. That's why we eventually said "nope!" to all the custom stuff and went with boring, plain vanilla ECS, which is itself too much now that we've started using fly. Don't ever want to even think about having to hand-write a peers.json file to recover a cluster, boot things up, and pray to the ancient gods that it works. We don't have time for that nonsense. Please, take my money, Fly/Render/everyone else. Your costs are a margin of error compared to what I had to pay a devops person to build our own stack. (I'm not even exaggerating. It was six figures. DevOps people are worth every penny but they cost many, many pennies.) Ultimately, we never used the infra. I want to focus on building solutions for my customers and not fiddling with weird server stuff.
- MuffinFlavored 4y ago> You can spin up a Postgres database, or a whole cluster, with just a couple of commands. Sign up for Fly.io and launch a full-stack app in minutes! What is the HackerNews opinion on: who is their actual customer? Obviously lots of companies pay for managed databases. It's not an uncrowded market for a reason. But like... pricing wise... it seems so expensive? What is the HackerNews take on the valuable proposition specifically for hosted databases? Is the answer basically 1:1 with anything cloud hosting related?
- rozenmd 4y agoLast I checked, a 4 vCPU/16GB RAM/1TB storage configuration costs around $80 USD per month at VPS hosts like Hetzner. It's $762 USD per month on RDS. There are trade-offs of course (https://onlineornot.com/self-hosting-vs-managed-services-deciding-how-host-your-database https://onlineornot.com/self-hosting-vs-managed-services-dec...) between the two options. I've been hoping fly.io builds a strong automated middle ground for a while now!
- w-ll 4y agoAm I reading https://instances.vantage.sh/rds/?min_vcpus=4&cost_duration=monthly https://instances.vantage.sh/rds/?min_vcpus=4&cost_duration=... wrong. It seams you can get that as a db.t4g.xlarge for ~$188/month. Still more than 2x the cost, but not nearly $760
- ithrow 4y ago$188 without storage and bandwidth.
- some_developer 4y ago> for ~$188/month. Note that this is still the on-demand price. Although it's still not Apple to Apple, no one (sane) would useRDS 24x7 without RI. That probably brings you down another 50 bucks or so.
- mnutt 4y ago
- stuff4ben 4y agoUsing Stolon for PG is a poor choice. Up until just very recently, they haven't had any significant updates in a year. We've abandoned our use of it in favor of EnterpriseDB.
- tptacek 4y agoThere's a bunch of our own engineering going into this. But: the good news about this whole situation is, if you have a clustering solution you like better: you can just use it. We "automate" Fly Postgres, but we don't "manage" it. Fly Postgres is using features of Fly.io that are available to anybody's application, not just ours.
- garganzol 4y agoDatabase is a thing I never want to deal with at such low level. Such approach will be brittle, very brittle. Managed data storage with SLA guarantees is not easy and there are quite a few companies specializing on that for a reason. As a proof, check out fly.io forums. They are full with posts about suddenly broken Postgres instances.
- craigkerstiens 4y agoFWIW, managed providers can fully plug into fly just fine. Here's actually a profile of performance times [1] of Fly with various providers and configurations, along with a repo [2] to reproduce/create the same setup yourself. 1. https://webstack.dancroak.com/ https://webstack.dancroak.com/ 2. https://github.com/croaky/webstack https://github.com/croaky/webstack *Disclaimer I work at one of those fully managed database providers.
- Scarbutt 4y agoWhy are the response times (api checks) so high?
- lukeasrodgers 4y agoYeah those response times are in the realm of “why would I even consider that”. If I am trying to tune my db queries to have p95 latency of 1ms (for example) there’s no way I would choose an architecture that then threw that all out the window with ~100ms network latency. Hopefully I am misunderstanding those numbers somehow.
- sb8244 4y agoI saw this earlier and thoroughly didn't understand what is going on with that. I can't make any sense of why that would be so high. Feels like a better test would be for the healthcheck to time the query RTT and report that back. Completely remove the web request from the equation.
- samwillis 4y agoI would love it if Crunchy Bridge was available on Fly. Have you considered offering it alongside the other clouds you run on? I could see a Crunchy Data managed version of Fly Postgres's doing super well. Combining the best aspects of both your experience managing HA Postgres and backup, with the distributed read replicas from Fly.
- smallerfish 4y agoThis is tangential - anybody have a reasonably clean way of running JVM apps on fly? Given that the deploy model of a jettified jar is so clean in comparison to the various hipster stacks ;) that they do have primary support for, I'm not sure why they don't have much in the way of documentation for it. I did find a support thread that refers to https://archive.is/o9YE1 https://archive.is/o9YE1, which seems like a fairly gross and opaque sequence of steps. @tptacek, what are the chances that somebody at Fly could produce a more streamlined recipe and/or documentation for JVM apps?
- garganzol 4y agoIf you avoid buildpacks altogether and just go straight to Dockerfiles then life suddenly becomes good. Speaking from my own experience.
- Scarbutt 4y agothey support docker as a second class citizen.
- michaeldwan 4y agoDocker (or rather OCI) is our first class citizen. The launchers for phoenix/rails/etc just generate dockerfiles and config by inspecting source code.
- jacktheturtle 4y agothe site is down now
- chrisweekly 4y agoI always enjoy Fly.io's blog posts; the friendly, casual tone coupled with real-world experience and strong technical chops of the staff add up to HN gold in my book. I'll def revisit their "treat pg as an app" approach next time I engage w stuff in that realm.
- fifanut 4y agoImproving Postgres is solid and boring. Solid and boring is often a good choice. I'm glad to see startups in this space. What's the latest on adoption of Spanner-like databases?
- phamilton 4y agoThis was on HN a few months back: https://github.com/losfair/mvsqlite https://github.com/losfair/mvsqlite While not Spanner, it is essentially an open source db like AlloyDB or Aurora, pushing replication and scale out to the storage layer (in this case via FoundationDB). The most interesting bit of mvsqlite is it's multi-writer capabilities, using FoundationDB to perform page-level locks. I'm neither the creator nor using it in production, but I'd love to see more DBs using FoundationDB as storage. It's a pretty cool solution.
- ignoramous 4y ago> This was on HN a few months back: https://github.com/losfair/mvsqlite https://github.com/losfair/mvsqlite The lead developer on mvsqlite has since joined Deno...
- satvikpendem 4y agoAny way to run Fly with multiple locations at the same time? For context, I want to build a simple uptime service which tells me when my website is down, but for that I don't want to use a single VPS, I want to load balance between several locations and servers in case any one of them goes down. Am I supposed to be looking for serverless deployment? Deploy a master/parent version of the app on one main VPS then several sub versions on other VPSes distributed around?
- tptacek 4y agoYou can scale a Fly app to an essentially arbitrary number of instances (`scale count x`), and deploy in any number of regions (`regions set nrt syd sin fra`). We'll load balance between instances in nearby regions.
- satvikpendem 4y agoThanks Thomas, glad to see you here. For the regions, will I be able to know programmatically in which region a particular instance of the app is running? For reference, I want the users of my service to know that we tested their site in regions A, B and C, and that it's currently up in A and B but down in C and show them that in the web dashboard.
- tptacek 4y agoYep, it's in the environment: `FLY_REGION`.
- satvikpendem 4y agoThanks, appreciate it! By the way, did you start at Fly recently? I seem to recall you had your own company or something like that before.
- tptacek 4y agoI've been at Fly.io since early summer 2020. I've had a couple of companies before that. None of them as fun as this one! :)
- atonse 4y agoI do have a question about this – (as I've said in another comment here) after investing a lot in building our own terraform/aws stack, we've started moving all our smaller apps to fly and it's downright delightful to the point where we are considering moving even our HIPAA-compliant software to Fly as soon as I saw the BAA option in your pricing, (and the only hesitation is that there is no Vanta integration and we'd probably have to fill out a bunch of stuff manually to make auditors happy). But to my actual question: I remember seeing somewhere that Fly Postgres is not a managed database and isn't as fire-and-forget (somewhere in the docs), and that honestly scares me a bit. It shouldn't, because at the end of the day, RDS is probably downright hairy and ugly under the hood, it's just hidden from us. I've also seen a handful of reliability issues on the forums around postgres. So what is Fly's position on actually running these clusters? Are you guys feeling pretty good about its stability for critical workloads? On a related note, these folks that are continuing to use RDS via fly... what are the latencies like with a us-east-1 RDS? I know you're in Ashburn so it must be sub 5ms? That itself would be worth keeping our RDS and moving everything else to Fly to be honest.
- sb8244 4y agoI recently setup fly -> ec2 -> rds using their wireguard + pgbouncer template. I'm getting 7ms RTT on us-east-1. Of course I messed up my setup the first time and had RDS in Ohio... That was 30ms. You would probably get sub 5ms if you did public RDS connection. Don't know if that would ever pass an audit though.
- atonse 4y agoYeah we’d never even consider adding a public ip to our databases (or web servers)
- langsoul-com 4y agoFly is pretty cool whilst free. Very easy to setup stuff. Would recommend using prepaid credits because there's no capped billing. Getting closer and closer to heroku deploy and auto config handles everything.
- dikei 4y agoNever used Stolon, but I've hand good experience using Patroni to auto-manage the fail-over of multiple PostgreSQL clusters.
- solarkraft 4y agoI have a somewhat unusual use case: A publicly available Postgres server with PostGIS for (infrequent) use with QGis. I understand why Fly doesn't want customers to expose the database to the rest of the internet (ecosystem and such), but am very happy that Railway allows it.
- DAlperin 4y agoWe allow this now: https://community.fly.io/t/new-proxy-handler-pg-tls-postgresql-sslmode/8788 https://community.fly.io/t/new-proxy-handler-pg-tls-postgres... :)
- tobase 4y agoWe are actually migrating from Fly because of bad performance and unreliability on the pg services. Still love the company tho but can’t support it more :/