4 ms·
People are willing to put in way too much work just to avoid using Prometheus or InfluxDB, aren't they?
by starttoaster 3y ago
People are willing to put in way too much work just to avoid using Prometheus or InfluxDB, aren't they?
- avereveard 3y agoTheir solution is zfs on the write master can't wait for the next blog post on how they found their data corrupted
- philkrylov 3y agoPostgreSQL does not use SEEK_DATA/SEEK_HOLE so they're ok
- bfung 3y agoOr how after they do a writer failover, they start seeing duplicate data.
- ikiris 3y agoOk, I'm out of the loop, whats the problem with zfs here?
- avereveard 3y agohttps://hn.algolia.com/?dateRange=all&page=0&prefix=false&query=Zfs%20corruption&sort=byDate&type=story https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu... the mean time between corruption article on zfs is two years
- philkrylov 3y agoLooking at your search results, there's just one recent ZFS corruption case with SEEK_DATA/SEEK_HOLE (in several HN reflections), a 2-year old Ubuntu-only buggy patch story, and some 2008 [Open]Solaris corruption.
- ikiris 3y agoMost of those links are about bad memory. If you blame bad memory for filesystem issues I don't really know what to tell you. Ignoring the poor state of the native encryption code, ZFS has had 1 corruption bug in like 10 years. Thats one of the best records for modern filesystems. I still wouldn't trust my data to btrfs by comparison.
- WJW 3y agoPosts with the basic messgage "Use nothing but postgres for everything from pubsub to background job queues, it's the best thing since sliced bread and will solve all your problems" have been a HN staple for at least a decade. It's no surprise that sooner or later people would start believing it.
- Ozzie_osman 3y agoIn all fairness, postgres still works for their workload even if RDS didn't. You can get a lot of workloads out of postgres with the right hardware, the right replication, and the right extensions (eg Citus has a columnar extension that probably would have been a pretty good fit for that).
- eximius 3y agoEh, it's a reaction against people making or reaching for the wrong tools or the right tools but at the wrong scale. Postgres is very very good. The vast majority of use cases work with it with very minor effort. People would, in general, be better off investing in thoroughly understanding a general tool like postgres (or similar dbs, just pick one to learn, but there are reasons why you would pick postgres over, say, oracle). There are still reasons to use more specialized DBs. But the push for postgres is because very often the people reaching for those specialized DBs do so in error. 20M rows is practically an in-memory dataset, for example.
- tracker1 3y agoTo be fair, we're in an age with servers that can handle hundreds of simultaneous threads on a single system with terabytes of RAM and storage faster than RAM a few generations back. You can scale up a lot with a general purpose RDBMS like postgres on a single server, and a read replica today. It's not perfect, or even ideal for many workloads or even all environments... But it probably can be good enough for most application needs. It's knowing when it isn't, why it isn't, and what to use instead that counts in those instances. But I hold no blame for starting with what is probably one of the better known and understood solutions to start with.
- zilti 3y agoOr just TimescaleDB
- wenc 3y agoUnless the time-series was only for simple monitoring or querying, I would stay away from key-value databases like Prometheus or InfluxDB which have limited joins and limited analytics capabilities. A fully relational time-series database like Timescale (a plugin db built on Postgres) gives you full SQL analytics, including aggregations and full relational joins with other data, which is where a lot of the value-add usually is. This also opens up the field to building multivariate machine learning models.
- Spivak 3y agoWell yeah, of course. I can't understand why people reach for a bunch of bespoke databases where you need a whole other ecosystem of tooling and libraries to use and monitor, can't have transactions across them, another single point of failure, your ORMs don't mix if you use one of those. The amount of work you need to do to make it not worth it is quite high assuming it can be done which they seem to have accomplished pretty easily by DIYing their own provisioned iops rds (not sure why they didn't try that).
- golergka 3y agoIt's always a good idea to use a jack of all trades like PostgreSQL that you know well for a first version and migrate parts of your service to a specialised tool that you have to research later, after you're sure that you have a good product.
- hobobaggins 3y agoIt's not that pgsql wasn't appropriate, it's that the neutered AWS RDS managed instance was inappropriate. Whether more appropriate non-pgsql solutions existed seems to have been outside the scope of the article.