5 ms·
That's just a really really bad write-up on the real problem on running a database on k8s. You need ha because k8s should run already with automatic node upgra
by Felminor 4y ago
That's just a really really bad write-up on the real problem on running a database on k8s.
You need ha because k8s should run already with automatic node upgrades.
You need a pod disruption budget to make sure it is running and switching over when a node fails or gets upgraded.
You want to either totally Oberprovision on memory or look into keep 2400 to make sure to fine-tune memory before k8s starts to throw your database out constantly.
K8s is not a VM.
If you use k8s and still don't take care of application migration strategies you still don't understand what cloud native means.
There are still other things missing here but still...
Of course excluding hobby people playing with k8s.
Memory and upgrading nodes are the two single most issues will see which disrupts service.
Otherwise k8s is a dream come true.
I still would try to use a db managed if it's critical.
Additional points:
Zalando postgres operator is great and shows the real magic of k8s and operator.
Use a helm chart and just bring your own little database for dev test and e2e tests.
You can easily use Auto scaling for node profiles. No noisy neighbors. If your db is too small for normal nodes you don't have a problem anyway.
- samokhvalov 4y ago> Use a helm chart and just bring your own little database for dev test and e2e tests. dev, test, and e2e tests should be done against full-size db clones
- e12e 4y ago> dev, test, and e2e tests should be done against full-size db clones Real customer/sensitive data should not exist outside prod (and backup). So generally no, not full-size clones. I'd argue instrumentation in prod should give information on performance - for some tests/development you might need prod-size fake data.
- holografix 4y agoCouldn’t agree more. Having full sized and fully speced dev and test DBs is wasteful and not realistic across several independent teams. Monitoring prod closely and understanding what could constitute a costly workload/query which would in turn require a temporary test env with similar sized dataset is the correct approach.
- samokhvalov 4y agoThis is too reactive approach, isn't it? Imagine a system that could do this in CI/CD pipelines: 1) first, run all the tests on a tiny DB as usual 2) extract queries 3) run them against full-size DB branch/thin clone (thin provisioning, CoW; PII is not there of course, wiped out for security/compliance) -- auto-guessing parms, that's the trickiest part, but assume it's solved 4) collect all the details about performance, focusing on IO numbers (rows, block read/writes) 5) if some queries are off – say, you forgot LIMIT – post a warning, block the change, do not allow deploying it, letting backend dev fix it. This would be proactive. And it's becoming possible with modern tools.
- Proven 4y ago[dead]
- Blackthorn 4y agoYou think I'm going to clone a multiple petabyte database just to run some tests?
- samokhvalov 4y agonot sure about petabytes (yet), but for dozens of TiB, we are fine – DB branching, thin clones, CoW are to the rescue.
- Proven 4y ago[dead]
- axlee 4y ago> dev, test, and e2e tests should be done against full-size db clones that's cute, what is your "full-size"? I don't have 2 days to run a test, and I'm pretty sure every single compliance requirements we are following would get obliterated the second someone hears about us doing that
- samokhvalov 4y agoAgreed on compliance part – of course, in many cases (not in all though) PII must not be in dev/test envs. Although, cannot agree with the former part. If your tests are running 2 days on a full size clone, and it's an OLTP case, what about users, do they suffer from long query duration too? It sounds like it's time to optimize queries and/or redesign test sets (or both). If it's bad in testing, it will be bad in prod. That's the idea of testing. Example: how do you check schema changes?
- eddsh1994 4y agoYou have a tool to keep the structure of data but anonymize it, I’ve seen this a few times in healthcare regulated systems
- Foobar8568 4y agoI am laughing each time I hear this so call anonymous process /data. Each time, there were different ways to link back data or things were badly scraped.
- solatic 4y ago> k8s should run already with automatic node upgrades This is difficult to impossible to do with databases; even if your database has a built-in recovery method for when a primary is taken offline, in such a way that allows for zero-downtime in theory, the reality is that such mechanisms depend on the secondary staying online until the failover mechanism is complete. If you turn over control of node upgrades to the cluster provider, the node under the secondary will get rebooted in the middle of the failover process, and you will get downtime at best, data loss at worst. What kubernetes teaches us is that databases aren't tied to the literal VM they're running on (which is now cattle), but rather on the availability of that node. If you run databases on kubernetes, you need to have a mechanism to slow down node upgrades. Source: helped run hundreds of Elasticsearch and Kafka nodes on kubernetes in production at one point in my career
- AtlasBarfed 4y agoOnline lossless zero downtime upgrades? I've done it with Cassandra...and yeah Kafka can do it I've heard. But those can be 30 hour operations even with you ducks in a row, and you better have backup strategies ready. Fun story, Amazon said rds would be always be zero downtime upgrades. But then came a major version upgrade and .... Surprise it wasn't.
- anecdotal1 4y agoAdd a new RDS replica, wait for it to sync, promote it to master? Zero downtime achieved
- AtlasBarfed 4y agoI didn't do the upgrade, I'm not a postgres MySQL person, but the best they could do was a third party tool that dropped it to a couple minutes.
- redrove 4y agoYour claim about needing the primary online to failiover to the secondary is untrue, at least not for all Postgres operators. Cloudnative PG rebuilds the secondary during failiover from the streamed WAL to an S3 endpoint. No primary needed.