3 ms·
We operate with 80 Tb of data ATM. It is laying in several nodes and meta nodes (this is our own terminology). All Postgres. Recently we need to move data from
by xenator 3y ago
We operate with 80 Tb of data ATM. It is laying in several nodes and meta nodes (this is our own terminology). All Postgres.
Recently we need to move data from one DB to another, about 600M records. It is not biggest chank of the data, but we need it on different server because we use FTS a lot. And don't want to interrupt other operations on previous server. It took 3 days and costs 0.
- saisrirampur 3y agoThanks for context! Totally understand where you are coming from. Postgres can be moulded to work for many use-cases. However it could take good amount of effort to make it happen. For example in your case building and managing a sharded Postgres environment isn't straightforward. It requires quite a lot of time and expertise. Citus automated exactly this (sharding). However it wasn't a fit for every workload. https://docs.citusdata.com/en/v12.1/get_started/what_is_citus.html#when-citus-is-inappropriate https://docs.citusdata.com/en/v12.1/get_started/what_is_citu...
- xenator 3y agoTbh I never bothered by these mental restrictions. I can't say that we have effort that requires more than normal human brain can handle (and we aren't smartest people on the planet). I used to work in one project where we process big part of all shop's cash receipts in one of the biggest european country. We don't use any of these products. And it was done by one person. Only stupid idea we had was to use AWS. Learned helplessness push people to change best product on the market but without salesman who tickle your balls. Postgres is one of the best product on the market. But so much FUD makes a new space for "problem solvers" for the problem never exist. I'm not about Citus, I'm about idea that it is require much effort to build something around Postgres.
- quadrature 3y agoIs your workload an analytical or transactional workload ?.
- xenator 3y agoBoth, we process all the data in several blockchains and make public reports available for our users. Some of data can be processed once, others need to be recalculated on the schedule. So we process every block, every log entity in transaction nature. And also have a lot of data that processed several ways 2nd/3rd/etc times. We also slowly evolve our internal analytics/intelligence. It is not something that generates high load, but will be at some point. Imagine something like dune.com.