4 ms·
For larger volumes of events, we wouldn't recommend using Postgres. The nice thing about single-tenancy is that in reality lots of users have small enough data
by timgl 7y ago
For larger volumes of events, we wouldn't recommend using Postgres.
The nice thing about single-tenancy is that in reality lots of users have small enough datasets that scaling isn't a problem. Heap et al have to scale to all of their users combined (as you said, terabytes), we just have to scale to the biggest user. Postgres also allows you to get started very quickly and do lots of queries yourself.
In our docs we explain our thinking more. Postgres is great for the vast majority of use-cases, and we're working hard to optimise those queries.
Once users get beyond Postgres, we have integrations with databases that can scale well across many hosts, and we provide services around this to help people size their servers correctly.
- malisper 7y ago> Heap et al have to scale to all of their users combined. The hard part wasn't scaling the system to handle all users combined. The hard part was designing the system such that when an individual user runs a query, they would get their results back in a reasonable amount of time. Having every user in a single cluster made this easier because an individual customer could make use of the computer power of a cluster that was sized to fit the data for everyone in it. In other words, if Heap doubled the number of customers, Heap would get twice as fast for everyone. That's not true for PostHog. > Heap et al have to scale to all of their users combined (as you said, terabytes), we just have to scale to the biggest user. A decent sized Heap customer had multiple terabytes of data with the largest being well beyond that. You're going to have to figure out how to scale PostHog to that point without the benefits of multi-tenancy. > Once users get beyond Postgres, we have integrations with databases that can scale well across many hosts, and we provide services around this to help people size their servers correctly. I think a cluster of servers that could churn through terabytes of data in seconds would be prohibitively expensive for any individual customer to purchase.