Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
mslot
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
61.
▲
by
mslot
6y ago
It is a custom format that was originally derived from ORC, but is very different at this point. For instance, all the metadata is kept in PostgreSQL catalog tables to make changes transactional.
62.
▲
by
mslot
6y ago
Correct, though it depends whether you are CPU-bound or I/O-bound. We see the latter a lot more often for large data sets. Columnar storage for PostgreSQL is especially relevant in cloud environments. Most database servers in the cloud
63.
▲
by
mslot
6y ago
It always depends on the data, but we've seen 92.5% and more: https://twitter.com/JeffMealo/status/1368030569557286915
64.
▲
by
mslot
6y ago
The access method approach followed in Citus is indeed lower level and more generic, which means it can be used on both time series data and other types of data. For time series data, you can use built-in partitioning in PostgreSQL. It'
65.
▲
Citus 10: Columnar for Postgres, rebalancer in open source, single-node, & more
(citusdata.com)
7 points
by
mslot
6y ago
|
0 comments
66.
▲
by
mslot
6y ago
Sharding a matching engine is indeed pretty hard, and requires redundancy and very deliberate data modelling choices. That does seem like a fun exercise :). The general rules of the game are: You can only scale up throughput of queries/
67.
▲
by
mslot
6y ago
One caveat to keep in mind when using the binary format is that arrays of custom types are not portable across databases because the serialized array contains the OID of the custom type, which may be different on the other end. The other th
68.
▲
by
mslot
6y ago
The important thing to remember is to use COPY .. FROM STDIN, not insert, for bulk loading into PostgreSQL. Most PostgreSQL drivers support COPY from the client, though it's not always available through generic database APIs. COPY comm
69.
▲
by
mslot
6y ago
Author here. What's kind of unique about the distributed functions feature in Citus that's described in the blog post is that you can scale stored procedures horizontally if procedure calls query co-located shards, but it doesn&#x
70.
▲
by
mslot
6y ago
They are currently not compatible. You can use Citus with pg_partman for time series data: http://docs.citusdata.com/en/latest/use_cases/timeseries.htm... pg_partman does not have all the features of Timescal
71.
▲
Evolving pg_cron together: Postgres 13, audit log, background workers, job names
(citusdata.com)
2 points
by
mslot
6y ago
|
2 comments
72.
▲
Evolving pg_cron: Postgres 13, audit log, background workers, & job names
(techcommunity.microsoft.com)
2 points
by
mslot
6y ago
|
0 comments
73.
▲
by
mslot
6y ago
A reason could be that the yields on other forms of investment are going down. If buyers are willing to accept a 1% yield because they cannot easily get it anywhere else then the share price can go to $1 trillion, or even further if yield i
74.
▲
by
mslot
7y ago
My feet are currently around 4 meters below sea level in the Netherlands. They are still dry. On its own, even worst case sea level rise is unlikely to have meaningful impact on the Netherlands over the next century, unless we stop maintain
75.
▲
by
mslot
7y ago
PostgreSQL 12 has REINDEX CONCURRENTLY.
76.
▲
by
mslot
7y ago
Stored procedures can also be a great way to scale out transactional workloads. In Citus (a sharding extension for Postgres by Microsoft) we recently introduced a stored procedure call delegation feature. If your tables are distributed by p
77.
▲
by
mslot
7y ago
Why even memcached? It's not going to be 10x faster than your typical LRU cache (e.g. Postgres' shared buffers). It might scale a little better, but you should only consider that when it's worth it.
78.
▲
by
mslot
7y ago
Looking up a row of JSONB data by a primary key in a 3GB PostgreSQL table on my laptop takes 0.4-0.5ms, since it easily fits in memory and Postgres caches the data.
79.
▲
by
mslot
7y ago
Yes. A generated column can be included in a foreign key and be referenced by a foreign key. Nothing too weird happens. It's basically like having every update and insert specify the column value.
80.
▲
by
mslot
7y ago
I doubt that's a significant factor in global green house gas production while plastic bags are a significant factor in plastic pollution. This seems like a basic optimization problem. Yes, this approach uses a tiny bit more memory, bu
81.
▲
by
mslot
7y ago
At some point people realized servers are prone to failure. They then started deploying their system redundantly to multiple servers in the data center (AZ) to increase availability. This helped, but created consistency issues. To fix this
82.
▲
by
mslot
8y ago
If you want to build something like this yourself, Postgres has a Trigram index: https://www.postgresql.org/docs/current/pgtrgm.html
83.
▲
by
mslot
8y ago
Citus is also used for large time-series / analytics use cases e.g. https://www.citusdata.com/customers/heap There's a question of what you actually want to do with the time-series data. If you don't exp
84.
▲
by
mslot
8y ago
PostgreSQL has excellent time series capabilities. It can load millions of rows per second, efficiently scan by time range, build rollup tables, has expressive SQL with excellent support for time (timezones, ranges, timestamps, intervals, c
85.
▲
by
mslot
8y ago
Citus is an open source plug-in that you load into vanilla PostgreSQL, similar to PostGIS and other extensions.
86.
▲
by
mslot
8y ago
> When we asked the core PostgreSQL devs about this, they explained that they did this because sorting out the appropriate locks was a hard problem, and that they saw this scenario as so unlikely for OLTP that they instead directed their
87.
▲
by
mslot
8y ago
> Spanner runs consensus for each key range and does not need all nodes to be available to make progress also my understanding since again it has a leader for each key range writes scale better. Spanner uses Paxos (consensus) for replica
88.
▲
by
mslot
8y ago
I don't think you can get around using 2PC in a distributed OLTP database, e.g. Spanner also uses 2PC for distributed transactions across shards. Fortunately, the overhead is not really that high because the prepare and commit messages
89.
▲
by
mslot
8y ago
Replicating data in consensus groups gives some nice guarantees, in particular it can guarantee monotonic read consistency even if you're not reading from the leader, at the cost of a network round-trip. Switching between master-slave
90.
▲
by
mslot
8y ago
Google translate.
More ›