4 ms·
When a comment on HN has more merit than the article. The only problem with postgres is that inserting has some interesting scaling problems. Putting a queue b
by datadeft 3y ago
When a comment on HN has more merit than the article.
The only problem with postgres is that inserting has some interesting scaling problems. Putting a queue between the event sources and the db is usually recommended.
- withinboredom 3y ago> Putting a queue between the event sources and the db is usually recommended. Emphasis on _usually_. If your db is at 2% CPU utilization with/without queues ... you probably don't need a queue.
- asah 3y agosee my reply above - one classic case is db system maintenance.
- refset 3y ago"normalize until it hurts, denormalize until it works" is evergreen advice for scaling both reads and writes. Synchronously enforcing referential integrity and other forms of normalized constraints is what gets expensive. Pat Helland has some really good writing on this stuff, e.g. https://pathelland.substack.com/p/i-am-so-glad-im-uncoordinated https://pathelland.substack.com/p/i-am-so-glad-im-uncoordina...
- ldng 3y agoAnd in between, at the coma, revise seriously your indexing policy and don't hesitate to remove unused and underused index (even on foreign key if you don't need them that much). People too often underestimate the impact of rebuilding and index on large inserts.
- refset 3y agoAbsolutely, there's no beating the RUM Conjecture.
- asah 3y agothanks! forgot to mention queuing (e.g. SQS), which is SUPER valuable, for example when you want to do large scale maintenance on the database (major version upgrade where the on-disk format can change)
- varelaz 3y agoSQS has eventual consistency, you can get the same message twice on 2 different intances for example (which was often the case for my projects). I would rather suggest Amazon MQ.
- icedchai 3y agoThis is often possible to work around by adding a UUID to the message, on the sender side. Then either handle dupes at the DB level by ignoring duplicate inserts, or using something like redis. In practice, even with other queue systems, you can wind up with dupes due to bugs, timeouts, retries, like duplicate sends when the first one actually went through but wasn't acknowledged. I worked on systems sending thousands of messages/second through rabbitmq, originating from various third parties, and dupes would happen often enough that we needed to work around it on the receiving "worker" side.
- withinboredom 3y agoSQS is weird in that you can get the same message in Japan and Germany at the exact same time and "race" to process the message. It's annoying af. I do not recommend SQS unless you are at some kind of scale (in company/team size, not revenue) where you can deal with it properly.
- icedchai 3y agoI used extensively at a previous company for longer running background tasks. It was simpler to use SQS than dealing with standing up our own RabbitMQ cluster. Their Amazon MQ service did not exist at the time. Our system was built to tolerate duplicates and it worked well enough. For something higher volume I'd definitely use RabbitMQ though.
- layer8 3y ago> Putting a queue between the event sources and the db is usually recommended. That depends on the nature of the events and whether you can live with the database being out-of-date while the events are still in the queue.