6 ms·
You could run Postgres over UNIX sockets although you will still get higher latency than SQLite's in-process model. Also, running a Postgres on every app instan
by benbjohnson 4y ago
You could run Postgres over UNIX sockets although you will still get higher latency than SQLite's in-process model. Also, running a Postgres on every app instance on the edge probably isn't practical. Postgres has some great advanced features if you need them but it's also much more heavy weight.
With LiteFS, we're aiming to easily run on low resource cloud hardware such as nodes with 256MB or less of RAM. I haven't tried it on a Raspberry Pi yet but I suspect it would run fine there as well.
- pphysch 4y agoSQLite almost certainly is the better edge RDBMS than Postgres, if only because it has less features taking up space. However, "local SQLite vs. remote Postgres/MySQL" remains a false dichotomy when talking about network latency.
- yencabulator 4y agoThere's still plenty of overhead in serializing data over a unix domain socket to a different process, waiting for that process to be scheduled, waiting for it to serialize a response, waiting for the client process to be scheduled. SQLite avoids all of that.
- clord 4y agoIt's an application design choice. It's perfectly reasonable to consider those two options when designing a system. The constraints of the system determine which is better for the application. I'd pick a centralized network-reachable database with a strong relational schema for write-heavy applications, and a lighter in-process system for something that is mostly reads and where latency matters. It's not a false dichotomy — but more like a continuum that certainly includes both extremes on it.
- cbsmith 4y agoIsn't this all talking to S3 anyway, not to mention the network trips intrinsic to the system before it gets to S3? I mean, I'm sure there's some performance win here, but I'm surprised it's so significant. It's not like the address-space separation is without benefits... heck, if it weren't, you could simply have embedded the whole application inside Postgres and achieved the same effect.
- simonw 4y agoYou may be confusing Litestream and LiteFS. Litestream writes everything to S3 (or similar storage). LiteFS lets different nodes copy replicated data directly to each other over a network, without involving S3. In either case, the actual SQLite writes and reads all happen directly against local disk, without any network traffic. Replication happens after that.
- cbsmith 4y agoI think I was definitely confusing it with Litestream as the blog post made reference to it (and I did find that confusing). That said, unless I've misunderstood the LifeFS use case, you're still going over the network to reach a node, and that node is still going through a FUSE filesystem. That would seem to create overhead comparable (potentially more significant) to talking to a Postgres database hosted on a remote node. It just doesn't seem that obvious that there's a big performance win here. I'd be curious to see the profiling data behind this.
- simonw 4y agoThe performance boost that matters most here is when your application reads from the database. Your application code is reading directly from disk there, through a very thin FUSE layer that does nothing at all with reads (it only monitors writes). So your read queries should mostly be measured in microseconds.
- cbsmith 4y ago> So your read queries should mostly be measured in microseconds. You should check out the read latency for read-only requests over unix domain sockets with PostgreSQL. You tend to measure it in microseconds, and depending on circumstances it can be single-digit microseconds. Regardless of whether your FUSE logic does nothing at all, It sure seems like there's intrinsic overhead to the FUSE model that is very similar to the intrinsic overhead of talking to another userspace database process... because you're talking to another userspace process (through the VFS layer). When the application reads, those requests go from userspace to the FUSE driver & /dev/fuse, with the thread being put into a wait state; then the FUSE daemon needs to pull the request from /dev/fuse to service it; then your FUSE code does whatever minimal work it needs to do to process the read and passes it back through /dev/fuse and the FUSE driver, and from there back to your application. That gets you pretty much the same "block and context switch" overhead of an IPC call to Postgres (arguably more). FUSE uses splicing to minimize data copying (of course, unix domain sockets also minimize data copying), though looking at the LiteFS Go daemon, I'm not entirely sure there isn't a copy going on anyway. Memory copying issues aside, from a latency perspective, you're jumping through very similar hoops to talking to another user-space process... because that's how FUSE works. There's a potential performance win if the data you're reading is already in the VFS cache, since that would bypass having to through the FUSE filesystem (and the /dev/fuse-to-userspace jump) entirely. The catch is, at that point you're bypassing SQLite transaction engine semantics entirely, giving you a dirty read that's really just a cached result from a previous read. That's not really a new trick, and you can get even better performance with client-side caching that can avoid a trip to kernel space. I'm sure there's a win here somewhere, but I'm struggling to understand where.
- Scarbutt 4y agoTL;DR "Postgres doesn't fit well our edge services business model" which sure fine but the article is indeed biased/misleading by completely ignoring and not mentioning the option of running postgres and app in a single server. The gains in latency with sqlite won't matter as soon as throughput starts to dominate.