6 ms·
Because everyone insists on insanely distributed architectures, most people will never really see the point of Redis, which is that if it is running on the same
by pjs_ 2y ago
Because everyone insists on insanely distributed architectures, most people will never really see the point of Redis, which is that if it is running on the same machine as the application, it can respond in much less than a millisecond. That lets you do stuff in the application that you just can't do with Postgres. Postgres kicks ass obviously but it is not running in memory on the same machine as the application.
If all you want to do is queues and whatnot, then sure, you don't need an in-memory KV store.
The point of an in-memory KV store is to do stuff that needs the performance characteristics of RAM. You obviously can't get the performance characteristics of RAM over a network connection. This is like, a tautology.
- Eikon 2y agoYou know what performs better than redis in this setup? An hashmap. Why Postgres wouldn’t run on the same server as the app? It’s actually pretty common.
- secondcoming 2y agoThere's more to redis than just being a K/V store.
- Eikon 2y agoThere’s more data structures in your favorite language std lib than an hashmap, at least, I hope so :)
- ummonk 2y agoLike?
- strken 2y agoPerhaps you could elaborate? It would be helpful to understand what Redis can do that cannot be done easily with local memory. Acting as shared memory for an inherently single-CPU language like JS is one I can think of. However, I don't use Redis, so you'd be better placed to drive the discussion forward with examples.
- shakow 2y ago> Perhaps you could elaborate? IPC, I suppose?
- quantadev 2y agoI have Docker running in swarm mode where it will spin up multiple load balanced instances of my web app (where requests can get routed randomly to any instance). So I use Redis to store "User Session Information" so the session can be read/written from any instance, supposedly faster than using a DB.
- mperham 2y agoRedis provides low-level persistent data structures which can be used to implement business logic distributed safely across a number of machines. That’s a LOT harder than in-memory in-process. My Sidekiq background job system runs entirely on top of Redis. Structures like Sorted Sets become the basis for indexes. Lists provide extremely fast queue behavior and Hashes map easily to persistent objects. Databases, traditionally, have not performed well when used as queues. Those are the big 3 structures necessary to implement anything: trees, lists and maps.
- chipdart 2y ago> There's more to redis than just being a K/V store. That's perfectly fine. You can compare Redis to other specialized tools just like you can compare it with Postgres and SQLite.
- secretmark 2y agoWhat if there are multiple processes that require access to a shared cache?
- mbreese 2y agoAnd I think this is the main use case they were looking for. If you have a web app where each request is a separate process/call (not uncommon), and you don’t have a good shared global state, Redis is a great tool. It is an in-memory data structure store that can respond to requests from different processes. I always considered it an evolution from memcached. If you only have one long lived process or good global variable control, then it is much less appealing in the single-server scenario. Similarly, if you require access from multiple hosts, it becomes a less obvious choice (especially if you already have a database in the loop). And redis is also overkill is you’re using it only as a cache.
- xxs 2y ago>shared cache As in performance improvement - cache should never be considered a datastore, e.g. you can pull the plug and nothing else happens (aside losing performance). It'd be a lot more beneficial all the processes to have a local cache, themselves. The latter is at least 4 orders of magnitude faster than redis. Now you may like some partitioning, too.
- nomel 2y ago> but it is not running in memory on the same machine as the application. What's the overhead of postgres vs redis, when run locally? Why do you think postgres isn't run locally? There's nothing special about postgres. It's just a program that runs in another process, just like redis. For local connections, it uses fast pipes, to reduce latency, and you get access to some faster data bulk transfer methods. I've used it in this way on many occasions.
- threeseed 2y agoPostgreSQL does not have the concept of in-memory tables. Except for temporary tables which are wiped after each session.
- nomel 2y agoThat's simplifying things a bit. Postgres has a shared memory cache, which can be set the same as redis, so your operations will all happen with in memory, with some background stuff putting it onto disk for you in the case your computer shuts off. Storage won't be involved. BUT, postgres still has ~6x the latency [1], even when run from memory. [1] https://medium.com/redis-with-raphael-de-lio/can-postgres-replace-redis-as-a-cache-f6cba13386dc#:~:text=The%20performance%20comparison%20shows%20that,observed%20for%20Postgres'%20unlogged%20table https://medium.com/redis-with-raphael-de-lio/can-postgres-re....
- threeseed 2y agoBut the whole point here is that PostgreSQL will be used for other tasks e.g. storing all of your business data. So it will be fighting for the shared cache as well as the disk. And of course storage will still be involved as again you can't have in-memory only tables. And having a locking system fluctuate in latency between milliseconds and seconds would cause all sorts of headaches.
- sgarland 2y agoIf you are both small enough that you’re considering cohosting the app and DB, the odds are good that your working set is small enough to comfortably fit into RAM on any decently-sized instance. > And having a locking system fluctuate in latency between milliseconds and seconds would cause all sorts of headaches. With the frequency that a locking system is likely to be used, it’s highly unlikely that those pages would ever get purged from the buffer pool.
- hkon 2y agoBut then you can have it in memory in your app.
- roncesvalles 2y agoIf the setup is that only one local process will use on-machine Redis as an in-memory cache, you're better off using the data structures available in your programming language.
- teaearlgraycold 2y agoThis is the vanillajs.com of data stores.
- threeseed 2y agoNot really. Because then once you restart that process you lose everything. And it's far more likely you are continuously upgrading your process than Redis.
- bastawhiz 2y agoIf you need that, you can use an embedded data store like leveldb/rocksdb or sqlite. Why bring another application running as its own process into the equation?
- nine_k 2y agoRunning in its own process, and, better yet, in its own cgroup (container) makes potential bugs in it, including RCEs, harder to exploit. It also makes it easier to limit resources it consumes, monitor its functioning, etc. Upgrading it does not require you to rebuild and redeploy your app, which may be important if a bug or performance regression occurs is triggered, and you need a quick upgrade or downgrade with (near) zero downtime. Ideally every significant part should live in its own universe, only interacting with other parts via well-defined interfaces. Sadly, it's either more expensive (even as Unix processes, to say nothing of Windows), slower (Erlang / Elixir), or both (microservices).
- LtWorf 2y agoAt the cost of requiring a lot of IPC and memory copying.
- halfcat 2y agoDjango has caching built in with support for Redis, and it also has an in-memory caching option which they label as “not for production” (because if you have multiple instances of Django serving requests, their in-memory caches will diverge which is...bad). But for lots of cases, especially internal business tools, we can scale up a single instance for a long time, and this in-memory caching makes things super fast. There’s a library, django-cachalot [1], that handles cache invalidation automatically any time a write happens to a table. That’s a rather blunt way to handle cache invalidation, but it’s wonderful because it will give you a boost for free with virtually no effort on your part, and if your internal business app has infrequent updates it basically runs entirely in RAM, and falls back to regular database queries if the data isn’t in the cache. [1] https://github.com/noripyt/django-cachalot https://github.com/noripyt/django-cachalot
- maxbond 2y agoOverengineering/premature distribution is a real problem, but Redis stands for "Remote Dictionary Server." The purpose is very much not to run it locally! (Though that's a legitimate design choice, especially if your language's native dictionary doesn't support range queries.)
- freedomben 2y agoThe purpose may not be primarily or originally to run it locally, but that has definitely become a common use case. That said, anitirez renaming it to lredis or reldis would be epic and one of my favorite moves of all time
- hinkley 2y agoOver the network became feasible when HDD got markedly slower than NICs. It’s a nearer thing with NVMe. I want a “redis” with something akin to the Consul client - which is a sidecar that participates in the Raft cluster and keeps up to date, cheaper lookups for all of the processes running on the same box. The few bit of data we needed to invalidate on infrequent writes went into consul, and the rest went into the dumbest (as in boring, not foolish) memcached cluster you can imagine. But as you say there was the network overhead, and what would be lovely is a 2 tier KV store that cached recent reads locally and distributed cache invalidation over Raft. Consistent hashing for the systems of record, broadcast data on put or delete so the local cache stays coherent.
- elcritch 2y agoI always liked the idea of a distributed sidecar like DB. I wonder if something like Cockroach DB might even work for small clusters.
- codr7 2y agoIt very well could be running on the same machine, and communicating with the app using unix sockets, which is a hell of a lot faster than TCP. But no one seems to be doing that much either. I feel that the virtualize and distribute everything to hell and back-trend might actually be about to break, there are signs, and G knows it's about time. The amounts wasted on cloud providers for apps that would run everything just fine on a single server, the effort wasted configuring their offerings, surreal.
- sgarland 2y ago> [Redis] can respond in much less than a millisecond. I have no idea how fast Redis can get, but it is entirely possible for an RDBMS to execute a query in well under a millisecond. I have instrumentation proving it. If everything is on the same machine, I would wager that IPC would ultimately be the bottleneck for both cases.
- chipdart 2y ago> Because everyone insists on insanely distributed architectures, most people will never really see the point of Redis, which is that if it is running on the same machine as the application, it can respond in much less than a millisecond. I don't think this is a realistic scenario at all. If you need a KV store, you want it to be the fastest by keeping it in-memory, you want it to run on each node, and you don't care how much it cost to scale vertically, then you do not run a separate process on the same node. You just keep a plain old dictionary data structure. You do not waste cycles deserializing and serializing queries. You only adopt something like Redis when you need far more than that, namely you need to have more than one process access shared data. That's at the core of all Redis usecases. https://redis.io/docs/latest/develop/interact/search-and-query/query-use-cases/ https://redis.io/docs/latest/develop/interact/search-and-que...
- nrdvana 2y agoI develop and maintain multiple applications that use a worker pool, and are small enough to run on a single host. We used pg for the user sessions, which get read and written on every single page request. Some of our apps are Internet-facing, and web crawlers can create sessions that get read and written (recording recent pages) as they browse the site. We switched to a redis service on the same host as the app and saw 3 main benefits: faster session loading and saving, less disk activity on the Pg server (so all other queries run faster) and less writes to the Pg WAL, so our backups require drastically less GB per day of retention. After the significant success of the first conversion, we've been working to convert all the rest of our apps. And no, host language data structures aren't useful because they aren't in shared memory between all the worker processes, and even if we found a module that implemented them in shared memory, we like to be able to preserve the sessions across a host restart, and then we'd need a process to save the data structures to disk and load them back, and by the time we did that we'd have just reinvented redis.
- machine_coffee 2y agoThis is the best response so far. Session churn creates lots of db activity but lots of it is of low business value. Better to offload to a separate process. Also session data is often Blobs which db's don't process as efficiently as columnar data.
- bufferoverflow 2y agoRedis is not just in-memory.
- xxs 2y ago> which is that if it is running on the same machine as the application in that case just use a regulator hashmap - it has nanoseconds performance compared to the sub-millis.