18 ms·
Streams: a new general purpose data structure in Redis
- erulabs 9y agoFantastic news, congrats Salvatore! Cannot _wait_ to replace some hacky Kafka uses with tried-and-true Redis4! :)
- ykler 9y agoIn what sense is Kafka (or your use of it) hacky? I have never used Kafka, but I have always thought of it as being more solidly engineered than Redis but also more complicated and perhaps tricky to deploy (based on blog posts I read).
- sidlls 9y agoIn any context its used where the demand (by whatever measure you care to use: bandwidth, throughput, message durability, etc.) doesn't justify it or isn't a good use case of Kafka, for starters. That happens all the time, because every data and infrastructure engineer in the Bay Area wants to put Kafka on his resume.
- deleted 9y ago[deleted]
- hueving 9y agoBut that doesn't explain why Kafka has any minimum the output required. Does it have usability issues? A good tool should be able to be used at any scale.
- sidlls 9y ago> A good tool should be able to be used at any scale. I don't necessarily agree that a tool should be used at any scale even if it's technically possible to. Cost (multiple dimensions, including money and engineering effort) factors in.
- erulabs 9y agoKafka has very poor tooling in my experience (a folder full of fairly buggy bash scripts...), and due to ZooKeeper requires a lot of operational care. For example, it's extremely easy to destroy a Kafka cluster by bringing a new, empty ZK server online with newer but incorrect data in its volume. ZK will happily trash the entire cluster thinking it has new instructions. So network isolation is key, which, while obvious, is another source of potential failure. Kafka also has the JVM, which requires a lot of love to scale in my experience. I do not want my programmers messing around with GC options when writing to what should (to them) be exposed just like a regular file handle (except distributed across many systems). I strongly prefer to avoid Java applications at all costs - in my experience it takes years and years and years for Java based infrastructure to become relatively stable & reliable (see ElasticSearch 5.0, or ask anyone who has been oncall for a Tomcat based application). This is almost certainly personal bias, but it's my bias regardless. Redis also has a _massive_ number of tooling / monitoring / ecosystem advantages, including hosted options, and can run on a single instance without configuration changes from the developers perspective. I also have personal reasons to prefer Salvatore's work over the work of Confluent.
- mioelnir 9y agoAs someone that has both Kafka and Redis in use without issue, for years, (and is about to replace a lot of misused Redis instances with Kafka) I really fail to follow your points. So, a Zookeeper cluster can't survive accidentally injecting just the right malicious data that will make it keel over. I'm sorry, how do you accidentally achieve that? Do you also accidentally configure your Redis Sentinel to replicate from /dev/null? As a matter of fact, this announcement comes at a very inopportune time for me. antirez had the epiphany of reading on IRC about replicated logs instead of looking at the opening paragraphs of the Kafka documentation, and all the Redis evangelists at my job will now try to shoe-horn the wrong usecase back into Redis because Redis!1cos(0)!. Sigh.
- takeda 9y ago> For example, it's extremely easy to destroy a Kafka cluster by bringing a new, empty ZK server online with newer but incorrect data in its volume. ZK will happily trash the entire cluster thinking it has new instructions. How does that happen? I mean a new, empty ZK server with never data than the rest of the cluster? Also, please note that ZK is not meant to be a database, but a coordination service, it's guarantee is to have all nodes being always in consistent state and neither of its nodes allow to make any changes if there's no quorum. So if a new node somehow has more recent data with higher serial number it's expected that remaining nodes will sync to that.
- philjackson 9y agoI've had a break from Antirez' blog for a while - it's fun to go back and see how good his English has gotten!
- bbrunner 9y agoI've been using redis + resque[1] for a few side projects and I have to say I'm glad that streams are getting first class support in redis. I was always a little wary of hacking this sort of functionality on top of redis lists. It worked, but it sort of seemed a little bit fragile. [1] https://github.com/resque/resque https://github.com/resque/resque
- hyades 9y agoUnder what circumstances would one prefer Redis streams over Kafka and vice versa?
- zedpm 9y agoOne that immediately comes to mind is cases where Kafka is overkill. Kafka is a great tool, but there's a lot of overhead in setting up and maintaining it (e.g. Zookeeper), so if your throughput needs are low, it's a poor fit. Spinning up a Redis server is dead simple, and if you're already using Redis for other things, then there's no need to bring an additional tool into the mix.
- chicagobuss 9y agoGenuine question - Why does everyone seem to think running a zookeeper cluster is so hard? You can run it on three small VMs and basically forget about it. We didn't have any zookeeper experience at my last startup before we started using it for Kafka and we used a very simple puppet module to install it on three instances in each of our AWS regions. It really never gave us many problems in the several years since. Also, all the tooling around it is quite mature - there are great monitoring and management tools for probing at the internals which helped when we were dealing with more exotic kafka surgery.
- hueving 9y agoBecause nothing in the current hype cycle depends on zookeeper so people use that as a crutch for following hype.
- takeda 9y agoFor some reason Zookeeper is unjustly seem as uncool technology. I even seen it being blamed for issues that it had nothing to do with. People say that setting ZK cluster is a huge issue, yet they don't see a problem spinning etcd, or sentry nodes in case of redis. When I learned about ZK I was skeptic, didn't like that it was written in Java, but ZK proved to be extremely robust.
- 9y ago
- jkarneges 9y agoVery cool! We've been doing time series with Redis using sorted sets, referencing items by timestamps and integer offsets, and using the clock shift workaround described in the article. Having this kind of thing consolidated down into a few Redis commands would be handy. The API looks clean, too.
- notamy 9y agoAfter reading this I'm still not sure I understand. Time-series data I can see, but is there a use-case outside of this?
- noncoml 9y agoPoor-man’s Kafka?
- wyc 9y agoProjects tend to gain more and more functionality to match the new workloads they're being used to accomplish, and it must certainly be a difficult decision for project visionaries. Do I listen to my users and implement features that will solve their new woes, but in return accept increased complexity and higher learning barriers? Complexity sucks, but it's even harder to say no to users in pain. I wonder if this opens such projects up to disruption, in the original sense of the word. I've seen the same teams forgoing apache for nginx, then forgo nginx for haproxy when it matches their needs. With additional layers of complexity, there may be accompanying opportunities for "simple but good" projects to gain traction.
- barrkel 9y agoIt's also why software is a pop culture. Sophistication and completeness is seen as complexity and cruft by each successive generation, who start something new and simple. I don't think it's very avoidable. Tech is genuinely getting incrementally better, but it's usually in a sawtooth pattern.
- freshhawk 9y agoI'll agree that sophistication and completeness is often seen as complexity/cruft but it also always comes with actual cruft as well since improvement is incremental and breaking APIs is annoying. My favorite aspect of this cycle is when some features in the complex software become seen as so useful as to be required and standard, so when the new simpler version is created they have to figure out a novel way of providing that useful functionality in a simple and elegant way. And they do it, sometimes knowing they have made a significant advance and sometimes without knowing.
- scott_karana 9y agoDo you have any examples of that second case? Sounds too good to be true... but I'd love to be proven wrong :)
- jankotek 9y agoLook at the source code. Redis is very simple, single threaded, primitive networking... I would hardly call it complex even with this extra feature.
- poorman 9y ago> A final important thing to note about XRANGE is that, given that we receive the IDs in the reply, and the immediately successive ID is trivially obtained just incrementing the sequence part of the ID, it is possible to use XRANGE to incrementally iterate the whole stream, receiving for every call the specified number of elements. Love it. No need for an expensive SCAN command.
- orixilus 9y agothere's video[1] at the end explaining this new feature [1] https://www.youtube.com/watch?v=ELDzy9lCFHQ https://www.youtube.com/watch?v=ELDzy9lCFHQ
- femto113 9y agoI have a confusion about ID structure/format: The ID is composed of two parts: a millisecond time and a sequence number. The number after the dot is the sequence number, and is used in order to distinguish entries added in the same millisecond. Does this mean for example that 1506872463535.11 comes after 1506872463535.2 (because 11 > 2)? If so that means treating these as decimals (which will be easy to do inadvertently) will yield the wrong order (as would sorting them lexicographically). If so it seems like something other than a decimal point would be a better separator (colon perhaps).
- ihattendorf 9y agoWhat about regions that treat the comma as a decimal separator? I agree with the other commenter that this is no different than using periods in IPv4 addresses.
- anyfoo 9y agoAs a member of such a region, I suspect that those regions a well aware that the prevalent notation throughout programming uses '.' as a decimal separator. IPv4 addresses usually have 4 components and do not have an inherently fractional unit as their first component.
- mijamo 9y agoDisagree. '.' is used for version numbers everywhere, even when it's only 2 components (ex: Wordpress versioning, Django etc.) and it seems pretty clear that 1.11 > 1.2 there. Once the specification is stated clearly as it is right now it is not a problem. Another symbol could be used (for instance #) but I really don't see a need for that.
- IanCal 9y ago> What about regions that treat the comma as a decimal separator? A good reason not to change it from a period to a comma, not really relevant to whether to change it from a period to something else. > I agree with the other commenter that this is no different than using periods in IPv4 addresses. IP addresses have no useful concept of "before" or "after", whereas these do.
- luhn 9y agoI'm very excited about this. I've been eying HTTP EventSource for a while now, but there hasn't been a good solution for the backend broker. Kafka is overkill and Amazon Kinesis' pricing isn't viable if you have lots of topics. This fills the need perfectly and Redis is already part of my stack.
- hakanito 9y agoJust wanted to chip in and say HTTP EventSource has been really nice to work with, we've been using it in production for 1+ year
- bonkabonka 9y agoWhich polyfill do use for IE/Edge?
- rakoo 9y agoI don't know if you've ever looked at it, but CouchDB has had an EventSource endpoint for a long time now (http://docs.couchdb.org/en/2.1.0/api/database/changes.html?highlight=eventsource#db-changes http://docs.couchdb.org/en/2.1.0/api/database/changes.html?h...). CouchDB is extremely easy to install, use and maintain, and there's a number of public providers out there if you don't want to host everything yourself. As a more generic solution, there's also pushpin (http://pushpin.org/ http://pushpin.org/), which is the backend of fanout (https://fanout.io/ https://fanout.io/), so that may also be a nice addition to your stack if you want a more direct redis->clients link
- baileymiller182 9y agoCheck out nchan.io, an nginx module that does everything you need to connected EventSource to redis backed pub/sub.
- denkmoon 9y ago+1, good to see other developers out there using EventSource and nchan. We've been happily using both in production for almost a year now.
- grandalf 9y agoThis looks very cool. Can't wait to try it out.
- bootcat 9y agonice to know, they have added a new data structure !
- twic 9y ago> However a special ID of “$” means: assume I’ve all the elements that there are in the stream right now, so give me just starting from the next element arriving. I can already see lazy users just repeatedly reading $, and then dropping messages when they arrive faster than they read them. Might it be safer to instead have command to ask what the latest ID in the stream is? You'd start off by using that to work out where the streams are, then construct an XREAD command to read from there. To construct your next XREAD, it should be easier to update the IDs from the ones you just read, rather than fetching the latest IDs again. Maybe.
- jonny_eh 9y agoThe clear writing/documentation, concise API design, and clever implementation of antirez and the rest of the Redis team continues to amaze me. Redis is easily my most admired OSS project.
- EGreg 9y agoFunny, we also built a lot of our technology around the Streams concept. https://github.com/Qbix/architecture/wiki/Internet-2.0 https://github.com/Qbix/architecture/wiki/Internet-2.0
- Sinjo 9y agoIt's been a long time since I looked into this: is there now a way to configure a cluster of Redis instances such that you won't lose messages on node failure? If not, all the nice at-least-once delivery (or "effectively once" when you add message dedupe) you get with something like Kafka/Kinesis/GCP PubSub is gone. If not, either people's messages don't matter /that/ much (which is fine, just not great for most of my usecases at the moment) or everyone's in for another round of "oh shit, where did the data go?" Edit: Just in case we end up in CP vs AP datastore wars, please go read https://martin.kleppmann.com/2015/05/11/please-stop-calling-databases-cp-or-ap.html https://martin.kleppmann.com/2015/05/11/please-stop-calling-... At-least-once delivery requires neither CAP consistency (linearisability) nor CAP availability (any non-failed node must return a response in a non-infinite time), but is a very useful property!
- deleted 9y ago[deleted]
- atombender 9y agoLast I checked, neither Redis Sentinel nor Redis Cluster were linearizable systems; you get neither C, A or P. Redis Cluster failed Aphyr's Jepsen tests back in 2013. I not sure what the current status is, but I don't think the fundamental architecture has changed since then. With vanilla Redis master/slave replication, I believe the best way to avoid data loss is to set replication to be synchronous (it's async by default) so that slaves are always guaranteed to be in sync with the master, in case you need to promote (using Sentinel) a slave to master.
- antirez 9y agoHello, the streams have basically the same characteristics as any other Redis data structure, that is, from the POV of a local node, you can configure strong persistence on disk, but on node failures, you have basically different tunable amount of best effort consistency, it means that you cannot guarantee no messages are lost. So basically this means that you can: 1. Use the default asynchronous replication, and live with the fact (if the use case permits this) that on failover, the message did not yet received the slave that will be promoted. 2. Use WAIT to force synchronous replication to N slaves. This will not still make Redis ensure you in mathematical terms that the failover will pick a slave that received the message, under complex partitions, but narrow the real world failure models leading to losing data to more "unlikely" cases. Yet you have just best effort consistency but with better real-world outcomes. So Redis streams will be good choice if one of the above is acceptable.
- vikiomega9 9y agoWhy not just use sequence ID? I'm confused about why a timestamp is important. The sequence ID gives us ordering, is always guaranteed to be increasing.
- deleted 9y ago[deleted]
- antirez 9y agoBecause with the way stream IDs are conceived you also get time-based range queries for free. With time series this is very important in many use cases.
- vikiomega9 9y agoI see, so this composite structure is in lieu of having two distinct fields exposed in the API?
- ralusek 9y agoI love redis, and this API looks amazingly simple. I'm sure I can think up a good use case for this, but the only problem I have with it is that time-series log data of this nature is increasingly becoming the defacto source of truth in the various models resembling some version or another of event-sourcing. Obviously the general thinking is that event sourced time series data allows you to treat every other data source as derived state from the log data. If it gets written to the log, it's safe and that's the only real data that cannot be considered a redundant read layer. A common structure might look like: 1.) Event Sourced/Time Series Layer: Kafka/Kinesis>S3 takes in and saves log data 2.) Operational State Layer: RDBMS Constraints / Application Logic determine how operational state is derived from log data 3.) Indexing Layer: Query optimizations occur with redundant read layers in what is essentially all just various forms of indexing. This can be RDBMS index, ElasticSearch, MapReduce, Redis, etc. Redis' place has historically been at #3, for many reasons. Whether an application has an event-sourced layer in which their operational state is derived from log data or is actually considered the primary source is something that is hugely variable. I would say that most applications do not make a distinction between #1 and #2, and just write state directly to an RDBMS. But while the operational state may or may not be considered redundant, depending on the application, the indexing layer is almost guaranteed to be a redundant layer. The redundancy of the index layer means that Redis operating purely in-memory allows the whole thing to be blown away with no consequence. To move it two steps down the data model to the defacto source of truth is a monumental shift in responsibility. Redis as an in memory caching layer has only ever had me have a cursory awareness of its capabilities in terms of saving to disk, but I would think that a fundamentally different use case like this will have me taking a serious look at where that functionality is at today. All of this being said, there are plenty of use cases with kafka/kinesis which are done today which actually don't even save the log data at all, and just use them as an intermediary buffer to have multiple consumers on an event stream. There's also nothing stopping us from just having one of the consumers of this stream to be saving it to S3/Disk ourselves.
- dbattaglia 9y agoI would imagine this is best used for the "store a ton of events with some capped maximum size" Kafka use case (ie - realtime analytics, IoT data, etc). I just can't imagine using Redis as your source of truth for an event sourced system, especially without the partitioning and log compacting features that Kafka has. Still this seems like a pretty amazing feature to have in Redis, can't wait to start playing with it.
- callumlocke 9y agoMirror: https://webcache.googleusercontent.com/search?q=cache:https://blog.even.com/anti-perks-4e2e49e80983 https://webcache.googleusercontent.com/search?q=cache:https:...
- ricardobeat 9y agoCould the difference between `MAXLEN ~ 1000000` and `MAXLEN 100000` be handled internally, by marking the overflowing items as deleted until a whole block can be removed? Looks like this tombstone functionality is already planned, would make the API simpler.
- philsnow 9y agoI would suggest making the efficient behavior the default and let people use `= 1000000` when they know they really want the expensive but exact behavior.
- ricardobeat 9y agoSlightly off-topic, but could the blog be adjust a little for mobile reading? <meta name="viewport" content="width=device-width, initial-scale=1"> #content { max-width: 800px; } // replaces width: 800px seems to do the job, and also behaves better in narrow desktop browser windows.
- richardjennings 9y agoI am looking forwards to trying this out as an easy access cqrs / event sourcing entry point.
- Gigablah 9y agoSurely any discussion about the concept of logs and mention of Kafka should have a reference to this excellent article by Jay Kreps on LinkedIn: https://engineering.linkedin.com/distributed-systems/log-what-every-software-engineer-should-know-about-real-time-datas-unifying https://engineering.linkedin.com/distributed-systems/log-wha...
- eternalban 9y agoIt [OP] was a bit cringe worthy, frankly. Obviously the Kafka papers about 'reconsidering the log' and 'unified view' and all that have been out there for years now. The Kreps article is quite excellent and well worth the read.
- philsnow 9y agoIs there a reason to choose millis as the granularity instead of micros or nanos? Is it because there's a stronger expectation of machines in a cluster agreeing on what milli it is "now" vs the other granularities? I'm kind of thrown by the idea of putting the timestamp / stream-id in the XADD command, I would have thought the server would assign that, since one of the strengths of redis's single threaded nature is consistency: what's in redis is the truth. If you allow clients to specify timestamp, what happens when ntpd isn't running on some? I probably misread or misunderstood it. Could you allow specifying `$` as the timestamp to tell the server you want it to use whatever it thinks the current time is as the timestamp / stream-id?
- antirez 9y agoHello, the stream implementation does not need for the different servers (for instance master and its slaves) to agree about the time. Simply the server that receives the XADD command will generate the ID (and the time part of the ID) to attach to the item. All the other participants in the replication will accept the same ID, because clients will use "" to specify the ID, while the command is rewritten to slaves with a specific ID. Example, I run into the master: 127.0.0.1:6379> xadd stream * a 1 b 2 1506977609865.0 But this is replicated as (output of redis-cli --slave): "xadd","stream","1506977609865.0","a","1","b","2" So XADD allows to specify an ID just for replication / AOF pruposes, not because clients should actually specify an ID normally. However of clients really want to do that, they could but at the risk of getting errors, for instance: 127.0.0.1:6379> xadd stream 10.0 a 1 b 2 (error) ERR The ID specified in XADD is smaller than the target stream top item Redis will anyway not accept any ID which is smaller than the current top-item ID. The reason why it was chosen to use milliseconds instead of nanoseconds is because, for most applications to query for sub-millisecond ranges is likely not useful, so to see even larger numbers in the ID maybe is just unpleasant if not useful, however we are still in time to change this if there are good motivations. But being the time the one produced by the local host, after a failover the IDs are generated by another host. Milliseconds can still more or less match with good time synchronization, but nanoseconds? So it's like if this additional precision will be just used to store non-valid info.
- 9y ago
- tuna 9y agoAmazing ! nit: Change COUNT to LIMIT while it still time. Also can this primitive be used to replicate redis data instead of sentinel/cluster ?
- Goopplesoft 9y agoWhat sort of compression do the blocks undergo? E.g. does periodicity of the timeseries help reduce the the space of the 64/128bit timestamp? Gorilla[1] style compression would be great, although it'd likely make sub block level range queries tough. [1]http://www.vldb.org/pvldb/vol8/p1816-teller.pdf http://www.vldb.org/pvldb/vol8/p1816-teller.pdf
- antirez 9y agoHello, yes IDs are delta compressed so they use actually just a few bytes per entry (often just 2) instead of 16.
- silverwind 9y agoThe use of `+` and `-` in XRANGE seems inconsistent. Why not use `0` and `-1` like LRANGE?
- antirez 9y agoBecause they do not mean a position, but a special ID.
- lsiebert 9y agoIt seems that "$" is a special ID for the last message, as opposed to the last possible message. I would humbly suggest that "^" would be a suitable symbol for the first message in a stream. ^ and $ are used in regex (and vim) in a similar way. That way you could write "XREAD BLOCK 5000 STREAMS newstream ^" and get all the messages in a stream from the beginning, and then block until a new message comes in all with a single command. You would still be able to add a count if needed, to prevent client flooding.
- antirez 9y agoExactly, $ is the last message ID, + the greatest, - the smallest. I chose the dollar exactly because of regex assonance. However the corresponding ^ is kinda useless because with XREAD we specify the last ID we got, so it would result in not returning the first element of the stream. It means that it's more useful to specify just 0 in that case.
- jrochkind1 9y agoOne thing I like about this post is the story of how the feature came to be: Someone who understood redis very well, thinking about the problem over literally years, eventually resulting in a more targetted and goal-driven thinking, and even that "specification remained just a specification for months, at the point that after some time I rewrote it almost from scratch in order to upgrade it with many hints that I accumulated talking with people about this upcoming addition to Redis." I've been thinking about this sort of thing for a while, wanting to maybe call it "slow code" (like "slow food"). This is how actual quality software that will stand the test of time gets designed and made, _slowly_, _carefully_, _intentionally_, with thought and discussion and feedback and reconsideration. And always based on understanding the domain and the existing software you are building upon. Not jumping from problem to a PR to a merge. (And _usually_ by one person, sometimes a couple/several working together, _rarely_ by committee).
- manigandham 9y agoThought/discussion usually leads to quality. Time has nothing to do with it.
- dvt 9y agoOf course it does: thought/discussion take time.
- manigandham 9y agoIt seems people are confused. Of course everything takes physical time, but slowly thinking about something isn't any objective sign of better outcomes.
- dvt 9y agoI think the salient point is that thinking takes more time than not thinking. At least in my experience, the policy is generally to push features, not to think (slowly or otherwise).
- 9y ago
- vtuulos 9y agoRedis Streams could be a nice real-time counterpart to http://traildb.io http://traildb.io: For instance, use Streams in Redis to record data in real-time and periodically store it in TrailDBs for long-term archival and analysis.
- erulabs 9y agoI have one follow up question - is TTL a planned feature? Being able to set a TTL on the _stream itself_ and -also- on the messages would be extremely nice. While MAXLEN prevents a queue from being extremely large, I also want to remove "stale" data after a configurable time period. Use case: A log of network latencies, where a user might currently `XREAD` with a timestamp 10 minutes in the past, would be able to save on memory usage by expiring log entries > 10 minutes, and then being able to `XREAD STREAMS strm 0` and let Redis (and therefore the infrastructure, not my code, manage data retention). Also, how does this work re: evictions? Say a node is at max memory, will entire _streams_ be evicted, or (I hope) the oldest messages in the LRUed or LFUed queues.
- plasma 9y agoAre there plans for a "unique count" over an XRANGE? I currently use multiple Sorted Sets (one set every 5 minutes) and Union 30-60 worth to produce a "rolling window" of uniques. I can see an alternative where I just request a unique count of elements within an XRANGE.
- itaifrenkel 9y agoTwo comments on effectively once stream processing. 1. Consider adding an example for a stateful event stream processor client that saves the last read stream offset in redis, together with its current state and continues reading from that offset as an atomic operation. For example, a client that sums a stream of numbers, in order to have effectively once semantics would need to persist to redis the sum and offset together. 2. Consider adding a stream read deduplication example to mitigate clients that reinserted the same event twice. It is not clear how the client should behave if it didn't get an ack and it resents an event. What is the correct resending semantics so the reader would effectively dedup? What is the right data structure used to dedup message ids without consuming too much memory, etc...?
- itaifrenkel 9y agoThe consumer groups proposal breaks the FIFO abstraction of a stream by allowing multiple clients to process a single stream. Have you considered adding a semantic layer inside streams that allows each client to consume a substream? In effect the stream becomes multiplexed substreams. If substreams makes the design too complex... have you considered server side stream 403 semantics? When a stream is manually deprecated it enters an immutable state and provides a redirect response with a link to another stream. This would allow multiplexing and demultiplexing streams without changing the client implementations too much. For completeness I would state the obvious when fifo grouping is needed: 1. Scaling stateful event processing by splitting streams and adding more clients (CPU limit) 2. Scaling cross region replication by splitting streams and adding more tcp connections (network limit) 3. Handling more throughput by splitting a stream into two redis nodes (disk I/O limit)
- manigandham 9y agoDon't use consumer groups and every client will get a complete copy of the stream. What is broken with that?
- itaifrenkel 9y agoThe use case I'm referring to is when the client must be sharded to avoid CPU or network or disk bottlenecks.
- manigandham 9y agoThen that's exactly what consumer groups help with but it sounds like you want partitioning then - which is exactly what Kafka does but with a little more automation. Run multiple Redis instances and use a simple hash based on whatever key you want to route messages and get the throughput you need. It's probably never going to be part of the core Redis logic but should be possible as a module to do the routing when used in a Redis cluster.
- itaifrenkel 9y ago
- yawniek 9y agohere is an approach i successfully used before: since timestamps are read from clock the epoch could be a persisted value and a few nibbles could be used for the instance id and a sequential number. with that you get 64bit numbers for easy use and computation without loosing the time information at the cost of a simple transformation function. it simplifies the interface much and generally makes things faster. many clients will even not need that timestamp anyways.
- baystep 9y agoThis may actually solve an issue I was about to tackle, which is high-speed notification delivery. Currently I was going to do a nasty wrestle of PUB/SUB with Lists and blocking keys to try and get a cluster of message processing servers to digest notifications when they get added to the "queue". My purpose in this is notifications, as in actual push notifications for mobile/web/etc. This seems like it fits perfectly with what I had planned. Especially since this ensures message delivery instead of fire-and-forget as you mentioned. As a question though... is the XACK commands working in your current branch? Since that is key to my usage of ensuring message consumption.
- baystep 9y agoEssentially, I was going down this route if anyone's curious.... http://code.flickr.net/2012/12/12/highly-available-real-time-notifications/ http://code.flickr.net/2012/12/12/highly-available-real-time...
- deleted 9y ago[deleted]