4 ms·
Since this is going to get some attention and might attract some welcome naysayers, it might be a good time for me to ask for resources if I do want to use Kafk
by ashkankiani 6y ago
Since this is going to get some attention and might attract some welcome naysayers, it might be a good time for me to ask for resources if I do want to use Kafka, specifically edge cases that I won't see as a newcomer until I actually try to use it in production.
I've already done my due dilligence and I consider Kafka to be one of the potential solutions I could use for a service, and while I've gotten started using it fairly easily in a dev env, there's obviously things I could miss. Any pointers from experienced users would be appreciated.
- gkop 6y agoIn case you haven't already read it, do read the classic LinkedIn "Log" article: https://engineering.linkedin.com/distributed-systems/log-what-every-software-engineer-should-know-about-real-time-datas-unifying https://engineering.linkedin.com/distributed-systems/log-wha... I just learned Kafka this year to augment an existing system at work. One thing I have found is, there's a good amount of smarts in the Kafka clients. Look carefully at the Kafka clients available for the runtimes you work with, and evaluate them for robustness, features, and development activity (sadly, for example, kafka-node seems to have lost all momentum).
- mr_tristan 6y agoYeah, I found the article to be a little annoying in that it's mostly attacking WeWork. My recommendation for anyone new to the area, is the streaming systems book: http://streamingbook.net/ http://streamingbook.net/, which came after two excellent blog posts: https://www.oreilly.com/radar/the-world-beyond-batch-streaming-101/ https://www.oreilly.com/radar/the-world-beyond-batch-streami..., and https://www.oreilly.com/radar/the-world-beyond-batch-streaming-102/ https://www.oreilly.com/radar/the-world-beyond-batch-streami... Basically, if you have an unbounded data problem, and you might have something that is _practically_ unbounded, that book really helped me gain insight to describe what we used Kinesis for (which is really similar to Kafka in a lot of ways). Thinking about event vs processing time, ordering and windowing requirements, etc, were really helpful in thinking about my own data problems. In the end, the author of this article claimed that WeWork just simply didn't have much of an unbounded data problem. Which I generally agree with. But simply saying "you can do this in PostgreSQL" isn't really a great takeaway, and, it really feels that was the message being reinforced here.