5 ms·
Being on-call/carrying the pager for a complex and unstable distributed system is a great way to understand how to build a good one. Building a distributed syst
by i0exception 7y ago
Being on-call/carrying the pager for a complex and unstable distributed system is a great way to understand how to build a good one. Building a distributed system isn't hard per-se. It's almost always the non-determinism and the associated difficulty in debugging errors that's a problem.
- deleted 7y ago[deleted]
- gbuk2013 7y agoThis has been my experience. The pieces are simple, but being to able to understand what is going on in the whole system is not. Detailed, quality structured logging is very important, sent to a central location that can be queried (e.g Elastic Search / Kibana). I have come across several experienced developers who really don’t like he idea of verbose logging but in my experience it has been critically important for dealing with a mesh of micro services.
- stingraycharles 7y agoYep, you never know what logs you need until you need it. If there’s an economical factor coming into play (because proper telemetry will cost a noticeable amount of money), you can always choose to sample the more verbose logs.
- ignoramous 7y agoHonest Q: Isn't that's where tracers like AWS XRay and Google Dapper come in? And somewhere down the line, solutions like Envoy/istio try to make the whole ordeal manageable?
- gbuk2013 7y agoNothing cloud is relevant when your production network has no outbound access to the internet due to PCI. ;) Not sure about the other two (will check it out) but we have have apps in 4 programming languages across several teams. Structured (JSON) logging is the easiest way of sending data to one place from literally everywhere.
- dropofwill 7y agoYou can implement distributed tracing with logs if every player in the system is propagating the headers appropriately. Tools like NewRelic APM and XRay will take care of that bit mostly for you. Service mesh data planes have a wider scope, but they do make it much easier to implement distributed tracing (I think Istio includes this out of the box). In the end though these tools can’t replace app logs because they only let you reconstruct what happens in between services, not internal to them.
- gbuk2013 7y agoAgree 100% - logging is expensive, in terms of performance hit and storage, but it is totally worth it.
- QuickToBan 7y agoIn the medium to long term, the absence of logging is always more expensive than its presence.
- alexbanks 7y ago"How does this even get invoked" and "Where would I go to find these logs" are the two questions that, if the answers are not obvious/obviously documented, there is a big problem with that particular system. My current company for whatever reason does not value logging at all, to the point where for lots of our systems there just aren't logs/nobody has ever bothered to look for them. It's pretty astonishing to me, a person who has valued logging as one of the highest-value mechanisms to increase tracibility.
- gbuk2013 7y agoPrior to moving to software development, I worked 5 years in Support and consequently I have a murderous hatred of anyone writing software that is hard to debug and logging is a huge part of that. I was actually greatly inspired by a product I used to support which had such detailed logs that you could troubleshoot almost any issue from the (rather huge) log dump. I try to do the same now.
- yjftsjthsd-h 7y agoI really wish all developers had to do support at the beginning of their careers and so they understand how to build things that the ops team can work with. Devops/SREs are one approach to this, but I'd really like to see it be more widespread.
- alexanderdmitri 7y agoI share your enthusiam for good logs, so am curious what made this product's logging great from a support perspective?
- cbetti 7y agoI wonder if they simply haven't seen a strong application of logging, and therefore don't know what they're missing. Is there an opportunity in your daily work to sneak some good logging into a complex subsystem to show how valuable it can be?
- QuickToBan 7y ago
- Quekid5 7y ago> Building a distributed system isn't hard per-se. I think this may be a bit overstated. It's certainly true that most of the algorithms, etc. are -- if not necessarily simple -- at least understandable/understood generally. IME the problem is really all of the engineering around the algorithmic stuff. One thing which I don't think is really well understood yet is different levels of Consistency. There are a lot of trade-offs to be made here, but generally I find that it's really hard to help people understand what those trade-offs are and how they could impact the UX and business. (After that there's things like how to handle configuration changes safely, how to properly dispose of nodes, etc. etc.)
- bitL 7y ago> Building a distributed system isn't hard per-se. Well, we just have some ugly proofs that building a proper distributed system is impossible, leading to many trade-offs that trigger one or the other customer at times... There aren't that many areas in computer science which are as difficult as building reliable scalable distributed systems.
- pfranz 7y agoWhat you're saying is true, but I think a big problem in general is that while bespoke problems often require bespoke solutions, there's a lot of common problems out there and using a common solution goes a long way. They make the pitfalls and limitations obvious, help greatly in staff turnover and growing your team, and make things more intuitive and consistent. I see a lot of systems that work pretty well when the author maintains it, but someone else jumps in and has no idea what server some part is running on, where the credentials might be, or where the log file is.
- fma 7y agoThis is why I roll my eyes when I see resumes, especially from senior engineers who job hop every year or two. They may have never deployed their system or even they did, have enough support time to learn how supportable their system is, how their system performs, how modular it is when requirements change, etc. Resume looks great, though.
- QuickToBan 7y agoI have to dispute this. A few months is sufficient to appreciate the stated concerns. A lot can happen in a few months. I am proof of the same. I think under two months is too short, however.
- afarrell 7y agoI think successfully doing this could be a good way, but I don’t think just carrying the pager would teach you much. I think unless you are able to understand and debug the system and handle alerts effectively, this experience would just leave you frustrated.