6 ms·
Monitoring demystified: A guide for logging, tracing, metrics
- FrontAid 6y agoRecently, I was searching for a service which offers those functionalities on a very basic level. I tried several options and was really disappointed with all of them. The only one that I found to be usable was https://logdna.com/ https://logdna.com/. I've now been using it for a couple of weeks and it works OK. It offers logging, alerts, metrics/dashboards, and some other things. And all that for a reasonable pricing.
- buro9 6y agoA lot of excellent information in that blog post and linked from it... but if you're wondering where to start: 1. Write good logs... not too noisy when everything is running well, meaningful enough to let you know the key state or branch of code when things deviate from the good path. Don't worry about structured vs unstructured too much, just ensure you include a timestamp, file, log level, func name (or line number), and that the message will help you debug. 2. Instrument metrics using Prometheus, there are libraries that make this easy: https://prometheus.io/docs/instrumenting/clientlibs/ https://prometheus.io/docs/instrumenting/clientlibs/ . Counts get you started, but you probably want to think in aggregation and to ask about the rate of things and percentiles. Use histograms for this https://prometheus.io/docs/practices/histograms/ https://prometheus.io/docs/practices/histograms/ . Use labels to create a more complex picture, i.e. A histogram of HTTP request times with a label of HTTP method means you can see all reqs, just the POST, or maybe the HEAD, GET together, etc... and then create rates over time, percentiles, etc. Do think about cardinality of label values, HTTP methods is good, but request identifiers are bad in high traffic environments... labels should group not identify. Start with those things, tracing follows good logging and metrics as it takes a little more effort to instrument an entire system whereas logging and metrics are valuable even when only small parts of a system are instrumented. Once you've instrumented... Grafana Cloud offers a hosted Grafana, Prometheus metrics scraping and storage, and Log tailing and storage (via Loki) https://grafana.com/products/cloud/ https://grafana.com/products/cloud/ so you can see the results of your work immediately. If it's a big project, you have a lot of options and I assume you know them already, this is when you start looking at Cortex and Thanos, Datadog and Loki, tracing with Jaegar.
- kostarelo 6y ago> Don't worry about structured vs unstructured too much, just ensure you include a timestamp, file, log level, func name (or line number), and that the message will help you debug. if you include all these information and the logs are not structured, you won't get much information out of them.
- deepGem 6y agoFrom the post "The key, he says, is using the right transaction identifiers so that calls can be traced across components, services, and queues". I think this is a key feature not many people implement especially in today's world of over blown micro services, having a transaction id from the time the request hits the reverse-proxy till the database write is so helpful in debugging, saves a ton of time.
- KaiserPro 6y ago100% if you manage to get opentrace to work, it is a brilliant debug tool
- eatonphil 6y agoYou can also just propagate a uuid throughout your system. I've used both uuids and opentrace. Just don't get hung up if opentrace seems too complex.
- dkarl 6y agoI agree with this wholeheartedly. You can even define a standard and let downstream services opt in over time. Simple wins like this should not be put off because "someday" we're going to implement a complex distributed tracing solution.
- dillonmckay 6y agoSo, in terms of full implementation, is the constraint the dev time to implement, or getting various factions to agree to something? I guess, is it political, or technical? Asking for a friend. Thanks!
- KaiserPro 6y agoA few things I have learnt along the way: Logs are great, but only once you've identified the problem. If you are searching through logs to _find_ a problem, its far too late. Processing/streaming logs to get metrics is a terrible waste of time, energy and money. Spend that producing high quality metrics directly from the apps you are looking after/writing/decomming (example: dont use access logs to collect 4xx/5xx and make a graph, collate and push the metrics directly) Raw metrics are pretty useless. They need to be manipulated into buisness goals: service x is producing 3% 5xx errors vs % of visitors unable to perform action x Alerts must be actionable. Alerts rules must be based on sensible clear cut rules: service x's response time is breeching its SLA not service x's response time is double its average for this time in may.
- say_it_as_it_is 6y agoThe difference in metrics seems to be proportional to the level of understanding about how the organization works.
- darkwater 6y agoOn the rest I agree but on > Raw metrics are pretty useless. They need to be manipulated into buisness goals: service x is producing 3% 5xx errors vs % of visitors unable to perform action x I think in general the business goals metrics are OK but you still need to keep lower level metrics as well, otherwise it would be more difficult to pinpoint the exact failure, you will just know that a % of visitors is unable to perform action X. In a moderately-complex system a user-level action X is probably composed by several low-level services.
- KaiserPro 6y agoI agree wholeheartedly. I was trying to get across that just because you collect metrics it doesn't make them useful. I encourage people to generate metrics for everything, we can always join them together later to make something useful. I think what I should have said is: "Collect metrics for everything, but be sure to display them is a way thats relevant to the customer"
- 6y ago
- say_it_as_it_is 6y agoIs there an open source solution for processing streams of structured and unstructured logs and routing then onward? I see solutions for moving logs to elastic or Kafka but nothing for evaluating the log.
- malechimp 6y agoMaybe you're looking something like this https://docs.tremor.rs/ https://docs.tremor.rs/
- dig1 6y agoRiemann [1]. You can create custom endpoint for accepting almost any kind of messages, logs or data. [1] http://riemann.io/ http://riemann.io/
- ysoft 6y agoI haven't found anything, we are moving to hosted Humio really soon. It uses kafka
- onefuncman 6y agoThe "OG" in the space is collectd, which is still my favorite choice if you are responsible at the operating system level: https://collectd.org/wiki/index.php/Chains https://collectd.org/wiki/index.php/Chains https://github.com/elastic/logstash https://github.com/elastic/logstash was one of the first modern approaches. I started using it less the more often I ran into JRuby related bugs. https://github.com/trivago/gollum https://github.com/trivago/gollum is my pick from the golang ecosystem. There are many more variants depending on how much complexity you are trying to apply. If you need to apply machine learning models, for example, you're probably going to end up with something similar to Apache Storm, though I don't know if it's operational story has improved enough to consider it over other alternatives, I lost track years ago between Apache Spark and the half dozen other stream processing projects.
- a10c 6y agohttps://vector.dev/ https://vector.dev/ sounds pretty close
- 6y ago
- waihtis 6y ago> Logging is critical to detecting attacks and intrusions. Yes, but not universally - and just collecting logs will not take you far. Logging everything and trying to approach security via the ’collect all data’ is both expensive and inaccurate, and one of the major inefficiencies in modern cyber.
- onefuncman 6y agoThis is done efficiently at scale by both Cylance and Crowdstrike, but is certainly only one part of a defense in depth strategy. There are viable products around human threat hunting which would be impossible without a 'collect all the data' component.
- waihtis 6y agoYou are correct, and this is the key part - what % of organizations have money, skills and people to build a robust enough capability around threat hunting, for example? I’ve been super lucky to meet various orgs and their security in all geographies and many industries and my gut feeling is 1 out of 10 teams.
- EricE 6y agoSecurity Onion does an amazing job at collecting and correlating, especially for an open source product. The traditional trade of with Open Source is there - a bit of up front effort for longer term value.
- notmalc 6y agoNice
- secondcoming 6y agoWe log extensively. Here are some of my thoughts it - at least in C++, the requirement to be able to log from pretty much anywhere can lead to messy code that either passes a reference to your logger to all classes that might possibly need it, or you've got an extern global somewhere. Yuck. - logging can enable laziness. Being able to log that something weird happened can be considered a sufficient substitute for proper testing. - logs are only as useful as the info they contain. This can mean state needs to be passed around all over the place just so that it can all be eventually logged on one line (it saves your data team from having to do a 'join') - if your logger doesn't support cycling log files it's useless. If something goes wrong you can easily fill a disk.
- viraptor 6y agoI'd disagree with 2 and 4. 2. Given a large enough system you will encounter situations where the only action you can take is to log "this really shouldn't happen" and try to roll back as cleanly as possible. This may be due to either complexity or a bug manifesting in a layer completely different than where it occurred (I've seen a null reference crash on "if(foo) foo->bar();" in the past) 4. I believe loggers should ideally know as little as possible about your logs. Logs can be rotated externally, can be buffered and sent to other hosts without touching the disk, can be ignored. Ideall, the system should care, not the app.
- mmkos 6y ago> I've seen a null reference crash on "if(foo) foo->bar();" in the past References can't be null. Regardless, that's a valid check for a null pointer and I don't think what you wrote is at all possible (unless maybe in some multithreaded scenario?).
- bradstewart 6y agoExpanding on your second point, logging is also not a substitute for proper error handling.
- gnufx 6y ago> the requirement to be able to log from pretty much anywhere can lead to messy code Ah, Milewski's example of insight from the supposedly useless mathematical stuff: https://bartoszmilewski.com/2014/12/23/kleisli-categories/ https://bartoszmilewski.com/2014/12/23/kleisli-categories/ (and the corresponding lecture video).
- kasey_junk 6y agoIt’s weird to see the stuff by Jay Kreps (of Kafka ~fame~) listed in the logs section. His writing is specifically _not_ about logs the observability tool, but logs the data structure such as you’d see at the heart of a database.
- rollulus 6y agoVery true. Jay Krep's log is completely unrelated to the topic of this article. This added to my feeling that this "guide" is rather a collection of fragments put together without a real understanding of the subject from the author.
- aloknnikhil 6y agoNo. The original Kafka paper does talk about logs in the observability sense as a premise to solve the aggregation problem. https://cs.uwaterloo.ca/~ssalihog/courses/papers/netdb11-final12.pdf https://cs.uwaterloo.ca/~ssalihog/courses/papers/netdb11-fin... > There is a large amount of “log” data generated at any sizable internet company. This data typically includes (1) user activity events corresponding to logins, pageviews, clicks, “likes”, sharing, comments, and search queries; (2) operational metrics such as service call stack, call latency, errors, and system metrics such as CPU, memory, network, or disk utilization on each machine. Log data has long been a component of analytics used to track user engagement, system utilization, and other metrics. > We have built a novel messaging system for log processing called Kafka [18] that combines the benefits of traditional log aggregators and messaging systems....Kafka provides an API similar to a messaging system and allows applications to consume log events in real time.
- kasey_junk 6y agoA quote from the LinkedIn blog post linked in the article: “But before we get too far let me clarify something that is a bit confusing. Every programmer is familiar with another definition of logging—the unstructured error messages or trace info an application might write out to a local file using syslog or log4j. For clarity I will call this "application logging". The application log is a degenerative form of the log concept I am describing”
- xondono 6y agoAm I the only one that can’t reach the “save and exit” privacy button on mobile? It’s hard for me to think that this is not intentional when the “Accept all” is usable but the alternative isn’t...
- anderspitman 6y agoIf you don't need all the fancy metrics, and just want something simple to keep an eye on your services, alert you if they fail, and automatically restart them, check out my stealthcheck service. It's all of 150 lines of free range, 0-dependency go: https://github.com/anderspitman/stealthcheck https://github.com/anderspitman/stealthcheck
- gnufx 6y agoI never see an important system management principle brought up: If you get a user complaint (for some value of "user") and not an alert, you should fix the monitoring system so that you don't get another occurrence of it or related problems. Obviously that's within reason, depending on the circumstances; the effort might not be worth it.
- dig1 6y agoThe Art of Monitoring [1], covers most of these stuff in a unified manner. You are introduced to some basics (push vs. pull monitoring), then proceeded with simple system metrics collection (cpu, memory) via collectd, then goes to logs ingestion and ends up extracting application-specific metrics from jvm and python applications. I highly recommend it, even for seasoned professionals. [1] https://artofmonitoring.com/ https://artofmonitoring.com/