3 ms·
Disclaimer: I work for Lightrun, a dynamic instrumentation (read: add logs at runtime) tool. Logging is a surprisingly pricey bit of observability at scale. I
by tomgs 4y ago
Disclaimer: I work for Lightrun, a dynamic instrumentation (read: add logs at runtime) tool.
Logging is a surprisingly pricey bit of observability at scale.
If you're doing anything highly-transactional you're pretty much guaranteed nowadays to get major observability bills (for things like ingestion, transmission, storage and analysis/querying).
It also creeps up on you - you're used to getting billed incessantly for cloud stuff, monitoring doesn't sound too pricey. Until, well, it is.
Can't help but point out that this is static logging, meaning logging added during development. This stems from an the approach colloquially referred to as "log everything, analyze later", rather than more disciplined, case-specific logging.
We're used to add "just-in-case-logs" to ensure that we've got ourselves covered when the sh*t hits the fan, but we rarely look at them.
An alternative approach would be using a tool, like the one Lightrun[0] builds (see disclaimer above) to add Logs in real-time to applications.
This means that instead of logging IN ADVANCE (i.e. during development) you can add logs when and where you need them, in real-time. The tool works using a variety of dynamic instrumentation techniques (depends on the runtime), and currently supports Java, Node.js & Python (.NET soon).
It can also pipe these logs into Loki, Elastic or SigNoz since they can be dumped like normal logs, right into the stdout.
In any case, we're seeing mass reduction in static logging (up to 60% in volume, 40% in cost) by going dynamic instead of mostly static. You'll always have to log some stuff for forensics and traceability, but dynamic instrumentation removes a lot of the dead weight.
[0] https://lightrun.com https://lightrun.com
- rocmcd 4y agoCan you explain in more detail how Lightrun works? It sounds like monkey-patching in real-time from what I can gather, which is neat but surely comes with some kind of overhead.
- wstuartcl 4y agoyeah even if they are wrapping with near noop guards all over the ast/codebase as its being compiled and the secret sauce is just to tickle the guard to log at those specific points dynamically you would think this would have huge overheads for the no log states as those guards get bypassed for hot code blocks. (kind of like running in a debugger env). They note that this is patented -- I wonder how. Every way I can imagine this being implemented in the languages they support seems to be clearly prior well known art. Patent office granting more "RED -- we own compressing raw" type patents? Looking at the patent seems like they are monkey patching either dynamically or into ast injected nodes at compile time. what the heck is the PO doing. also who in their right mind would open their prod services to an external party for code injection lol.
- tomgs 4y agoGreat questions form all ends. Let me try and clarify. This is not monkey-patching or hot-swapping, but a different approach - we call it dynamic instrumentation. This amounts to having an agent (an SDK/library, in essence) perform the addition of Lightrun Logs, Snapshots and Metrics, then (potentially) pipe the information where it needs to go. We've got a nice diagram here [0]. I think that the mechanism itself - which changes by runtime, naturally - is well explained in the link above. However, the core security mechanisms we have are enabled by another component of the agent we dub the Sandbox. In essence, we've got a (patented) way to verify everything that we do at runtime. That means we ensure that each evaluated expression does not have side effects (like changing the value of a member of an array, or editing a variable value), every metric we could think of is throttled and rate-limited (that includes our usage of CPU, RAM, the rate of I/O and a bunch of other things). Given this sandbox mechanism, and the way the networking requirements look like (again, look at [0] - no need to open a debug port / inbound ports, and a pretty agnostic deployment model) I think we've got a pretty robust defense layer against a variety of failures. Also see my comment a few comments above regarding data security. [0] https://docs.lightrun.com/more-about-lightrun/#how https://docs.lightrun.com/more-about-lightrun/#how
- srikanthccv 4y agoDo you have SDKs or some client libraries for lightrun on GitHub that I can look at?
- tomgs 4y agoGot a free tier you can play around with at [0] , and a few examples you can check out over at [1]. We've also got a zero-config (in-browser!) version of the whole thing [2], using code-server [3]. [0] https://lightrun.com/free https://lightrun.com/free [1] https://github.com/lightrun-platform/lightrun/tree/main/examples https://github.com/lightrun-platform/lightrun/tree/main/exam... [2] https://playground.lightrun.com https://playground.lightrun.com [3] https://github.com/coder/code-server https://github.com/coder/code-server
- krembo 4y agoIsn't that equivalent to setting the debug level in the code and scaling in/out the amount of data as needed with a push of a button?
- tomgs 4y agoWell, what if there was no log there to begin with? If you've got a bunch of logs labelled as INFO, WARN and ERROR in your system, and you tweak the debug level - you get more detailed debug logs as need be. This will stream EXISTING logs to stdout and then to your favorite APM. However, if the exact bit of information you wanted isn't there (that variable isn't logged, that specific piece of code isn't instrumented so you can't know it was reached, etc...) you're stuck. In addition, sometimes (often) it's hard to correlate the exact path the code took, since it's not clear which condition or class or package or server were actually involved in the process of execution. This tool enables you to conditionally log just what you need in real time - basically "paint a path" through the code at runtime. Hope that's clearer, can elaborate more if need be.
- twic 4y agoCan Lightrun dynamically add logging to an app in the past? If not, i don't really see this value of this.
- tomgs 4y agoLightrun add logs at runtime, which means you'll add a log and it will be integrated into the stream of logs emitted from the application (or streamed to your IDE plugin) - it's not a time-travelling debugger, if I understand the question correctly:)
- ketchupdebugger 4y agoThanks for this! This is really cool, I have a few questions. Can Lightrun prevent certain fields from being logged? Some fields shouldn't be logged such as credit card numbers, secrets, username etc. How can we prevent devs from accessing things they are not supposed to? Can this replace continuous profilers as well? Whats the reason this is not getting more adoptance? AFAIK most companies are either eating the cost of logging or managing it with sampling. This seems way better than either of those options, so whats the downside?
- tomgs 4y agoPer your first question - protecting sensitive information - we've got PII redaction and Blocklisting that cover a variety of potential cases of data leakage [0]. Regarding profiling - we're not actively a profiler, but we do offer a set of code-level metrics you can use to do performance analysis and detect bottlenecks [1]. Regarding adoption - we're doing OK:) I think this is a new approach, one that is quite different than what developers are used to. But it's also one of those things that - as you mentioned - make so much more sense than the alternatives, and with costs rising it's really even more evident than ever. If you're interested in the cost of logging specifically we've got it broken down in a study [2] we recently released. [0] https://docs.lightrun.com/data-security/ https://docs.lightrun.com/data-security/ [1] https://docs.lightrun.com/metrics/ https://docs.lightrun.com/metrics/ [2] https://lightrun.com/resources/lightruns-economic-impact-on-enterprise-logging-observability-costs/ https://lightrun.com/resources/lightruns-economic-impact-on-...