7 ms·
Structured logs are the way to start
- m-a-r-c-e-l 2y agoIMO manual logging and especially being consistent is hard in a team. Language skills, cultural background, personal preferences, etc. Much better is a well thought of error handling. This shows exactly when and where something went wrong. If your error handler supports it, even context information like all the variables of the current stack is reported. Add managable background jobs to the recipe which you can restart after fixing the code... This helps in 99.99% of all cases.
- gianarb 2y agoCiao! This is all gold thank for sharing! I agree that consistency is the key. But it is the key in many fields this is why well-designed abstractions or continuous integration exists. To enforce consistency. Error handling as well can be very helpful to communicate what your system is doing, but errors are not the only state you want to look for. In theory, but it is something that I didn't see used too much a logging library can be wrapped into an abstraction where useful to enforce consistency. For example if wrap your library in something that conventionally sounds like "ThisLoggerIsCriticalDontMoveItAsYouWillDoWithOtherLogs(logger)" you are communicating something more about how that.
- tail_exchange 2y agoContextual logging makes structured logging even more powerful. For example, you can attach an ID to the contex of an http request when you receive it, which then gets logged in every operation that is performed for serving that request. If you are investigating what happened for a specific request, then you can just search for its ID. This works with any repeatable task and identifier, like runs of a cron job, and user ids.
- ManuelKiessling 2y agoAdding userId, clientId, sessionId, and requestId to every log line and every event in the data warehouse is one of those „one weird tricks“ that actually does improve your life.
- bigmattystyles 2y agoUntil GDPR comes for you.
- XorNot 2y agoExcept like, not really? If you need to remove someone's data for GDPR reasons, then "if match == userId then delete" is pretty straightforward in your log aggregation store.
- jcrites 2y agoMany log aggregation stores are not optimized for performing row-level updates or deletes like this. In my experience, the majority of log aggregation stores are immutable and support primarily time-based retention only. (Though perhaps one can meet compliance needs by keeping these logs only for a fixed maximum period of time, e.g. 30 days, and keeping only appropriately anonymized data longer.)
- hipadev23 2y agoA log aggregation store that can’t handle deletes in 2024 is a product that shouldn’t be utilized. GDPR and similar redaction laws are not new. Efficient or fast is not a requirement for GDPR, so it can happen slowly and in the background just fine.
- randerson 2y agoA log aggregation store that can handle deletes is a security and compliance problem. Try proving to an auditor that a hacker couldn't have hacked in and then covered their tracks by deleting the logs.
- gnabgib 2y agoServer seems to be struggling? https://archive.is/mhm4u https://archive.is/mhm4u
- hipadev23 2y ago6-paragraph text post can’t show due to a database connection. Come on guys.
- bgdkbtv 2y agoFamous HN hug of death
- hipadev23 2y agoIt’s wild to me how quickly poorly configured servers succumb to a bit of read-only easily-cached traffic.
- cqqxo4zV46cp 2y agoIf this is “wild” to you then you are incredibly, incredibly green behind the ears. This is a very common reality. What you’re really saying is “I could do so much better than this”, which…good for you?
- denysvitali 2y agoThe problem here (IMHO) is using Wordpress to serve a blog. WordPress is fantastic for its WYSIWYG features, but it's probably one of the worst in terms of performance (especially w/o caching). This article could have been an HTML (or even Markdown) page...
- curt15 2y agoWhy should a DB even be involved here? Can't the content be displayed by a static website?
- cqqxo4zV46cp 2y agoWhat do you want someone to say to this? It’s incredibly easy to sit back and suggest changes to infrastructure. Smugness aside, nothing being said here is going to make the site come up.
- akira2501 2y agoMy habit lately has been to have a "request event" object that picks up context as it works it's way through the layers and then is fully saved to disk referenced by it's unique event number. In Go this is usually just a map. These logs are usually very large and get rotated into archive and deletion very quickly. Then in my standard error logs I always just include this event ID and an actual description of the error and it's context from the call site. These logs are usually very small and easy to analyze to spot the error and every log line includes the event ID that was being processed when it was generated.
- bob1029 2y agoI think SQLite is maybe the best option if you can get a bit clever around the scalability implications. For instance, you could maintain an in-memory copy of the log DB schema for each http/logical request context and then conditionally back it up to disk if an exception occurs. The request trace SQLite db path could then be recorded in a metadata SQLite db that tracks exceptions. This gets you away from all clients serializing through the same WAL on the happy path and also minimizes disk IO.
- siddharthgoel88 2y agoFor Java applications, we built a structured logging library which would do a few things - - Add OTel based instrumentation to generate traces - Do salted hash of PII (injected in plain text by API Gateway in each request) like userid, etc to propagate internally to other downstream services via Baggage - Inject all this context like trace-id and hashed PIIs into log - Have Log4j and Logback Layout implementations to structure logs in JSON format Logs are compressed and ingested to AWS S3 so it is also not expensive to store so much logs to S3. AWS provides a tool called S3Select to search structured logs/info in S3. We built a Golang Cobra based cli tool, which is aware of the structure we have defined and allows us to search for logs in all possible ways, even with PII info even without saving. In just 2 months, with 2 people we were able to build this stack and integrate to 100+ microservices and get rid of Cloudwatch. This not just saved us a lots of money on Cloudwatch side but also improved our capability to search to logs with a lot of context when issues happens.
- e3bc54b2 2y agohey, we're in pretty similar place logging wise, and I would really like to know more about your solution. If at all possible, I'd like to understand your rationale and implementation architecture more.
- siddharthgoel88 2y agoNext month I will be publishing a blog on this topic. I will share the link here as well.
- foota 2y agoI was just talking to some acquaintances the other day where I was asking them what they used for structured logging and they looked at me like "what's that" and I remembered people don't do it everywhere.
- jdwyah 2y ago"A breakpoint for logging is usually scalability, because they are expensive to store and index." I hope 2024 is the year where we realize that if we make the log levels dynamically update-able we can have our cake and eat it too. We feel stuck in a world where all logging is either useless bc it's off or on and expensive. All you need is a way to easily modify log level off without restarting and this gets a lot better.
- rendaw 2y agoThat's probably true for some uses of logging, but for information about historical events you're stuck with whatever information you get from the log level you had in the past.
- langsoul-com 2y agoIs there anyway to replay an error at the moment of a crash? I found just logs alone aren't that useful, it still takes eons to find out wtf happened. Almost like a local breakpoint debugger on crash, but for prod.
- gettodachoppa 2y agoIt's trivial in C/C++ due to GCC being a first-class citizen in Linux, but idk how it's done for interpreted languages, Java, etc. If anyone can chime in I'm curious. In C or C++, you just run 'ulimit -c unlimited' in your shell before running your program. When it crashes, a GDB-friendly core dump is generated. Then you can load it in gdb ('gdb myexecutable mycoredump'), and it takes you to the exact line where it crashed, including showing you the stack trace, letting you view local variables at every frame of the stack, etc. Every C++ IDE supports loading a core file, so it's literally an interactive debugger at the time you most need it. It's a life-saver. Keep in mind you have to compile with debug symbols enabled to be able to make sense of the coredump. However, you can then strip your binary, as long as you keep an unstripped copy around to help you with debugging.
- SassyBird 2y agoThis has nothing to do with GCC being a first-class citizen in Linux. It’s a kernel feature. The kernel doesn’t care which compiler or debugger you’re using. You can dump core of any process regardless of the language it’s written in. Every modern OS supports that.
- telotortium 2y agoOne great thing about Go is its built-in structured logging package "log/slog" since Go 1.21. Not only can you output in multiple structured formats, but Go also widely uses its `context.Context` type to pass request-level information, so you can easily attach requestID, sessionID, etc.