6 ms·
I think some of the posts here miss a little bit of the context as to why things like this happen in the first place. It's only in the last handful of years th
by polygotdomain 6y ago
I think some of the posts here miss a little bit of the context as to why things like this happen in the first place. It's only in the last handful of years that a stack for logging has really become mainstream. Chances are a lot of these types of logging solutions predate that and used whatever persistence technology was readily available. Writing to files on web servers can be a pain, and these logs will have to be queried at some point, so storing it in a database is not a bad idea, especially when better options have only recently become available.
The problem is that relational SQL is bad for logs, but by the time it gets to the scale where it's problematic that there's too much volume in the logs to make anything "easy". Simultaneously there's a lot of business value in that log data that you don't want to lose.
Yes, SQL's a poor fit for logs, but it's a better fit then a lot of other things, including not logging at all. Better solutions exist, but they don't exist in a bubble, and there's a cost to integrating them and migrating to them. A lot of these comments seem to be judging a technology decision based solely on hindsight without realizing that there are legitimate reasons for logging to SQL.
- amichal 6y agoTotally agree. I would be curious if anyone who is saying why didn't you use XYZ has actually migrated 500 billion rows of queryable log data from one tech to the other. For everyone complaining about PII. The article shows a 6 year log retention which fits with many regulatory requirements for data retention. We forget that a lot changes in 6 years on what makes sense to do. ELK stack according to elastics history page became a real thing in 2015.
- user5994461 6y agoI've done with a hundred billion rows. It's not too difficult but it takes a lot of time. Note that elasticsearch scales really easy and really well, I have an article on a small setup many years ago for 12 TB in a small company https://thehftguy.com/2016/09/12/250-gbday-of-logs-with-graylog-lessons-learned/ https://thehftguy.com/2016/09/12/250-gbday-of-logs-with-gray... The challenge is, a migration would require a lot of hardware resources (storage mainly) to setup a new cluster and the poster couldn't get one damn hard drive to work with, so any migration would have been a death march from the start. Knowing that and that the original system was out of disk and had no backup of any sort. I would personally question the effort to try to keep all these old logs.
- jeffbee 6y agoI disagree with the premise of your statement. It's typical that a log will be accessed zero times. Collecting, aggregating, and indexing logs is usually a mistake made by people who aren't clear on the use case for the logs.
- j88439h84 6y agoWhat is the use case for logs?
- jeffbee 6y agoThere isn't a universal one. If you don't have a concrete one in mind, you shouldn't produce the log at all.
- nickpeterson 6y agoI appreciate the zen-like nature of this advice, but I think you also know how unreasonable it is most of the time, unless by 'concrete' you allow something as vague as, "troubleshoot production issues".
- jeffbee 6y agoAd-hoc production troubleshooting is a reason to keep, at most, 7 days of logs. Usually you want the most recent minute or hour. Troubleshooting usually does not need collection, aggregation, and indexing because either the problem is isolated to a host or the logs of a single host, pod, or process are representative of what is happening in the rest of the fleet. Even if you want to access all logs, it's still better to leave them where they were produced and push a predicate out to every host; your log-producing fleet has far, far more compute resources than your poor little central database, no matter how big that DB is.
- j88439h84 6y agoMay I ask what kind of production environments you have in mind? Are these large-scale FAANG-style deployments or something else?
- bob1029 6y agoLogging to a database is probably a mistake if you haven't thought about how you'd use that data after you write it to disk. If you have good ideas about how you'd want to use that data, then it is probably a fantastic idea. We log to SQL so that we can instantly obtain a full list of log entries that pertain to a specific user action trace id. We can go from a collection of user actions over a larger business process and then for each action we can pull all of the log entries that were generated. All of this is exposed through a nice web interface with full text search capability over the log messages. Without logging to SQL, we would have a hell of a time building something similar. In order to keep this from exploding out of control, we have a strict 90 day expiration policy on all persisted business state and log entries. Our log table indicies are: - User Action Trace Id (16 bytes) - Timestamp (64 bits) - Message (Variable - FTS) We store all log entries in a single SQLite database (logs.db) contained in each environment. These are queried over HTTP from centralized management tools.
- oneplane 6y agoWe use similar setups but with ElasticSearch instead of an RDBMS and we use cluster connections to query against disparate environments at the same time. They are not allowed to communicate to each other, but a 'placeholder' or 'empty' cluster that is allowed to talk to the others can give you exactly what you want for analysis and auditing. Another big benefit is that environments and other 'boxing' formats are isolated and cannot influence another. Same with lifecycle management, they do their own hot/cold/archive/deletion. We usually end up querying based on session signatures (which is something like user agent + JWT), on plain user ids or on trace ids. This gets you a pretty neat timeline with references for full records if you need to drill down on a specific event. Makes everything super fast and when you need more data or want to read PII you can drill down (depending on your access roles).
- tutfbhuf 6y ago> I think some of the posts here miss a little bit of the context as to why things like this happen in the first place. There were always a bunch of bad decisions and a few good one in hindsight. I think it's completely fair to discuss solutions (including more recent tech like ELK) to the problem even more so if the problem still exists. > Yes, SQL's a poor fit for logs, but it's a better fit then a lot of other things, including not logging at all. No I disagree, not logging 40 TB of data to SQL would have been at least an equally good if not better choice, based on the input we have from the blog post. Just keep the past few weeks of raw traffic logs and be fine.