4 ms·
The problem with logging since time immemorial seems to be log producers throught about it as 'How do we put records from one SQL table into strings?' IMHO, th
by ethbr0 3y ago
The problem with logging since time immemorial seems to be log producers throught about it as 'How do we put records from one SQL table into strings?'
IMHO, that's not even the right mental model, because it assumes a constant schema. Log entries are more like types: the fields in one entry may look nothing like the next.
The consequence of using the wrong model is its inevitable collapse into 'Shove everything custom we need into a single field that's somewhat appropriate'.
In the base case that devolves into everything in a one line string. Yet even in structured systems you see the behavior in {use of boring, standard fields correctly} + {everything else in an 'error' or 'customdata' field}.
Thankfully, some tools and ecosystems seem to have grokked this and moved forward.
And honestly, it's mostly an ecosystem problem (producers and consumers), so really needs to be solved via more programmatic self-declaration of data structures.
- dale_glass 3y agoYeah, logging doesn't fit well into a RDBMS because fields can change from one message to another. My imagined ideal logging model is logging two things: 1. A structured tree of data. With field names, types, and content in native format. 2. An end-user hint about which of those fields are of most interest to an user. This is used by a log viewer to generate something like a traditional log file by default. So an HTTP log entry might contain the entire HTTP request with all the headers and metadata, and the hint then says "Of this, the admin probably want to see timestamp, URL and status code". You get the convenience of paging visually through what looks like a traditional log, but also everything is available for easy analysis.
- ethbr0 3y agoTo me, GraphQL is a great base model for what logs should be. The producer declares all possible data. The consumer requests the specific data it wants, at runtime. Because fundamentally, there's a mismatch of knowledge. Log consumers won't inherently know everything that can be in logs. Ergo, modern API-style ecosystem models are probably better approaches.