4 ms·
This academic paper has some interesting insights about what distributed systems (e.g. Hadoop) log in order to be able to reconstruct what happened after a prob
by serhei 9y ago
This academic paper has some interesting insights about what distributed systems (e.g. Hadoop) log in order to be able to reconstruct what happened after a problem occurs.
http://www.eecg.toronto.edu/~yuan/papers/zhao_stitch.pdf http://www.eecg.toronto.edu/~yuan/papers/zhao_stitch.pdf
In particular,
"- log a sufficient number of events — even at default
logging verbosity — at critical points in the control path so as to enable a post mortem understanding of
the control flow leading up to the failure.
- identify the objects involved in the event to help differ-
entiate between log statements of concurrent/parallel
homogeneous control flows. Note that this would not
be possible when solely using constant strings. For
example, if two concurrent processes, when opening
a file, both output “opening file”, without additional
identifiers (e.g., process identifier) then one would not
be able to attribute this type of event to either process.
- include a sufficient number of object identifiers in the
same log statement to unambiguously identify the ob-
jects involved. Note that many identifiers are naturally
ambiguous and need to be put into context in order to
uniquely identify an object. For example a thread iden-
tifier (tid) needs to be interpreted in the context of a
specific process, and a process identifier (pid) needs to
be interpreted in the context of a specific host; hence
the programmer will not typically output a tid alone,
but always together with a pid and a hostname. If the
identifiers are printed separately in multiple log state-
ments (e.g., hostname and pid in one log statement and
tid in a subsequent one) then a programmer can no
longer reliably determine the context of each tid be-
cause a multi-threaded system can interleave multiple
instances of these log entries."
Some of this may be duh-obvious, but it results in a naming scheme where all important objects in the system have unique, printable ids.
Couple more observations: Hadoop logs tend to fit on one line (they can be long lines, though), except in the case of exceptions. Exceptions are almost always logged, even if they represent an ordinary situation that is recovered from in the error handling code. (Many software failures in these systems occur because a standard, recoverable exception is thrown in one place, improperly handled in the error handling code, and then this causes a much more severe exception further down the line.) All exception logs include a multi-line stack trace.