4 ms·
JSON is a good choice if you have very tight coupling between producers and consumers of your logs. Twitter (disclosure: I manage/was early engineer on the ana
by squarecog 14y ago
JSON is a good choice if you have very tight coupling between producers and consumers of your logs.
Twitter (disclosure: I manage/was early engineer on the analytics infra team) used Json initially, and it quickly turned into a mess because JSON does not enforce schemas or type safety. There is no way for your consumers to find out what you are logging, other than to sample records. This makes evolving logging, deprecating fields, and doing sanity checking very complex (you basically have to build a separate metadata discovery system).
Use Thrift or Avro. Avro is extra nice because not only does it have a schema, it keeps the schema together with the data; however, the support for it is not as mature/wide. It's improving fast, though. Thrift, and, I think, Avro, have JSON protocols if you really love JSON -- you get the human-readable debugging mode, the metadata, the type safety, and compact binary representation.. happiness ensues.
- calibraxis 14y agoExcellent point. People should read your coworker Nathan Marz's book _Big Data_, for a lucid discussion of schemas.