3 ms·
I am no expert but my take was that in many cases, logs, metrics, and traces are stored in individually queryable tables/stores. Their proposal is to have a si
by cmgriffing 2y ago
I am no expert but my take was that in many cases, logs, metrics, and traces are stored in individually queryable tables/stores.
Their proposal is to have a single "Event" table/store with fields for spans, traces, and metrics that can be sparsely populated similar to a dynamodb row.
Again, I might have missed the point, though.
- strken 2y agoMy take on the underlying point is that there's a good reason not to use events for every system everywhere. That reason is cost: not cost as in "we can save a bit of money", but cost as in "collecting all our events would double our infrastructure budget". Traces (correlated bundles of events) and metrics (aggregated events, and also other stuff that has nothing to do with events, but let's ignore that) are an attempt to solve the problems caused when you've exhausted the limits of shoving events into a table. I think GPs point is that sure, you can shove events into a table, but this is just structured logging. It doesn't need to be called observability 2.0 because it's really observability 0.5: the baseline you start with before you need traces or metrics derived from events. All the observability 2.0 hype and the love hearts are a sales pitch from a CTO explaining why you shouldn't shoot yourself in the foot with an unnecessarily heavyweight observability implementation when your log volume is small.
- Veserv 2y agoYep, my post was just saying this is observability 0.5, not 2.0. However, I actually do agree that you should prefer observability 0.5 in most cases if you have a appropriate implementation. I am familiar with time travel debugging where you use automatic event instrumentation to stream GB/s per core to fully and perfectly reconstruct all program states. The problems and solutions I see being proposed in the distributed tracing space seem so grotesquely inefficient, yet actionable data anemic in comparison that I am pretty sure most of the limits preventing “just events” are due to inadequate implementations. However, if you really are running into limits, then it is important to have these known fallback mechanisms instead of just ignoring their existence and calling it a 2.0.
- spimmy 2y agoit's weird that you think it's a sales pitch, when i ended it by pleading for other people to share writeups of their similar solutions. i know they exist, and i know people who are desperate to use them. if it came across as a sales pitch, i def missed the target somehow, apologies.
- strken 2y agoApologies if that sounded too cynical. I didn't mean that the general argument sounded "sales-ey", just that a CTO writing about how their company's implementation is better is not taking a thousand-foot objective view of the upsides and downsides. In this case, I do think 1.0 and 2.0 are flawed names that imply one supercedes the other, when really it's a choice between scalability to levels most companies don't need at the cost of complexity most companies don't want.