3 ms·
begging people to recognize that a person who sells a solution is going to view these problems through the lens of being rewarded for applying their solution to
by zemo 3y ago
begging people to recognize that a person who sells a solution is going to view these problems through the lens of being rewarded for applying their solution to your problem, even if it's not appropriate.
> Yet, per my own experience it’s still extremely hard to explain what does Charity meant by “logs are thrash”, let alone the fact that logs and traces are essentially the same things. Why is everyone so confused?
Charity is not confused, Charity is incentivized. What she means by "logs are trash" is "I do not sell a logging product". (and, to be clear, I'm only naming Charity individually here because that's who the author named in their article.)
> When I was working at Meta, I wasn’t aware that I was privileged to be using the best observability system ever.
The observability system that is appropriate for Meta is not necessarily appropriate for your project. Those tools are cool but also require a pretty serious investment to build and operate correctly. It's very easy to wade into a cardinality explosion problem when tagging and indexing everything you can imagine, it's very easy to wade into problems regarding mixed retention policies when some events are important and others are less-important, it's very easy to wade into a latency-sensitivity issue if you're building a log/event collection infra that you don't allow to ever lose data, etc. As it turns out, observability is a large topic.
The idea that there's one "best" way to do observability is a little ridiculous. Like when I worked at Etsy some of the data was literally money, when I worked at Jackbox Games we made fart joke games (Quiplash, Drawful, Fibbage, You Don't Know Jack, etc) and the infrastructure was nothing but pure cost. The observability needs of those two orgs were phenomenally different, because the products were different, the revenue models were different, the needs of the users were different, etc.
Also this notion that "all you need is wide events" is the answer seems ... really shallow. A data point is an unordered set of key-value pairs? That's how ... a LOT of logging, metrics, and tracing infra expresses things at the level of an individual record/event. The difference is in the relationships between the keys and values, the relationships between the individual records, etc.
and "stop sampling" is just a bizarre marketing angle. If you have 1 million records or 10 million records and you get the same squiggly line out of analyzing it, congrats you have inflated the size of the data that nobody ever looks at. There is only one person who this benefits and it's the person who charges you for the pipeline, which is exactly why people who sell a pipeline are incentivized to tell you that sampling is bad: if you are sampling, you are sending and storing and querying fewer data points, so they are charging you less money. They are getting paid to tell you that sampling is bad. Sampling is not good or bad, sampling is sampling. The reality is that in a lot of these systems, the vast majority of the information will never, ever be looked at or used. Whether or not that matters is entirely context dependent.
- isburmistrov 3y ago> and "stop sampling" is just a bizarre marketing angle Wait, where did I mention stopping sampling? :) The opposite: the article is praising the native sampling Scuba has.
- phillipcarter 3y agoFor this: > There is only one person who this benefits and it's the person who charges you for the pipeline, which is exactly why people who sell a pipeline are incentivized to tell you that sampling is bad I largely agree, and I'll say that at least with Honeycomb (since it's mentioned by the author) we make sampling a key component of pretty much any deal before anything gets signed. For small stuff this clearly doesn't matter so much, but it basically boils down to: - Most of your data is probably uninteresting because it's uniform and represents success cases, so you just need a statistically significant sampling to get a sense of what "okay" means for comparisons - Whenever there's an error or high latency, there's almost always something interesting in there, and you probably want all of it And so this typically works out to generating an order of magnitude or two more data than you actually need to get an accurate view of what's going on any any point in time. And so when you do this, you can (and probably should?) pack your events/logs/traces/whatever-you-call-it with a bunch of data. There's some examples where you can't do this, though. Some people want to be able to do something like plug in a customer ID and dig up the exact trace that represents something they complained about. Or there's some compliance to adhere to, legal or non-legal, where it's still less money to pay for everything unsampled than it is to deal with the consequences of non-compliance. But for most organizations I'd say what I mentioned above holds true. ...but that's just one of the several ways that some folks will frame up Observability. The term has been kleenexed a bit so it now means whatever any vendor says it means, and they all say varyingly different things.