8 ms·
tbh that's not the flex. storing 100PB of logs just means we haven't figured out what's actually worth logging. metrics + structured events can usually tell 90%
by b0a04gl 1y ago
tbh that's not the flex. storing 100PB of logs just means we haven't figured out what's actually worth logging. metrics + structured events can usually tell 90% of the story. the rest? trace level chaos no one reads unless prod's on fire. what'd could've done better be: auto pruning logs that no alert ever looked at. or logs that never hit a search query in 3 months. call it attention weighted retention. until then this is just high end digital landfill with compression
- Spivak 1y ago> trace level chaos no one reads unless prod's on fire God why do we keep these fire extinguishers around, they sit unused 99.999% of the time.
- jiggawatts 1y ago“Just go back in time and turn on the specific log you will need!”
- hinkley 1y agoThat logging isn’t even free on the sending side, especially in languages where they are eager to get the logs to disk in case the final message reveals why the program crashed. And there’s a lot of scanning blindness out there. Too much extraneous data can hide correlations between other logs entries. And there’s half life in value of logs written for bugs that are already closed, and it’s fairly short. I prefer stats because of the way they get aggregated. Though for GIL languages some models like OTEL have higher overhead than they should.
- nijave 1y agoIn fairness, I think a lot of GIL languages already have high overload and I've never been under the impression OTEL was optimized for performance and efficiency.
- hinkley 1y agoIt really isn’t. The code reads like it was designed by SpringBoot users. You have to read three different docs to suss out how to use multiple calls together to get a desired approach, and some of the docs leave out critical details. I think people forget that folks use Google thinks is the top result isn’t necessarily what the creators would assume is the document people would find for a topic. I’ve been trying to explain this to the Elixir community for instance. “Can’t to X, doesn’t work.” “Look, it’s easy. Did you even RTFM? http://blah.example.com/doc/articleb#section2” http://blah.example.com/doc/articleb#section2” “Uh, no, because search engine took me to http://blah.example.com/doc/articleg#section7” http://blah.example.com/doc/articleg#section7”
- eddd-ddde 1y agoIs there any tools that does log/trace capture on error conditions? I.e. we capture all local events, but only upload them when something meaningful happens, like the server crashed / requests are returning 5xx.
- mdaniel 1y agoI love this idea in principle, but in practice I would guess it means one of two sub-optimal things: either the node caches them for a window of time, in order to know whether to really transmit them, or the logs are mutated post-delivery as kind of a "tiny expiry" Everything else I could write is just turning various trade-off knobs, which is why I'd guess you haven't seen an out-of-the-box offering that does what you're describing. There's not just one solution to it that would be reasonable for all audiences
- Macha 1y agoI've been in a bunch of companies that have pushed for reducing logs in favour of metrics and a limited set of events, usually motivated by "we're using datadog and it's contract renewal time and the number is staggering". The problem is, if you knew what was going to go wrong, you'd have fixed it already. So when there's a report that something did not operate correctly and you want to find out WTF happened, the detailed logs are useful, but you don't know which logs are useful for that unless you have reoccuring problems.
- __MatrixMan__ 1y ago> auto pruning logs that no alert ever looked at I'm sure someone somewhere is working on an AI that predicts whether a given log is likely to get looked at based on previous logs that did get looked at. You could store everything for 24h, slightly less for 7d, pruning more aggressively as the data gets stale so that 1y out the story is pretty thin--just the catastrophes.
- ethan_smith 1y agoThe "attention weighted retention" concept is brilliant. You could implement this with a simple counter tracking query/alert hits per log pattern, then use that for TTL policies in most observability platforms. This approach reduced our storage costs by 70% while preserving all actionable data.
- solatic 1y agoIf you work for a large enterprise, there are so many dev teams supporting so many products that "we haven't figured out what's actually worth logging" is just disconnected from the developer incentives in those teams (ship features fast, fix your problems even faster because nobody has time for that BS) as well as ops incentives (the servers ARE on fire, and the devs didn't log enough). FinOps comes last, if there's even cost tracking per team in the observability suite. You don't understand why DataDog has a $44 billion market cap. It's yet another instance of Finance complaining that the transition to The Cloud gave every engineer a corporate credit card with no spend controls or a way for Finance to turn off the spigot.
- CoolCold 1y agowut? > As you’ll read below, this saves us millions of dollars a year and allows us to scale out our ClickHouse Cloud service without having to be concerned about observability costs, or make compromises on the log data we retain. https://clickhouse.com/blog/building-a-logging-platform-with-clickhouse-and-saving-millions-over-datadog https://clickhouse.com/blog/building-a-logging-platform-with...
- behemot 1y agohey there! I work at ClickHouse. to clarify: the vast majority of this 100PB is structured events. in our case logs are supplementary.
- imiric 1y agoSure, but if the data is already there, it's a sifting and pruning problem, which can be done after ingestion, if needed. It's better to have all data and not need it, than to need it and not have it. Assuming you have the resources to ingest it in the first place, which seems like the focus of the optimization work they did.
- namanyayg 1y ago[flagged]
- hnlmorg 1y agoI’m of the opposite opinion. It’s better to ingest everything and then filter out the stuff you don’t want at the observability platform. The problem of filtering out debug logs is you don’t need them, until you do. And then trying to recreate an event you can’t even debug is often impossible. So it’s easier to then retrieve those debug logs if they’re already there but hidden.
- gavinray 1y ago"Better to have it and not need it; than to need it, and not have it..."
- jkogara 1y agoOr more succinctly, albeit less eloquently: "Better to be looking at it than looking for it."
- 9dev 1y agoUntil you’re working with personal information of EU customers, where the opposite maxime applies: "Only store what you absolutely need" Seriously, storing petabytes of logs is a guarantee for someone on your team writing sensitive data to logs, and/or violate regulations.
- jodrellblank 1y ago“You can’t have everything. Where would you put it?” - Steven Wright. “Better to have hoarding disorder than to need a fifty year old carrier bag full of rotting bus tickets and not have one” really should need more justification than a quote about how convenient it is to have what you need. The reason caches exist as a thing is so you can have what you probably need handy because you can’t have everything handy and have to choose. The amount of things you might possibly want or need one day - including unforeseen needs - is unbounded, and refusing to make a decision is not good engineering, it’s a cop-out. Apart from cost, the more time and money you spend indexing, cataloging, searching it. How many companies are going to run an internal Google-2002 sized infrastructure just to search their old hoarded data?
- nikolayasdf123 1y agoyeah, same thoughts. business events + error/tail-sampled traces + metrics ... and logs in rare cases when none of the above works. logs are dump of everyting. why would you want to have so many logs in first place? and then build whole infra to scale that? and who and how reads all those logs? they build metrics on top of that? so might as well just build metrics directly and purposefully? with such high volume, even LLMs would not read them (too slow and too costly).. and what would even LLM tell from those logs? (may be sparce/low signal, hard to decipher without tool-calling, like creating merics)