Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
isburmistrov
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
isburmistrov
3y ago
I agree that specialised DBs outperform a general-purpose OLAP database. The question is - what does outperform mean. In this area queries should not be actually ultra-fast, they should be reasonably fast to be comfortable. And so missing i
2.
▲
by
isburmistrov
3y ago
Thanks for sharing! > This change together with Apache Superset for the BI layer made a huge impact in the amount of internal users that could extract value from the data collected. Went from around 150 to 800 internal users. Really impr
3.
▲
by
isburmistrov
3y ago
And yet people use ClickHouse quite effectively for this very problem, see the comment here: https://news.ycombinator.com/item?id=39549218 There are also time-series databases out there that are OK with high cardinality: h
4.
▲
by
isburmistrov
3y ago
Thanks for sharing, didn't know about Motif
5.
▲
by
isburmistrov
3y ago
Does it support "native sampling" described in the article? This is really important to keep the cost low.
6.
▲
by
isburmistrov
3y ago
Never in the world I would have expected my post to cause the discussion about ODS flaws :D
7.
▲
by
isburmistrov
3y ago
I think the storage architecture in ClickHouse and Elastic are very different. And I think compression in ClickHouse can be really damn good. Don't have a good comparison at hands though, but a few random links on the topic: https:&#x
8.
▲
by
isburmistrov
3y ago
Haha, I think this term is originated by Honeycomb team actually. Why I prefer "wide event" over "structured log" as a term because it has this "wide" component that serves for 2 purposes: - it highlights the i
9.
▲
by
isburmistrov
3y ago
Yes, structured log exactly. Why I prefer "wide event" as a term because it has this "wide" component that serves for 2 purposes: - it highlights the intention of storing as much context as possible - it also hints on th
10.
▲
by
isburmistrov
3y ago
> Those are entirely different orders of magnitude both when it comes to size and how much usefulness you lose. In modern storage backends like Victoriametrics a counter gonna cost you around byte per metric per probe. And as you emit th
11.
▲
by
isburmistrov
3y ago
I don't think Meta's margins have something to do with this. Companies smaller in scale than Meta also have less data! And yes, Scuba is in-memory, but it's not the requirement. Check this video out on how Honeycomb implement
12.
▲
by
isburmistrov
3y ago
I think ClickHouse is becoming a default storage for observability nowdays: https://clickhouse.com/use-cases/logging-and-metrics And there are quite a few solutions on top of it. A couple of examples that seem to be in
13.
▲
by
isburmistrov
3y ago
Columnar storage stores data very efficiently, too - because it compresses data of a similar nature (columns). Check e.g. ClickHouse on this matter: https://clickhouse.com/docs/en/about-us/distinctive-feature
14.
▲
by
isburmistrov
3y ago
Sampling can be smart, e.g. based on some field all events have (can be called traceId, haha).
15.
▲
by
isburmistrov
3y ago
> and "stop sampling" is just a bizarre marketing angle Wait, where did I mention stopping sampling? :) The opposite: the article is praising the native sampling Scuba has.
16.
▲
by
isburmistrov
3y ago
I didn't say such tools don't exist. Honeycomb, mentioned in the post, is exactly Scuba fwiw. I said that over-focusing on traces / metrics and "logs" (in the classical understanding) hides the true power of wide ev