4 ms·
What sort of work are you referring to when you say it's a lot to keep up? We've been using OpenTSDB here in an extremely high traffic setup for about a year n
by Cidan 11y ago
What sort of work are you referring to when you say it's a lot to keep up?
We've been using OpenTSDB here in an extremely high traffic setup for about a year now, with absolutely no issues at all. It took a few hours at most to setup and figure out scaling.
- krenoten 11y agoI love OpenTSDB, but I think it's highly unusual for people to have "absulutely no issues at all". I ran OpenTSDB at the hundreds of thousands to low millions of ops per second range for several years. OpenTSDB has historically crapped itself when you have writes hit it for rows that it already ran its all-columns-for-an-hour-into-a-single-column compaction on. Now it just drops data on reads. If you have a large team of engineers writing data into it, they will sometimes abuse the schema by overloading single metrics with many tags. Every permutation of tag*value creates a new row per hour of data. This creates extremely hot shards when engineers inevitably do something like store a client IP address into a tag. When this happens, you may have to write some web-UI scraping code to figure out which regions are getting the most traffic (assuming you have "extremely high traffic setup" numbers of region servers), or script up something like misra-gries on a tcpdump of its inbound metrics to see where the hot shit is so you can get somebody's deploy reverted. I know of some companies that have forked OpenTSDB and prefixed each row with a hash such that it spreads overloaded metrics around the cluster much more evenly. KairosDB solves this by not using a lexicographically sharded database (Cassandra). When run properly, it runs great. But that takes some really painful learning experiences to learn how to do, generally.