4 ms·
Early in Prometheus' life (2011-12), I had really pined for an indexed sparse columnar data store similar to the semantics of what Bigtable provided Google inte
by matttproud 5y ago
Early in Prometheus' life (2011-12), I had really pined for an indexed sparse columnar data store similar to the semantics of what Bigtable provided Google internally for time series storage for many years.
Pre-public-Git repository attempts at Prometheus' data store had used Cassandra in prototyping, but the Thrift protocol, which Cassandra then used, was unfortunately riddled with bugs in the wireformat (how were enums indexed over the wire? turns out it varied by what kind of client you were using even in the same distribution release, which was to me incredulously bad), which resulted in a lot of byzantine bugs. I eventually gave up on Cassandra and settled on a different data model using LevelDB (Levigo bindings). It worked well and was stable and allowed me to not need to worry about:
https://danluu.com/file-consistency/ https://danluu.com/file-consistency/
and
https://danluu.com/deconstruct-files/ https://danluu.com/deconstruct-files/
I had a feeling in the back of my mind once Prometheus matured a bit it could have fallen back onto Cassandra for distributed series storage (again, ca. 2012 thinking), since it largely could have used an append-only storage model for series archival. Even the metrics metadata (metric families, fields, and indexing) could have been stored in such a way.
It would be fun to revisit these design problems with today's technologies and knowledge.