3 ms·
Hey @precisionemma, sorry for the late reply - that's a great question. So I'd say there's a few non-standard things that help: - Transport compressed blocks
by roskilli 8y ago
Hey @precisionemma, sorry for the late reply - that's a great question. So I'd say there's a few non-standard things that help:
- Transport compressed blocks over RPC storage nodes to query service (i.e. raw byte buffers instead of timestamp and float64 values) to reduce response payloads (and less serialization/deserialization steps)
- Auto batching of fetches, each batch can be fetched in parallel (i.e. on a 24 core box, a single query can become 24 subqueries and run concurrently)
- Series block level caching (fine grained cache), index caching, can enable async inserts which heavily reduces lock contention at the series map layer because the inserts get batched together (durable as written to WAL but series may take a few milliseconds to show up after a write finishes, so you can't always read your own writes if you're using this, which is ok for metrics)
- Ensuring bloom filters and index summaries are in mmap and not on heap, to reduce GC pressure, also scoped to each time window so helps for sparse time series because the disk has to be read a lot less if the in memory bloom filters tell you you only need to fetch part of one file volume (out of a lot of potential file volumes) before going to disk
- Object pooling of a lot of data structures and keeping things in mmaps as much as possible helps reduce the heap size and consequently GC pressure is less (even using bytes instead of strings for keys and IDs was a huge win, because byte buffers can be reused whereas strings are immutable and so cannot be reused)
I hope that helps, its not an entirely comprehensive list but definitely covers some things.