3 ms·
Downsampling might not save as much space as you think depending on how m3db works. https://github.com/thanos-io/thanos/issues/813 https://github.com/thanos-io/
by sciurus 7y ago
Downsampling might not save as much space as you think depending on how m3db works. https://github.com/thanos-io/thanos/issues/813 https://github.com/thanos-io/thanos/issues/813 goes in to why for a similar project.
- roskilli 7y agoWhether it saves space or not, looking at metrics over period of months or years when the data is raw is far more slow/expensive than looking at downsampled data. If you still want to be able to quickly graph and view old data, downsampling is the only way to keep your queries interactive. Take for instance 30s data vs 10min data. 20x more computation, network exchange and everything else of that nature needs to happen. Also if you want to keep only a subset of your data for a very long time, you need to have retention policies - otherwise you end up storing all that extra data forever. At large numbers (terabytes to petabytes) this stuff is impactful, at smaller numbers (gigabytes) my points here are far less relevant.
- cube2222 7y agoI know about the Thanos problem, and it's one of the things I didn't really like about it using the prometheus storage format. It does work in m3db. Following are the on-disk sizes of one replica: 2 weeks at 15s res: 90G 2 months at 1m res: 160G 3 months at 5m res: 60G EDIT: The only thing is that m3db doesn't really downsample. You just create one namespace (table) for each resolution, and set the m3coordinator up so that it writes to each at the wanted interval, then set up a different retention for it. (this way you have duplicates in recent data) M3 namespaces aren't set up for a specific resolution. The writer decides and you can write various series in different resolutions to one namespace in theory.