4 ms·
My two cents: monitoring RAM usage is completely useless, as whatever number you consider an “used/free RAM” is meaningless (and the ideal state is that all of
by dfox 2y ago
My two cents: monitoring RAM usage is completely useless, as whatever number you consider an “used/free RAM” is meaningless (and the ideal state is that all of the RAM is somehow “used” anyway). You should monitor for page faults and cache misses in block device reads.
- bayindirh 2y agoDepends. "free" reports the area used for disk buffers and programs, hence "available" and "free" numbers. On my servers I want some available RAM which means "used - buffers", because this means I configured my servers correctly and nothing is running away, or nothing is using more than it should. On the other hand, you want "free" almost zero on a warmed up server (except some cases which hints that heaps of memory has been recently freed) since the rest is always utilized as disk cache. Similarly having some data on swap state doesn't harm as long as it's spilled there because some process has ran away and used more memory than it should be. So, RAM usage metrics carry a ton of nuance and can mean totally different things depending on how you use that particular server.
- hinkley 2y agoOne of the older arguments I get to keep having over and over is No, You May Not Put Another Service on These Servers. We are using those disk caches thank you very much. I do not enjoy showing up to yet another discussion of why our response times just went up “for no reason”. Learn your latency tables people.
- bayindirh 2y agoYeah, people tend to think server utilization as black and white. Look, we're using just 50% of that RAM. Look, there're two cores that are almost idle. No & No. Rest of the RAM is your secret for instant responses, and that spare CPU resource is for me to do system management without you notice or to front the odd torrent of requests we have semi regularly (e.g.: /. hug of death. Remember?).
- hinkley 2y agoI need to find a really good intro to queuing theory to send people to. A full queue is a slow queue. You actually want to aim for about 65% utilization.
- bayindirh 2y agoAlso, there was a formula for determining the optimal cache size. I forget the name all the time. IIRC, in the end, caching most popular 10 items was enough to respond to 95% of your queries without hitting the disk.
- redxtech 2y agoIf the numbers from the phoenix project are to be trusted, a loose estimate is the time spent in queue is proportional to the ratio of utilized to unutilized resources. For example, 50% used & 50% unused is 50:50 = 1 unit of time. 99% used is 99:1 = 99 units of time.
- Version467 2y agoThis might be too basic, but I found this blog post to be an incredible introduction to queues: https://encore.dev/blog/queueing https://encore.dev/blog/queueing
- hi-v-rocknroll 2y agoCorrect identification but wrong prescription. Cache misses don't have anything to do with memory pressure, they're related to caching effectiveness. Production systems shouldn't have any page faults because they shouldn't be using swap. The traditional way Linux memory pressure was measured using a very small swap file and check for any usage of it. Modern Linux has the PSI subsystem. Also, monitoring for OOM events also means a system needs more RAM or a workload needs to be tuned or spread out.
- dilyevsky 2y agoPretty strong statement without any supporting arguments. How do you propose someone debug a memory leak issue using page faults and cache misses but no rss metrics?
- fulafel 2y agoYou should monitor meaningful memory metrics instead, eg memory usage of your processes, and system memory pressure.