4 ms·
I also had a serious problem with AWS Managed ES. In my case, some of the heap didn't free on every garbage collect which would eventually result in cluster fa
by arecurrence 5y ago
I also had a serious problem with AWS Managed ES. In my case, some of the heap didn't free on every garbage collect which would eventually result in cluster failure. This was likely a JVM misconfiguration and was most easily observed by viewing a shrinking sawtooth pattern on the memory graphs. This resulted in a multi day marathon of sleeplessness keeping the cluster alive by continually rolling it every few hours (We initially assumed we had done something wrong and investigated ourselves first... eventually I shunted traffic to both my own cluster and the managed cluster... my cluster did free heap as expected and we successfully switched over without downtime but wow was it ever hairy during high traffic periods).
We saved a bunch of money and gained performance by using our own cluster. That cluster hasn't gone down since... years later.
It's very difficult to debug these problems when you don't have direct access to elastic search's configuration... what would normally take minutes to verify can take hours to isolate.
- antonhag 5y agoFor what it's worth, this is most likely not a JVM configuration issue and more likely an ES/OpenSearch issue.
- arecurrence 5y agoI'm fairly confident it was because I ended up finding a ticket later where very similar behavior was isolated to a jvm.options configuration problem. Effectively a newer config file was lacking a jvm.options line that changed the behavior with an older machine setup. I would not be surprised if AWS deployed a new config to an old environment. Unfortunately, not having direct machine access, I could not confirm whether this was the case in the failing cluster.