3 ms·
Yeah, it takes a bit of tuning, depending on what you want to do with Elasticsearch. There are so many different use cases. I found that especially keeping the
by ralphm 13y ago
Yeah, it takes a bit of tuning, depending on what you want to do with Elasticsearch. There are so many different use cases. I found that especially keeping the heap size for field data down to 40% helped quite a bit, in my case.
- justinsb 13y agoI'd love to see a blog post that went into detail about those issues: what to measure, what to tune etc.
- ralphm 13y agoWhat I found most important is to monitor Elasticsearch while doing that tuning. That's when I set up Graphite and StatsD and https://github.com/ralphm/vor https://github.com/ralphm/vor. First of, you need to make sure Elasticsearch can lock a chunk of memory (using mlock). About half of the available RAM is a good size, as other system processes need some memory, too, and not everything is on the heap. You want to look for how many items you are indexing and how much of the heap the field data cache is using while doing queries. By default ES tries to keep the total heap size at about 3/4 of the allocated memory. The types of queries are important, too. E.g. if you do faceting or sorting on fields that have many different values, this will fill up the field data cache in no time. Is that the kind of information you're after? I can go into in more detail if you have more specific question.
- justinsb 13y agoThat's great stuff - thanks. Just trying to collect tips from those that have been there, so that when I get there I have something to start from!
- ralphm 13y agoYou're most welcome. Drop me a line any time.
- dc2447 13y agoI you want to monitor elastic search with Graphite then Collectd provides an excellent curl-json plugin which works really well with the ES health api https://gist.github.com/dc2447/6783658 https://gist.github.com/dc2447/6783658
- brasetvik 13y agoI wrote an article that covers some of this: http://www.found.no/foundation/elasticsearch-in-production/ http://www.found.no/foundation/elasticsearch-in-production/ It's a bit introductory on the memory-parts. One that goes in more detail on memory- and JVM-tuning is planned. (Full disclosure: I work for Found)
- rafekett 13y agoI worked on a similar system at somewhat larger scale (maybe 5-10x larger volume and dataset size). There's a lot of basic JVM tuning to be done. You have to monitor heap usage and GC pauses and keep manipulating heap size, perm gen size, heap usage that should initiate major GC (forget what variable this is) until you get things as smooth as possible and meet whatever resource constraints you want to meet on your hardware. In my experience, for large datasets ES really pushes the memory architecture of the JVM -- it doesn't really perform too well when your heap size is like 30g. After that, you'll have to tune merging. The underlying Lucene storage engine chokes when it tries to merge segments that are large, so you have to tune the max segment sizes, etc. Then, you'll have to tune your queries -- it's not feasible to do a full index scan on a large index so you have to get clever about how you pull data in chunks. If your data is timestamped and that timestamp is indexed, you can pull data for smaller time ranges which will be faster than pulling all data for an index at once.
- ralphm 13y agoI have understood that you should definitely try to keep the heap size under 32G as apparently the JVM can only compress pointers below that. Because we do event logging and our documents don't change after being indexed, some of the problems you mentioned (like merging) are not as pressing. And of course each log event has a timestamp and we can indeed limit our queries with them.