5 ms·
> Because Elasticsearch is not a reliable data store, organizations that use Postgres typically extract, transform, and load (ETL) data from Postgres to Elastic
by radpanda 2y ago
> Because Elasticsearch is not a reliable data store, organizations that use Postgres typically extract, transform, and load (ETL) data from Postgres to Elasticsearch
I’ll admit haven’t kept up with this but is it still the case that Elasticsearch is “not a reliable data store”?
I remember there used to be a line in the Elasticsearch docs saying that Elasticseach shouldn’t be your primary data store or something to that effect. At some point they removed that verbiage, seemingly indicating more confidence in their reliability but I still hear people sticking with the previous guidance.
- wordofx 2y agoES is just unreliable. Can be running smoothly for a year and boom it falls over and you’re left scratching your head.
- cyberes 2y agoThat exact scenario just happened to me a few days ago.
- fizx 2y agoA problem here is that ES (and Solr too) are pathological with respect to garbage collection. To make generational GC efficient, you want to have very short lived objects, or objects that live forever. Lots of moderately long-lived objects is the worst case scenario, as it causes permanent fragmentation of the old GC generation. Lucene generally churns through a lot of strings while indexing, but it also builds a lot of caches that live on-heap for a few minutes. Because the minor GCs come fast and furious due to indexing, that means you have caches that last just long enough to get evicted into the old generation, only to become useless shortly thereafter. The end result looks like a slow burning memory leak. I've seen the worst cases take down servers every hour or two, but this can accumulate over time on a slower fuse as well.
- nickpsecurity 2y agoCan this be fixed with alternative GC’s or tuning?
- fizx 2y agoIt might be better by now with the newer GC options (ZGC, G1, Azul). For a while those had their own problems, but I'm a little out of the loop. Tuning the older options wasn't really all that beneficial. We tried tuning, then running a custom build of OpenJDK to mess with the survivor space (which isn't that tunable via config), then ultimately settled on more aggressive (i.e. weekly) rolling restarts of servers.
- mannyv 2y agoThis seems to be a Lucene/SOLR problem. I used Lucene/SOLR years ago, and it died so often that we just auto-indexed nightly and had a re-index button in the UI.
- jillesvangurp 2y agoI've been using and supporting ES (and OS lately) for well over a decade. Mostly its fine but a lot of users struggle with sizing their clusters properly (which is costly). Elasticsearch falling over is what happens when you don't do that. It scales fine until it doesn't and then you hit a brick wall and things get ugly. Additionally, many companies learn the hard way that dynamic mapping is a bad idea because you might end up with hundreds of fields and a lot of memory overhead and garbage collection. I've fixed more than a few situations like this for clients that ended up with hundreds or thousands of fields, many shards and indices. Usually it's because they are just dumping a lot of data in there without thinking about how to optimize that for what they need. A properly architected setup is not going to fall over randomly. But you need to know what you are doing and there are a lot of clients that I help that clearly don't have the in house expertise to do this properly and are a bit out of their depth.
- fizx 2y agoIt got a lot better in the ~7 series IIRC when they added checksums to the on-disk files. I don't know if you still have to recover corruptions by hand, or whether the correct file gets copied in from a replica. The replication protocols and leader election were IMO not battle-hardened or likely to pass Aphyr-style testing. It was pretty easy to get into a state where the source of truth was unclear. Source: Ran an Elasticsearch hosting company in the 2010's. A little out of the loop, but not sure much has changed.
- willio58 2y agoI'm not sure what the author was referring to, but in our stack ES is the only non-serverless tech we have to work with. I know there's a lot of hate in HN around serverless for many reasons but for us for several years, we've been able to scale without any worry of our systems being affected performance-wise (I know this won't last forever). ES is not this way, we have to manage our nodes ourselves and figure out "that one node is failing, why?" type questions. I hear they're working on a serverless version, but honestly, I think we will be leaving ES before that happens.
- maxxxxxx 2y agoElastic’s serverless offering went GA recently: https://www.elastic.co/elasticsearch/serverless https://www.elastic.co/elasticsearch/serverless
- farsa 2y agoIt's still in technical preview.
- jrochkind1 2y agoWhat are you considering replacing it with for full-text search?
- easton 2y ago> that one node is failing, why We have a similar experience except we’re using AWS’ version, so any inquiry as to what happened ends with support saying “maybe if you upgrade your nodes this won’t happen again? idk”
- alecthomas 2y agoAnd that's why we pay AWS the big bucks.
- amai 2y agoJust read the article: „Elasticsearch’s lack of ACID transactions and MVCC can lead to data inconsistencies and loss, while its lack of relational properties and real-time consistency makes many database queries challenging.“
- IamLoading 2y agoThey publicly track their resillency efforts here https://www.elastic.co/guide/en/elasticsearch/resiliency/master/index.html https://www.elastic.co/guide/en/elasticsearch/resiliency/mas... >7 its become quite reliable. It comes down to how you manage and maintain the cluster
- jillesvangurp 2y agoI've been abusing it as a data store for many years. I don't recommend it but not because of a lack of reliability. It's actually fine but you need to know what you are doing and it's not exactly an optimal solution from a performance point of view. The main issue is not lack of robustness but the fact that there are no transactions and it doesn't scale very well on the write side unless you use bulk inserts. You can work around some of the limitations with things like optimistic locking to guarantee that you aren't overwriting your own writes. Doing that requires a bit of boiler plate. For applications with low amounts of writes, it's actually not horrible if you do that. Otherwise, if you make sure you do regular backups (with snapshots), you should be fine. Additionally, you need to size your cluster properly. Things get bad when you run low on memory or disk. But that's true for any database. If you are interested in doing this, my open source kotlin client for Elasticsearch and Opensearch (I support both) has an IndexRepository class which makes all this very easy to use and encapsulates a lot of the trickery and bookkeeping you need to do to get optimistic locking. I also have a very easy to use way to do bulk indexing that takes care of all the bookkeeping, retries, etc. without requiring a lot of boiler plate. And of course because this is Kotlin, it comes with nice to use kotlin DSLs. You can emulate what it does with other clients for other languages but I'm not really aware of a lot of other projects that make this as easy. Frankly, most clients are a bit bare bones and require a lot of boiler plate. Getting rid of boiler plate was the key motivation for me to create this library. https://github.com/jillesvangurp/kt-search https://github.com/jillesvangurp/kt-search