12 ms·
Elasticsearch 1.0.0 released
- Zilog 13y agoToo bad they have yet to address the split brain issue.
- r00fus 13y agoLink for the curious: http://blog.trifork.com/2013/10/24/how-to-avoid-the-split-brain-problem-in-elasticsearch/ http://blog.trifork.com/2013/10/24/how-to-avoid-the-split-br...
- AznHisoka 13y agoTrue, that's a valid issue. For me, it's not as I end up indexing the same document multiple times over the course of 2-3 days.
- chriscareycode 13y agoI haven't had a split brain on my 15 node cluster in over 6 months even though the cluster is split among multiple data centers which do drop connectivity from time to time. When the setting was wrong, it happened constantly. Tune it properly and it won't happen. n/2+1
- kailuowang 13y agoCongratulations to the team. This is a great library that we really appreciate.
- dabeeeenster 13y agoES is a fantastic project. Thank you thank you thank you for your contribution; truly standing on the shoulders...
- mtrn 13y agoElasticsearch is a really great piece of software because it makes the simple easy and the complicated possible.
- m0th87 13y agoIt was two weeks ago, and our startup was on the precipice of a major launch. We had completely rewritten our online publication site, which drives the bulk of our traffic. The product had to be shipped on-time - we had press releases, eager investors and a launch party dependent on it. A few days before launch, things were not looking good. As admins manipulated articles in preparation for the launch, the servers kept crashing. In a time-constrained major launch like this, a lot of nasty little hacks build up in the codebase. Our search system for admins was a complete mess. It was a custom solution that worked fine when admins managed a handful of database records, but now that they were managing thousands of articles, it was not scaling at all. At the 11th hour, we dropped elasticsearch into our infrastructure. It worked like a charm. The servers stopped crapping out, and we launched on time. Elasticsearch mostly "just works", and we didn't have to worry about complex schema definitions, working with giant complex XML files (hello Solr), or build anything on top to interface between the index and the queries themselves (Lucene). Thanks elasticsearch, you saved us!
- dc2447 13y ago> Elasticsearch mostly "just works", and we didn't have to worry about complex schema definitions, working with giant complex XML files (hello Solr) If you were using Solr there are a few operational modes to run in. Config file based or SolrCloud[0]. The latter is more akin the ES in terms of cluster management. I agree though from an simplicity of deployment perspective at scale ES is has a much lighter learning curve. [0] https://cwiki.apache.org/confluence/display/solr/SolrCloud https://cwiki.apache.org/confluence/display/solr/SolrCloud
- acdha 13y agoSolrCloud is nothing like ES in terms of management: you end up running a separate zookeeper service with even more files which all have to be configured correctly just to get it running and you have to micromanage shard allocation to ensure that you can add nodes in the future but also not have it intentionally deadlock when a server fails and you no longer have enough nodes for a quorum. All of this happens with the usual contempt for sysadmins where things you need to know (“refusing to process requests”) won't be logged but a bunch of startup boilerplate will be, and simply configuring logging correctly requires (IIRC) editing two XML files and a properties file. `java -jar elasticsearch.jar` does a better job and that's basically all it takes. I'm planning to switch as soon as https://github.com/elasticsearch/elasticsearch/issues/256 https://github.com/elasticsearch/elasticsearch/issues/256 lands.
- skarnik 13y agocongrats to the team!
- RyanZAG 13y agoElasticsearch is really awesome for searching, but what most people don't realize is that it makes a better MongoDB than MongoDB while giving you that searching too.
- sandGorgon 13y agoI had a live production logistics system running on top of Elasticsearch 0.6 (as a NoSQL database ) back in 2012. This powered one of India's largest ecommerce systems (at that time). Elasticsearch is brilliant as a NoSQL - and if you were already using elasticsearch as a search system, you dont need to introduce yet another component into your stack.
- mtrn 13y agoTrue. I evaluated Mongo, Couch and a couple of similar solutions, but ES being a search engine from the start really convinced me, that it can be a viable database for loosely structured data.
- kainosnoema 13y agoI'm surprised so many people miss this. Out of the box, Elasticsearch is a distributed NoSQL store with better write consistency (and arguably performance) than MongoDB offers in its default configuration. The major missing feature was backup snapshots and restores, which 1.0 delivers—along with aggregations that more than rival MongoDBs. The team has intentionally avoided marketing themselves as a NoSQL store (was told this directly by an employee), but they're aware of the potential and have customers using it as such.
- nkoren 13y agoIt's easy to miss. On the front page, the word "store" only occurs once, buried three page-scrolls down in the body text. Otherwise it very much gives the impression of being some kind of analytics dashboard for third-party datastores. And I didn't notice that until after I've visited the website, clicked through a few links trying to figure out what the fuss was about, then gave up and decided to read the comments here.
- lflux 13y ago> Easy to read, console-based insight into what is happening in your cluster. Particularly useful to the sysadmin when the alarm goes off at 3am and JSON is too difficult to read. It's these little details I love, when a project actually cares about operations and not just "well here's the API" I've been using ElasticSearch only for Logstash, but i've been blown away so far as how easy it is to deal with.
- buckbova 13y agoI didn't know what this was and looking at this link it was tough to tell. The github lays it out well. https://github.com/elasticsearch/elasticsearch https://github.com/elasticsearch/elasticsearch
- karterk 13y agoElasticsearch mostly "just works". The latest version of Solr has made clustering easier (requires managing Zookeeper), but before that, it was either ES or nightmare. Lucene is one of those projects which hardly has any real competition. That's surprising given how many real world software projects have a search requirement. While Lucene is excellent, it's not without flaws and competition is always great.
- swah 13y agoHmm, could that be because they have to compete with free?
- malaporte 13y agoLucene does have competition, mostly in the commercial world. I know, since I work for one of those companies :p Solr, ElasticSearch, etc. are mostly concerned about the index/search features, and they do quite a good job there. But this still leaves a huge amount of space for commercial offerings, as core search is only a part of the problem. I'm thinking about connectivity with complex enterprise systems, support for the specific security models of those systems, integration in other systems, etc. Believe me, those problems are not easy to solve. So, even if we have an index that can most probably match Lucene's feature for feature and quite a lot of things beside, we typically won't go after deals where simple search is the only requirement. Instead we focus on larger deals with more complex requirements. And we're doing quite well, thank you :)
- m0th87 13y agoFWIW, Elasticsearch builds on Lucene. It's just working at a much higher level of abstraction.
- dclara 13y agoI agree with you, almost every website needs a search server on the backend for people to search their document base, especially for enterprise intranet. Maybe enterprises are using commercial products, such as SharePoint. How about the rest of the small businesses and websites? Maybe the learning curve is steep for every website to adopt so far.
- NDizzle 13y agoI also took a few days a few weeks ago to setup elastic search after my mysql full text search fell apart. What I'm doing is slamming the full text output of OCRed PDFs into a MyISAM table, the entire document in a text field. What I'm afraid I'm not doing right is creating the web interface to search elasticsearch. What I'm using filters with the query string syntax[1] in the search box, pointing directly at that fulltext column. I'm also using the highlight functionality so that I can specify how many highlight blurbs to return with the result. The query string syntax works great with the OCR'd text, because most of it is near-garbage (as most ocr is) so you can search for something like "net sales"~50 to find those two terms within 50 words of each other. I think the results were something like: net sales 15,000 results "net sales" 120 results "net sales"~50 550 results Can anyone point me at a good web based search implementation using elasticsearch that explains how they're doing it? What I have works pretty good, I just want to... check my work, I guess. [1]: http://www.elasticsearch.org/guide/en/elasticsearch/reference/current/query-dsl-query-string-query.html#query-string-syntax http://www.elasticsearch.org/guide/en/elasticsearch/referenc...
- nzadrozny 13y agoI host and support websolr.com and bonsai.io and have seen a lot of search implementations. The main thing for good stability and performance is to be very good at batching your updates. You don't want to sling a ton of highly-parallel single-document updates at Lucene, lest you thrash the JVM and start garbage collecting like crazy. From there, on the query side, you'll want to get a good working knowledge of the different tokenization and analysis options. There are a lot of subtle and interesting combinations to be had in there that influence performance and relevance of your search results.
- NDizzle 13y agoDo you have a demo on either of those sites where I can input terms into a search box and look at results? What explanation do you give to users as to the options available when formatting the query?
- 13y ago
- gane5h 13y agoReally impressed with the pace of innovation in the last few months: cat api, aggregations, snapshots. The unfortunate side effect is that books and stack overflow posts written before 1.0 are outdated. Disclaimer: I’m the founder of a hosted Search As A Service and we use ES in a few critical parts of our infrastructure.
- bryanh 13y agoThe thing that worried me the most about Elasticsearch was how fragile it got around the limits of its performance. Run out of memory because of a nasty query? Boom, data corrupted. I hope you weren't using it as your primary persistence layer... Otherwise, we love ES. The other comment about it being a better Mongo than Mongo rings true. With the backup/restore API and the some of the circuit breakers, I'm hopeful that my fears will be abated.
- nzadrozny 13y agoDitto open file handles, which is easy to push when aggressively over-sharding. Not an uncommon mistake for the enthusiastic newbie. Having supported Solr/ES/Lucene in production for 4+ years now (websolr.com / bonsai.io) I would be pretty hesitant to trust Lucene in general as a primary data store. Beautiful for secondary indexing, but otherwise, Why Not Postgres?™ ;)
- RyanZAG 13y agoComplexity. Having two copies of the data means more dev time, more resources required to shift the data around, etc. Having just 1 data store that can also handle all your searching is like the holy grail. As you say, not sure if Solr/ES/Lucene are there yet - but they're definitely very very close. There is no theoretical barrier either - it just comes down to closing bugs, and the ES/Lucene team are very good at closing bugs. EDIT: I don't think MongoDB is there yet either. There are definite benefits and drawbacks between Postgres and ES, tipping heavily towards Postgres for structured heavy write data. But for ES and MongoDB? I think MongoDB falls a bit short there.
- nzadrozny 13y agoSure, that's a fair point. Data consistency reliability in ES and Lucene will only get better over time. But I personally suspect Lucene won't ever get away from the dreaded "just reindex." And to the larger point, I think recent resurgent interest in data stores and distributed systems have shown pretty clearly that there is no holy grail. No single data store can provide all the semantics necessary for all use cases. Maybe not even for most use cases. There are just too many tradeoffs to consider. Believe me, I earn a living hosting Elasticsearch, so I'd love to see it become a robust primary data store. There are some use cases where it actually does make sense—just look at the amazing traction ES is experiencing for storing and indexing time-series data. But as a general-purpose primary store, I'm not really holding my breath. Maybe I'm just becoming battle-worn and bitter. I would love to be proven otherwise over the next few years!
- jonhmchan 13y agoCongrats to the team - absolutely love elasticsearch. Having a lot of fun with it here at Stack Overflow.
- Argorak 13y agoBeyond the technology, Elasticsearch has a very mature, active and helpful community with users groups all over the world. We're well connected. Pick your favourite users group here: http://elasticsearch.meetup.com/ http://elasticsearch.meetup.com/ Full disclosure: I started and run the Berlin UG. We set ourselves apart by always providing a small introduction into ES for those that are completely new and would have a hard time following the main talk.
- shurane 13y agoIntros to ES and other technologies are useful. I don't see many tutorials covering usage of ES here: http://www.elasticsearch.org/tutorials/ http://www.elasticsearch.org/tutorials/ Could you maybe provide a link to yours?
- Argorak 13y agoThe introduction is in person, at the users group. Yep, tutorials is a huge problem, but there are people working on that.
- willcodeforfoo 13y agoCongrats! Elasticsearch is one of my favorite recent pieces of technology.
- philfreo 13y agoWe wrote a tutorial about how we wrote our search for Close.io using elasticsearch and pyparsing: "Sales data search: Writing a query parser / AST using pyparsing + elasticsearch" Part 1: http://blog.close.io/sales-data-search-writing-a-query-parser-ast-using-part-1 http://blog.close.io/sales-data-search-writing-a-query-parse... Part 2: http://blog.close.io/sales-data-search-writing-a-query-parser-ast-using http://blog.close.io/sales-data-search-writing-a-query-parse...
- alecco 13y agoWhy is it awesome? Why "it just works"? Is it just a mongodb-kind document store over Hadoop+Lucene? What makes it so special to have hundreds of votes and tweets all around within 2 hours? I don't understand. A DB engine engineer.
- gibrown 13y agoThere are a lot of features thoughtfully combined that make ES great. Top of my list would be: 1. It handles human written language. Any language. The same technology that let's it handle strings written in human language provides a lot of flexibility in handling string in other applications. Particular when handling logs. 2. Non-string data it also handles very fast and cleanly (numbers, dates, geo). 3. Lucene has an inverted index that has been optimized over many years. ES scales that pretty seamlessly across many servers. All decisions in the project seem to be made around whether a feature can scale to 100s of nodes. The devs have also been really smart to focus on the "out of box experience". Very well thought out defaults. More on our experience with ES at scale: http://gibrown.wordpress.com/2014/01/09/scaling-elasticsearch-part-1-overview/ http://gibrown.wordpress.com/2014/01/09/scaling-elasticsearc...
- buckbova 13y agoIs this accurate to elastic search since it is build on Lucene? https://lucene.apache.org/core/ https://lucene.apache.org/core/ "index size roughly 20-30% the size of text indexed" That seems excessive for an index.
- gibrown 13y agoNot sure how that's calculated. I assume it is accurate, but the index size is going to depend a lot on what kind of text you have and how it is separated into individual terms (or n-grams or all the other ways you can tokenize and filter to create individual terms). Personally, I think of disk space as cheap, and am far more concerned with having options to improve speed and quality of search results.
- ddorian43 13y ago
- mavelikara 13y agoES seems to have ability to run analytic queries. I have read about people using it as an OLAP solution [1], although I have not yet read anyone describe their experience. In that respect how does ES analytics capabilities compare against: 1) Dremel clones [2] like Impala & Presto (for near real-time, ad hoc analytic queries over large datasets) 2) Lambda Architecture [3] systems (where queries are known up- front, but need to run against a large dataset) Does anyone here have experience ES in such usecases, beyond the free text searching one ES is well-known for? [1]: https://groups.google.com/forum/#!topic/elasticsearch/iTy9IYL23as https://groups.google.com/forum/#!topic/elasticsearch/iTy9IY... [2]: http://static.googleusercontent.com/media/research.google.com/en/us/pubs/archive/36632.pdf http://static.googleusercontent.com/media/research.google.co... [3]: http://jameskinley.tumblr.com/post/37398560534/the-lambda-architecture-principles-for-architecting http://jameskinley.tumblr.com/post/37398560534/the-lambda-ar...
- zcrar70 13y agoI would also be interested in this.
- xutopia 13y agoI love when something I've been using in production for what seems like years just announces now that they've reached 1.0.
- brickcap 13y agoWell does it not make you feel glad that you took the risk? After all version is just a number :)
- sandstrom 13y agoThis gem is from the 'breaking changes' list: “Geo queries used to use miles as the default unit. And we all know what happened at NASA because of that decision. The new default unit is meters.” I like this release already.
- roryokane 13y agoLink to that page: http://www.elasticsearch.org/guide/en/elasticsearch/reference/1.x/_parameters.html http://www.elasticsearch.org/guide/en/elasticsearch/referenc...
- capkutay 13y agoI was vetting ES for a business critical search platform, had some concerns about write/read performance and how the lucene indexes are handled on disk. I read that it doesn't really perform as well a splunk...Instead of ES, I'm considering a solution using HBase to shard lucene indexes on HDFS.
- deleted 13y ago[deleted]
- axionike 13y agoES has performed very well for us as the backbone for the solution we deployed for a large government-sector customer. Had some GC issues initially, and were worried about user concurrency, especially since we were not restricting queries (i.e. users can do full-scale wildcard searches against the entire data set of 1BN+ records). But ES continues to shine. Congrats to the ElasticSearch team, and all the supporters around it. Once I get back into more of a coding role, I'll definitely be contributing back to the ES project.
- room271 13y agoThis may require a bit more lengthy answer than makes sense here, but I'm curious about what was causing your GC issues and how you fixed them (we have GC issues at the moment).
- polyfractal 13y agoNot the OP, but GC issues in Elasticsearch basically boil down to memory pressure (obviously), which is usually caused by facets. Facets eat a lot of memory, especially if you are faceting high-cardinality fields - think fields like "tags" or any analyzed field. High cardinality, analyzed strings is the easiest way to blow out the heap. There are other reasons, but that is like 90% of GC issues. To solve it, you need to make sure your faceted fields are configured well (usually not_analyzed) and assess how much memory is available. You may be able to index and even full-text search ten billion docs on a single machine, but faceting it may just be too much to ask for a single node. Omiting norms, disabling bloom filters on old indices and enabling doc values are other ways to help alleviate field-data pressure. Other GC culprits can be: too large bulk requests, unbounded threadpool queues, or something like parent/child/scripts/filter cache keys eating all your memory. Also don't go above 30gb heaps, the JVM becomes unhappy :)
- pron 13y agoWhat does Elasticsearch add on top of Lucene?
- lobster_johnson 13y agoA lot. Lucene is basically the inverted indexes, providing on-disk structures and a mechanism to query, as well as assorted bits like tokenization. ES adds distribution (multimaster-replicated cluster of nodes connected via a gossip protocol), sharding, defines a document model and schema (the mapping of arbitrary JSON documents to index structures), faceting, aggregation (ie., roll-up-type calculations), various types of scoring (eg., geographic distance), ETL ("rivers"), backup/restore, performance metrics, a plugin system (eg., for indexing different file formats) and a bunch of other things -- and of course a REST-based API on top of the whole thing.
- dreamdu5t 13y agoWe recently switched from using MixPanel + Crittercism + Sphinx to using qbox.io (hosted elasticsearch) and Kibana to do all our analytics, crash reporting, and search. I can't recommend qbox.io enough! Point-and-click scaling of managed elasticsearch clusters + Kibana == bliss.
- rartichoke 13y agoES is one of the few techs that I seriously love. The rails support for it is amazing too. The guy creating the rails integration lib is really talented and active.
- vhost- 13y agoI'd be curious to see how well Elastic Search holds up to Endeca. I'm currently stuck maintaining some Endeca instances and it's a nightmare. I wish I could go back to ES. At my last place of work, ES was beautiful and required little work to get a very fast, workable search in place.
- quicksilver03 13y agoFYI, at my shop we use Oracle Commerce (ATG) and we've seen Oracle's salespeople pushing Endeca to all current and new customers. For our current project we went with ElasticSearch and we're quite happy. One of the contributing factors was that one of our most experienced guys was unable to get the damn thing installed, even with the help of one Endeca consultant.
- hungryblank 13y agoAt Contentful in Berlin (Germany) we're looking for an elasticsearch/lucene expert, if you're excited by this tool and want to work full time with it get in touch. https://groups.google.com/d/msg/elasticsearch/Rb7Lei4gaaE/7IDPuPxQV-IJ https://groups.google.com/d/msg/elasticsearch/Rb7Lei4gaaE/7I...
- pyotrgalois 13y agoGreat news. In every new project that we create (in general REST JSON APIs made with nodejs, erlang or rails that are consumed by iOS and android clients) we always finish using postgresql, redis and elasticsearch. Great tools.
- elchief 13y agoAnybody know if elasticsearch does multiword synonyms properly? (Solr doesn't). Thx