15 ms·
Riot – Full-text search engine in Go
- dest 9y agoA comparison with Yacy would be interesting IMHO
- dpcx 9y agoNot that I'm against people building tools in their language of choice, but how does this compare to Sphinx (http://sphinxsearch.com/ http://sphinxsearch.com/)?
- didip 9y agoLast I used Sphinx, it is tied to MySQL. Is that still true? If so, then Sphinx could be a deal breaker to some.
- dpcx 9y agoSphinx can connect to MySQL or Postgres, but can also read XML and [TC]SV, among others. Does this do something better? The examples seem to be where the user is providing the data to the indexer, which Sphinx can also do.
- mrweasel 9y agoIt hasn't been tied to MySQL in the last 10 years, so I'm wondering if it ever was. You can connect to Sphinx using a MySQL client, use it as a MySQL storage engine or using MySQL as a data source. But it's not specifically tied to MySQL.
- pQd 9y agoIt's worth mentioning that the original Sphinxsearch project has been stagnant for the past year. There's a new lively fork - https://manticoresearch.com/ https://manticoresearch.com/ Beyond being a user of both I don't have affiliations with either of them.
- dpcx 9y agoGood to know. Any idea what happened?
- pQd 9y agoThat's what I've gathered based on the online discussions: Possible reasons why the original project stagnated: [0]. Mention of the manticoresearch as a fork project was removed from the sphinx forum [1] - so I can guess that developers who moved to the new project did not part on good terms with Andrew - the original author. that's the last reply in that thread from person involved in Manticore, posted on the 23rd of Oct 2017: " aditirex just replied to 'Sphinx search fork': ===cut=== > But your the people who are already using sphinx, why we should change? The open-source version of Sphinx received 5 code commits since November last year, from which 3 are related to building stuff. Last release was 12 months ago. There are also a lot of unresolved reported bugs (many of them are crashes) in the bug tracker. Andrew said a while ago that the open-source version would only receive fixes (which doesn't seem to happen either). No one wanted to do the fork, it was the only way several big users saw it in order to continue using Sphinx. Don't ask me how we got into this situation, I'm not the right person to answer to that. > What are the main benefits rather than using Sphinx? Manticore is pretty much continuing Sphinx. Last year we had 4 developers + Andrew working on the code, 3 of them are working now on Manticore. If Sphinx just works for you there is no reason to switch. But we're adding new features, fix existing bugs, the software is tested by some big users before getting released, you get a software that has support from it's developers. > Is Foolz\SphinxQL\SphinxQL working? Everything works as before, it's a fork, not a total new software. " [0] http://sphinxsearch.com/blog/2017/07/24/sphinx-2017/ http://sphinxsearch.com/blog/2017/07/24/sphinx-2017/ [1] http://webcache.googleusercontent.com/search?q=cache%3Ahttp%3A%2F%2Fsphinxsearch.com%2Fforum%2Fforum.html%3Fid%3D1&oq=cache%3Ahttp%3A%2F%2Fsphinxsearch.com%2Fforum%2Fforum.html%3Fid%3D1&aqs=chrome..69i57j69i58.1031j0j7&sourceid=chrome&ie=UTF-8 http://webcache.googleusercontent.com/search?q=cache%3Ahttp%...
- Jemaclus 9y agoAlso, for consideration: Bleve (https://github.com/blevesearch/bleve https://github.com/blevesearch/bleve) I'm in the process of building my own search engine (as a learning exercise, but also because it's related to my day job). I've learned that it's one thing to write a full-text search engine, like this one, and it's quite another to do field-specific searches with faceting support and so on, like Algolia and Lucene-based search engines do. That said, this is clean and simple. I like it. I can definitely learn from this.
- CSDude 9y agoLast time I checked, bleeve was very slow compared to Lucene, I really hope it gets better over time.
- markpapadakis 9y agoSupporting faced search and other functionality requiring access to per document field-values is just an extension over the core IR functionality. Tracking (document, field) values can be used for query by range or by geolocation primitives (that's what Lucene does, where it will index that data into a special tree-like structure, and for each query, it will build a custom 'iterator' and use it along with other iterators to match documents), and for static ranking of matched documents. BTW, Lucene and Algolia are vastly different in terms of the underlying architecture.
- lobster_johnson 9y agoWhat is Algolia's underlying architecture like? Are there any papers or code?
- markpapadakis 9y agoSee links about Algolia arch (and other related material) here: https://github.com/phaistos-networks/Trinity/wiki/IR-Search-Links https://github.com/phaistos-networks/Trinity/wiki/IR-Search-...
- 9y ago
- unicornporn 9y agoKind of taken https://about.riot.im/ https://about.riot.im/
- lylejohnson 9y agoNot to mention http://riotjs.com/ http://riotjs.com/.
- speps 9y agoWhat about https://www.riotgames.com https://www.riotgames.com ? More so because they have an engineering blog, which is very interesting : https://engineering.riotgames.com/ https://engineering.riotgames.com/
- literallycancer 9y agoMentioning Pendragon et al. in a discussion about originality? Rich.
- whyever 9y agoI think the company is a bit bigger than one person by now.
- krautsourced 9y agoAlso http://luci.criosweb.ro/riot/ http://luci.criosweb.ro/riot/
- throwawaymaroon 9y agono one in JS land gives a crap about riotjs though
- tenken 9y agoAlso for consideration: https://github.com/coleifer/scout https://github.com/coleifer/scout > RESTful search server written in Python, powered by SQLite.
- wolfgarbe 9y agoIt seems the whole index is kept in RAM. Thus the index size is limited by the amount of RAM available. This explains the impressive indexing and search performance (1M blog 500M data 28 seconds index finished, 1.65 ms search response time, 19K search QPS) The Persistent storage data is stored to the hard disk solely when the program closes. The data is then restored from the hard disk when the program restarts ( https://github.com/go-ego/riot/blob/master/docs/zh/persistent_storage.md https://github.com/go-ego/riot/blob/master/docs/zh/persisten... ). This is a limited approach compared to Lucene/Solr/Elasticsearch LSM which handle high-volume inserts to its indexes with a log-structured merge-tree (LSM) and where the index size is only limited by the available hard disk space.
- markpapadakis 9y ago1.65ms for what kind of queries? Also, is that 1M blog posts, all weighting 500mb in total size(characters)?
- ethanwillis 9y agoI wonder if they're using succint data structures. I'm in bioinformatics and the first time I implemented a wavelet tree to reduce the size of genomes in memory.. It was just breathtaking.
- ausjke 9y agovery interesting, can you elaborate and on it a little more? I need quick fuzzy search on a low-end embedded device that has limited storage(both RAM and HDD), was thinking about putting the index on a server with plenty RAM then do websocket or RPC for that.
- ethanwillis 9y agoThere's a very good blog post for the implementation details here: http://alexbowe.com/wavelet-trees/ http://alexbowe.com/wavelet-trees/ I had a decent implementation in Python, but it's on my old macbook that I would need to dig up. If you're interested you can add me on telegram: @rightcheek. Now to go with Wavelet trees you may or may not need to know about suffix arrays and optimal suffix array construction. Take a look at this: https://en.wikipedia.org/wiki/Suffix_array https://en.wikipedia.org/wiki/Suffix_array This is what's going to give you space efficiency in combination with a wavelet tree. And the wavelet tree also gives you good rank/select efficiency. Edit: Here's a suffix array construction algorithm implementation I did (not sure if it's fully correct) https://github.com/ethanwillis/comp7295_final/blob/master/saca.py https://github.com/ethanwillis/comp7295_final/blob/master/sa... It is based on this paper: https://local.ugene.unipro.ru/tracker/secure/attachment/12144/Linear+Suffix+Array+Construction+by+Almost+Pure+Induced-Sorting.pdf https://local.ugene.unipro.ru/tracker/secure/attachment/1214...
- conmarap 9y agoAh, this is awesome! So far I had to rely on Lucene++, and it could get complicated at times. Using go is something I had been wishing for.
- veni0 9y agoYep, go can be a simple deployment.
- wiremine 9y agoSidebar: I wish open source authors would think a bit harder about naming their projects. Here's some other projects already named riot: - http://riotjs.com/ http://riotjs.com/ - https://riot-os.org/ https://riot-os.org/ - https://github.com/vector-im https://github.com/vector-im
- KitDuncan 9y agoAnd riot games. Basically the worst name possible for SEO.
- KGIII 9y agoThen again, the language is 'Go.' I am not sure that was the best naming choice and I've heard complaints that it was initially difficult to search for it. My understanding is that 'golang' has helped as a query. So, I guess the situation has improved but I understand it was problematic at first. To be clear, I've never used the language. I actually dislike programming, though I've decided to get back into it because I have a couple of projects I want to poke at. I'm now deciding between Java and Python. Maybe I should do an 'ask HN' submission.
- cr0sh 9y agoIf you dislike programming, Python may be the saner choice.
- ShabbosGoy 9y agoBrainfuck is the sanest choice.
- KGIII 9y agoI'm not new to programming. I just haven't done much for a whole lot of years. Brainfuck, I'm familiar with it, is never the sanest choice. I've narrowed it down to Python or Java. I've done C, C++, BASIC, QBASIC, Perl, PHP, and even some COBOL. I've played with a few others.
- ct520 9y agoDoes anyone know anything like this that will search pdf text, or tiff?
- mikey_p 9y agohttp://tika.apache.org http://tika.apache.org Which I'm pretty sure can be embedded in Solr, has plugins for Elasticsearch and others.
- maxpert 9y agoI wish I could understand mandarin, such nice projects need some good translation.
- veni0 9y agoThanks for understanding, the document is improving.
- sAbakumoff 9y agoI mean, we have bleve[0]. What else do you need, really? [0] https://github.com/blevesearch/bleve https://github.com/blevesearch/bleve
- veni0 9y agoMulti-language and distributed support, simple.
- sAbakumoff 9y agobleve search does support multiple languages.
- veni0 9y agoBut like the Chinese must use plug-ins.
- billconan 9y agocan this replace elastic search? I'm looking for a light weight elastic search alternative.
- hardwaresofton 9y agoA little late but if you didn't know you could do reasonably fast Full Text Search with SQLite... Now you know: https://sqlite.org/fts3.html https://sqlite.org/fts3.html https://sqlite.org/fts5.html https://sqlite.org/fts5.html
- burntsushi 9y agoDoes it suffer the same limitation as PostgreSQL's fulltext search? i.e., It doesn't use corpus frequencies in its ranking function. (I skimmed the docs but couldn't immediately find my answer.)
- hardwaresofton 9y agoI'm not sure -- but I'm going to guess yes? I tried to look around and couldn't find an answer either...