4 ms·
Hey all, Fist project creator here. Just wanted to thank everyone for the support and the comments. I’ve been working on this project for a few months now mainl
by max0563 7y ago
Hey all, Fist project creator here. Just wanted to thank everyone for the support and the comments. I’ve been working on this project for a few months now mainly as a passion project. It’s awesome to see some people taking an interest.
I could definitely use some help on this. Whether it be giving advice or writing actual code.
If anyone would like to chat more about this I have setup a slack channel for be project. Here is the invite link
https://join.slack.com/t/fist-global/shared_invite/enQtNjcyNzY4MTUwMDg0LTRiYzM5ZWNkOTMwODYzODRjNDQzNThiYjdhNjgzZDUxZGYxODRjOTI4NTcwYmYzYmI5MTViYjFiNGFlNWEwYjY https://join.slack.com/t/fist-global/shared_invite/enQtNjcyN...
Thank you again for all the support, I’m excited to continue development.
- michelpp 7y agoAs others have pointed out there are some features considered essential in this space that would be good additions to the code: Stemming: (removing prefix/suffixes walked, walking, walker -> walk) The snowball library contains many language stemmers https://snowballstem.org https://snowballstem.org Stopping: removing many common words like the/and that almost every document will surely match anyway. Again, per language. Relevance ranking: Generally some variant of TF/IDF like BM25. The book "Managing Gigabytes" is an excellent intro to the subject and information retrieval in general: https://people.eng.unimelb.edu.au/ammoffat/mg/ https://people.eng.unimelb.edu.au/ammoffat/mg/ A Document/Term/Value model, this is how Lucence, Xapian, and other IR systems model the store. It's worth sticking with that same pattern. Good luck!
- max0563 7y agoThis is very helpful. I am working on all of this, but summarizing it all here helped me a lot. Thank you.
- aasasd 7y agoPretty much everyone here knows Lucene as the primary text-search engine these days, but there's a competitor Sphinx Search which is written in C or C++ and is much lighter, to the point of being feasible as bundled with desktop software. It however has a bit weird database structure since text indexing is the primary use-case―this is being corrected in the third version with support for secondary indexes, but v3 isn't open-source yet. Not to be confused with Sphinx the documentation builder.