5 ms·
Using the RSS feed for the content seems clever, and useful as a means to build the static search widget outside the content generation phase of site building (
by nprescott 9y ago
Using the RSS feed for the content seems clever, and useful as a means to build the static search widget outside the content generation phase of site building (or for adding after the fact).
I recently had a similar (less well fleshed-out) idea[0], but built the search corpus during the generation of my static site (this was made easy by the fact that I wrote the generator too). The results aren't nearly so impressive, but it was an interesting experience that lead to all sorts of interesting ways to improve searching, like stemming, lemmatization[1] and (removing) stop words (this was mostly to reduce the size of the text corpus).
[0]: https://idle.nprescott.com/2017/text-search-on-a-static-blog.html https://idle.nprescott.com/2017/text-search-on-a-static-blog...
[1]: https://nlp.stanford.edu/IR-book/html/htmledition/stemming-and-lemmatization-1.html https://nlp.stanford.edu/IR-book/html/htmledition/stemming-a...
- AndrewStephens 9y agoThat is a really nice way of doing it. If you are going to the trouble of statically generating your site, you might has well statically generate a well-structured index at the same time.
- pmlnr 9y agoI'm doing this as well, the python whoosh library is doing a reasonable job there.
- gjstein 9y agoYour post was a great read, and I might try doing something similar on my blog. A suggestion though: I typed in a query with an uppercase letter and it failed to find any matches. Perhaps .lower() your query if it's something you'd like to maintain? Thanks!