3 ms·
Actually the entire backend is on github. But let me explain it for you. There is as you have already discovered, a crawler implemented with scrapy. However I
by wallunit 14y ago
Actually the entire backend is on github. But let me explain it for you.
There is as you have already discovered, a crawler implemented with scrapy. However I don't use the scrapy server and pipelines. Instead I have a script that lets scrapy generate JSON files with the crawled recipes and builds the sphinx index from the crawled data. There is no RDBMS. Basically sphinx is my database. :)
Well and than there is the website. Its server-side is implemented with werkzeug and its UI with jQuery.
- tharshan09 14y agoOh i actually make scrapy do that to. Do you use the scrapy crawl with -o and -t json options? I did not know spinx can index json files etc. Any reason why sphinx was chosen rather than an alternative? Thanks for the info.
- wallunit 14y agoYes, I use "scrapy crawl <spider> -o <spider>.json" (see bootstrap.sh). Sphinx can not index JSON files directly, but it can index an XML stream written to stdout by a given command. So I have written a script (sphinx/xmlpipe.py), that reads the JSON files generated by scrapy and writes the crawled recipes in sphinx's XML format to stdout. There are alternatives to Sphinx? ;) * Sphinx is ridiculous fast, as you can see when searching. But even building the index takes only 340ms (from which 220ms are spend by the python script that generates the XML) for 1699 recipes on my 3 years old notebook. * Sphinx don't require a RDBMS to index documents from. It can index documents from any source. You just need to write a simple script that brings the documents in the XML format expected by sphinx. * Sphinx is not only a full text search engine. It is also a multi-value store. You can add extra information like the title and url to indexed documents. And so you don't need an additional database. * I need the ability to limit a fulltext search to sentence boundaries. I don't know if there are other fulltext search engines that can do that.
- tharshan09 14y agoI was quite surprised at how fast it was on the searches. It sounds really great and im sure I can put up with xml, I just dont like having to return HTML - but is that you preference and the way you have done it? It looks like its just return dict of results, so no reason why I cant form a json response I guess. How does a typical query look like? and how are you using the sentence boundaries (tbh im sure sure what it even is). I guess for fast text search this is perfect but im guessing any computation it cant do?
- wallunit 14y agoFor example, if you enter the ingredients "rye whiskey" and "vermouth", following query is generated: @ingredients (rye SENTENCE whiskey) | vermouth The SENTENCE operator, basically works like the & operator, just that both operands must occur in the same sentence. That is very helpful, since the index field "ingredients" contains a list of all ingredients of the recipe separated by an "!". However Sphinx can not only do full text search. You can also filter and sort by attributes and complex expressions, that involve any attribute, the relevance from the fulll-text search, arithmetics and some built-in functions. And Sphinx does that much faster than any RDBMS does. I never managed to generate a sphinx query that took a measurable amount of time, on my 3 years old notebook. ;) The sphinx index doesn't contain any HTML. However I return HTML, from the WSGI app, that serves the AJAX calls. So I guess that is what you are talking about. It just seemed to me, it would be simpler to generate the HTML for the search results on the server-side with Python than in Javascript. By the way if you want to use Sphinx for your own project, and you have the data to index already in a MySQL or PostgreSQL database, there is no need to write a script to generate XML. Sphinx can index data from MySQL and PostgreSQL databases directly. However I was just saying that, thanks to xmlpipe support, you don't need an RDBMS just to use Sphinx. And that thanks to index attributes, Sphinx can completely replace an RDBMS in a lot of cases.
- tharshan09 14y agoGreat I think you have me sold :) I want to give this a try since a standard db is what I would normally use to keep the data. This would be an interesting experience. I will be reading up on sphinx and mysql etc, and do any future data scraping straight to sphinx through xmlpipe. Last question :) - Do you have any useful links you can provide on the matter of sphinx etc? Thanks.