4 ms·
Simple search (non-cached) like "hello world" or "i love you" throw a bad gateway. It took 10+ seconds for "hello world" when it worked. For my name "sandeep"
by sangupta 6y ago
Simple search (non-cached) like "hello world" or "i love you" throw a bad gateway. It took 10+ seconds for "hello world" when it worked.
For my name "sandeep" it threw 0 results, however a fine line says 2870 URLs in Index. I assume 2870 is not the entire corpus of the index - if yes, then search is extremely slow.
- shakna 6y agoThat is the entire current corpus. It sounds like you were hitting the site whilst it was under a bit of heavy load. From my logs, about the time you posted here on HN, someone was tossing `siege` at the site, and it is not a heavyweight server. Running certain searches do seem to be able to trigger a pathological response, however, so I'll need to look into that a bit more. Likely to do with some of the nltk stuff it uses when it tries to handle searching summaries.
- sangupta 6y agoWill try it again tomorrow. I could not find any information on site as to what will make this search engine different from what we already have? Also, the corpus is now 3162, roughly 250 links in last 2 hours - which is way too slow for a real world scenario. I though like the page for its simplicity and the experiment.
- shakna 6y agoI don't imagine that I, as a lone person, can actually build an engine to ever compete with Google or Bing, etc. It is not link crawling the entire web. Instead it is trying to find information that is current. You'll find relevant and current information from the BBC, CNN, and NPR for example. As well as stuff from places like LowTech magazine, the CCC, and FSF, etc. I haven't really put any information on the page at all about it. But mostly it scratches my own itch. But I wouldn't see it becoming a competitor if you're looking for "anything & everything", ever. (Though that 2 hours is actually you watching the nginx cache expire. The database updates hourly.)
- sangupta 6y agoGot it. So it sources data from a whitelisted set of sites and updates every hour. If the curated list can scratch my itch, I would love to come back. You mention you used some sort of NLP (mention of nltk before) - is it for summaries or reducing noise, or for bringing context to search terms.
- shakna 6y agoYeah, pretty much. I'm using nltk for - generating most (not all) of the summaries, and for building most of the tags that get attached to each article. It's also being used when searching the text of a summary.