6 ms·
Crawling speed is between 100...1000 pages per second. We crawled about 4 million linked unique web pages The pricing for an Index like DeepHN would be $99/mon
by wolfgarbe 5y ago
Crawling speed is between 100...1000 pages per second.
We crawled about 4 million linked unique web pages
The pricing for an Index like DeepHN would be $99/month: While we are indexing 30 million Hacker news posts, for DeepHN we are combining a single HN story with all its comments and its linked webpage into a single SeekSorm document. So that we index just 4 million SeekStorm documents.
Yes, it would be possible to expand index with embeddings (vectors) and perform semantic search. This would we an auxiliary step between crawling and indexing.
- freediver 5y agoWhat is being used as a crawler and is it integrated with Seekstorm? The same article referenced here https://deephn.org/?q=how+to+be+productive&filter=%7B%22hashtags%22%3A%5B%22CPU%22%5D%7D https://deephn.org/?q=how+to+be+productive&filter=%7B%22hash... contains the phrase 'well-defined' Any idea why doesn't the article surface when searching for this ?
- wolfgarbe 5y agoThe crawler is a part of SeekStorm. "Well-defined": Just a guess: We are doing key text extraction, i.e. we try not to index boilerplate stuff an and menu items. As "well-defined" is within a short list item, it might be accidentally skipped. So, that is not yet perfect, and considering the diversity in web page structure it probably never will. But we will try to improve.
- freediver 5y agoYes, if the preview is the representation of what you have indexed, then half of that article is missing. You may have identified the weakest link in your stack - crawler/extractor (which is notoriously hard to do, would be good if you provided more detail eg. do you use headless browser or simple GET request, do you crawl PDFs etc). Little use of the advanced stack on top of it, if the data does not end in the index in the first place. Hope you provide an update on this in the future. I'd probably sign up for a plan.
- wolfgarbe 5y agoNo, the preview is NOT the representation of what we have indexed. The preview ist limited to about 200 words to be compliant with fair use legislation. Indexing is limited to 1 MB per document.
- WalterGR 5y agoCrawling speed is between 100...1000 pages per second. Stupid question, but you were crawling news.ycombinator.com, right? Its robots.txt (https://news.ycombinator.com/robots.txt https://news.ycombinator.com/robots.txt) contains Crawl-delay: 30 Why did you not follow that? (I'm not being accusatory. I'm both curious about web crawling in general and have personally been archiving the front page and "new" every 60 seconds or so... (Obviously there's no reason for me to retrieve them more often, but my curiosity persists.))
- wolfgarbe 5y ago>> you were crawling news.ycombinator.com, right? No, for retrieving the Hacker News Posts we were using the public Hacker News API, which returns the posts in JSON format: https://github.com/HackerNews/API https://github.com/HackerNews/API The crawling speed of 100...1000 pages per second refers to crawling the external pages linked from Hacker news posts. As they are from different domains we can achieve a high crawling speed while being a polite crawler with a low crawling rate per domain.