3 ms·
Scraping is as you might expect: it has to go fetch the page if it hasn't yet (URLs are only fetched once). One thing I can do to help here is pre-fetch/parse p
by jkkramer 16y ago
Scraping is as you might expect: it has to go fetch the page if it hasn't yet (URLs are only fetched once). One thing I can do to help here is pre-fetch/parse pages from popular sites. I'm already doing this for Food Network and Allrecipes.com.
The machine learning aspects aren't too bad. Most of the time is spent parsing the HTML DOM. I've done some micro-benchmarking but am very curious to see how it will handle under real load. I have a plan for moving from VPS (Linode) to AWS if necessary.