2 ms·
Thanks! Extracting info from the wild web is hard, and that's one of the reasons I did it -- because no one else was. Clojure + some basic machine learning made
by jkkramer 16y ago
Thanks! Extracting info from the wild web is hard, and that's one of the reasons I did it -- because no one else was. Clojure + some basic machine learning made it feasible.
- ambirex 16y agoHow processor intensive is the scrapping and machine learning? I guess how well do you think it could scale.
- jkkramer 16y agoScraping is as you might expect: it has to go fetch the page if it hasn't yet (URLs are only fetched once). One thing I can do to help here is pre-fetch/parse pages from popular sites. I'm already doing this for Food Network and Allrecipes.com. The machine learning aspects aren't too bad. Most of the time is spent parsing the HTML DOM. I've done some micro-benchmarking but am very curious to see how it will handle under real load. I have a plan for moving from VPS (Linode) to AWS if necessary.