4 ms·
I always like to read about things like this -- but if they think 1MM pages at 25gigs and 0.8 requests/sec is a lot then I can't wait for their service to get m
by coderdude 16y ago
I always like to read about things like this -- but if they think 1MM pages at 25gigs and 0.8 requests/sec is a lot then I can't wait for their service to get more popular. It's similar to someone saying they have a huge database with only a few million rows in it. They do have a good site design though. I wish them the best of luck. It only gets more fun from there. :)
- Swizec 16y agoIt is a lot in the sense that we're footing the bill for handling all of this data ourselves ... Also a million is such a nice milestone, I just had to write a blog about it :)
- coderdude 16y agoFrom your article: our scraper just isn't all that amazingly fast. It takes around two seconds to extract the meat of the article out of a decently sized and reasonably complex website. Since you're on AppEngine does that mean that your code to extract the article content is written in Python? If so, consider using the Python port of Readability.js: https://github.com/srid/readability https://github.com/srid/readability I haven't tested this one personally (yet), but before it came out I ported Readability to Python and it was very quick, less than a second to extract the content. I used lxml to parse through the X/HTML. You'll have to do a lot of tweaking to get the heuristics working more or less perfectly but it would be well worth it. (The following assumes you're downloading and processing the content using AppEngine) As for bandwidth concerns, you will probably want to just get a fat pipe into your place of business (or wherever you're operating from). Where I'm at it's only $100/mo for a 16mbit/down line. That would handle your growth for quite awhile. Once the content is all downloaded and extracted locally you can push it up to your AppEngine server.
- Swizec 16y agoActually we have tried python-readability. The problem is if you want to run it on AppEngine you can't use lxml (no C code) so you're stuck with BeautifulSoup. It was also hellishly expensive, I think that experiment cost us ~$10 an hour in CPU costs. As a result, it's better to stick with the javascript version of readability and run it in node.js. It's perhaps a dash slower than a python version would be, but the results are much much better and it's easier to keep up to date.
- coderdude 16y agoWow... that's pretty expensive. Better to just do all this stuff locally and save your money for serving up requests to your users IMO. Kudos for setting it up to work with node.js though, very neat.
- albertogh 16y agoI'm using lxml with a bit of C for the scrapper behind Printful. My algorithm is way more complex than Readability and I can still parse web pages really fast. For example, parsing a long article from NYTimes on a really old machine: fiam@raichu:~/printful$ cat /proc/cpuinfo processor : 0 vendor_id : AuthenticAMD cpu family : 6 model : 8 model name : AMD Athlon(tm) XP 2600+ fiam@raichu:~/printful$ python readable.py http://www.nytimes.com/2011/01/25/world/europe/25moscow.html http://www.nytimes.com/2011/01/25/world/europe/25moscow.html > /dev/null Took 0.193097171783 seconds for http://www.nytimes.com/2011/01/25/world/europe/25moscow.html http://www.nytimes.com/2011/01/25/world/europe/25moscow.html I'd say that hosting something like this on App Engine, which prevents you from using C extensions, is not a very good idea. Edit: formatting
- Swizec 16y agoWe're using node.js to run readability in. It's actually working out pretty well and is blazing fast on a dedicated computer (or run of the mill PC), the problem is really the overhead involved with virtualization. I think it takes around 0.5 seconds on a PC, a lot of this resulting from a pure javascript implementation of pretty much everything.
- mikeklaas 16y agoI found that about 40% of execution time was the the "double br to p's" step (running in mobile safari). Removing that speeds stuff up greatly without compromising quality too much. Also, the publicly available port of readability to python is ridiculously slow... I was about to make it about 10 times faster by re-implementing it myself. I can't release the source code, but the idea is to invert the algorithm: do one pass through the DOM collecting stats about children nodes, text length, and link density, instead of repeated queries at each step. You can get up to 10-20 docs/second on commodity hardware with this approach with an cElementTree DOM and python.