2 ms·
I believe Instapaper uses a custom parsing engine for stripping content from pages. There's a page here which lists the site specific rules Instapaper uses for
by msmithstubbs 15y ago
I believe Instapaper uses a custom parsing engine for stripping content from pages.
There's a page here which lists the site specific rules Instapaper uses for extracting content:
http://www.instapaper.com/bodytext http://www.instapaper.com/bodytext
I think this is why you see such good results from Instapaper.
- cskau 15y agoInstapaper used to use some wrapper code around Readability and do all the work in the client browser. However, I'm a bit surprised to see that Marco has moved away from this approach. If you look at the code loaded with the bookmarklet[1], you can see it seems to simply compress the page (with Base64) and send it to the server, where it must be processed now. Likely this is done to shift the workload off the user, and make it simpler to maintain and modify. [1] http://www.instapaper.com/j/VOvVpvxqOIFR http://www.instapaper.com/j/VOvVpvxqOIFR