3 ms·
Just took a look at this, here's my guess. - Pretend they're a crawler such as Google and pull down the HTML, potentially executing javascript - Once it's pul
by brad0 8y ago
Just took a look at this, here's my guess.
- Pretend they're a crawler such as Google and pull down the HTML, potentially executing javascript
- Once it's pulled down, clean it up using open source code such as readability https://github.com/mozilla/readability https://github.com/mozilla/readability
- Store that result as a document in a nosql database
Once they have pulled the article down once they don't need to get it again.