3 ms·
I had been experimenting with web crawling with a lot of technologies. (Python based and others). What most (uninitiated) developers do not realize is that web
by deostroll 5y ago
I had been experimenting with web crawling with a lot of technologies. (Python based and others).
What most (uninitiated) developers do not realize is that web crawling is not for mere mortals.
1. We are at the mercy of the webpage authors.
HTML is a great lanuage to encode information. But most developers (usually webpage authors) see it as a kind of tool for presentation only. Infomation can go anywhere in the document. And they are prone to changes.
2. The internet society frowns on web crawling
You look into any site's TnC, you might come across a clause which prevents you from crawling. The specific word may not be in the legalese, however, it implies any kind of crawling is denied. There is good reason for this - it is mostly done to encourage fair use of the service.
3. No body designs services to be crawlable.
Most big name companies do have some alternative. Like Facebook had "graphs". (Now obsolete). They allowed end users to extract data using simple queries - like " list friends of X who live in city Y, and who is not your friend". But "graphs" feature came a lot later after Facebook launch. Not at beginning.
Usually at the beginning stages of any services we are at the mercy of #1 and #2
For #1 no one ever designs page to have have information always at a standard location. It changes.
4. The tech isn't ripe yet.
This is my personal view. I had been experimenting with puppeteer and selenium behind a corporate environment. I wasn't that happy with the "net" developer experience. I found things like taking a screenshot or pdf buggy. For e.g. to get the latter I have to run my browser in non-headless mode. In headless mode my laptop system policy disabled some extensions important for the webpage to load correctly.
- peanutz454 5y agoYep, I wanted to crawl a job search website so that I could search for jobs at work, without going to the job site (don't blame me, my job back then sucked). It was impossible to find information because all the tags were generated in some sort of framework that obfuscated everything.
- missblit 5y agoYeah architecturally Chromium Headless is an "embedder" so it doesn't automatically get all the front-end goodies full-fledged Chrome supports unless someone puts in the work to plumb through the code. So stuff like extensions don't work at all.