5 ms·
Just wondering if the guys can automate the scraping part, seems unlikely this is usually information that needs to be handpicked or "moderated" by one or more
by SadWebDeveloper 9y ago
Just wondering if the guys can automate the scraping part, seems unlikely this is usually information that needs to be handpicked or "moderated" by one or more humans.
- gruturo 9y agoSpeculation: it's possible that the source sites are actively hostile to scraping and intentionally mess with the layout regularly, or may even have ruled out any kind of automation in their ToS - so if you don't want to get sued or blacklisted, you can do no kind of (detectable) scraping.
- sleepychu 9y agoMaybe you're in a better position to negotiate this if you're them though? Skyscanner scrape to keep their partner's honest (which has always sounded pretty hostile to me but I guess they're approaching a pretty user hostile marketplace) and with their market share they're able to force their partners to comply with reasonable demands.
- SadWebDeveloper 9y agoThat's my main issue with this type of business... can't it be automated? probably a huge NO without legal implications therefore the "human" expenses are high and if we start to think "globally" like the Silicon Valley startups type guys always do, the next logical step is to make a mutual benefit business partnerships with the Airlines. This will end in a war like Airbnb vs Travel Agencies vs Hotel Chains vs Everyone else, were the best price in town will be to go directly to the source rather than to 3rd parties.
- teej 9y agoI tried to build a flight search site a few years back and many sites have measures in place to actively mitigate scraping. It quickly became obvious that I was going to spend more effort getting data than working on the product, so I scrapped the idea. Like you said, scraping is against ToS so as soon as you get caught, you're cut off until you shell out $$$$ to buy a feed. I suspect the reason they've succeeded here is that they've stuck with a human-based approach.
- blevin 9y agoCan anyone recommend scraping adapters (businesses or tech) that are robust to this sort of thing? I'm talking about something higher level than, say, Beautiful Soup -- something you can configure to point at an endpoint, essentially request a sql row subscription from it, and not have to mind it too much. Both the traversal/retrieval and data-interpretation parts seem to have interesting aspects when you consider current website design. Some websites make themselves hard even for humans to read (consider why safari reader mode exists). This seems like a potentially valuable service, in the sense of being a schlep. I wonder how many places have home-grown scraping efforts as part of their business and how annoying it is for them to maintain.
- RussianCow 9y agoAt work we use Mozenda[0] with success, though I can't tell you much beyond that because I'm not involved with that project. I've also heard of Agenty[1]. [0]: https://www.mozenda.com/ https://www.mozenda.com/ [1]: https://www.agenty.com/ https://www.agenty.com/
- gorkonsine 9y ago>so if you don't want to get sued or blacklisted, you can do no kind of (detectable) scraping. So why would it be hard to make scraping undetectable, anyway, unless you do it particularly incompetently? In theory, it seems pretty easy: use a browser string that matches an existing popular browser, and make sure to not load anything faster than a human would.
- tomarr 9y agoHave you done much scraping in the past? There's normally a lot more to it when javascript is involved, captcha systems etc. This is obviously helped recently by the relatively new headless modes for Chrome & Firefox, but before that it was using buggy headless implementations or Selenium. These weren't well suited to operations at scale.
- deleted 9y ago[deleted]
- danjoc 9y agoWho says it's not automated? https://workplace.stackexchange.com/questions/93696/is-it-unethical-for-me-to-not-tell-my-employer-i-ve-automated-my-job https://workplace.stackexchange.com/questions/93696/is-it-un... ;) Seriously though, automation may not be a good value proposition, even if the companies with the data that SCF is scraping manually want to cooperate. It seems the human scrapers are a fixed cost. It's not getting harder to do manual scraping as new customers come on board. The business is already paying this fixed cost. The business is currently doing well. All the business really needs to do is continue growing and the fixed cost will continue to shrink as a proportion of the business. Why would the business care in that situation? Reduce from 12 mid-range salaries to 3-4 high salaries? They would do well just to break even on salary expense, so why bother? If it ain't broke, don't fix it.
- SadWebDeveloper 9y agoHuman labor has always been seen as an expensive "disposable" business resource (primarily in the US) and for bussiness to grow that "big" they will need to scale globally, so they have three paths "optimize" human performance, automate web scraping or outsource to a 3rd party country. From the last three m inclined to think that this will be their answer to scale and only have 3-4 teams with high salaries moderating what gets published and what not.