3 ms·
They have their problems but how else am I supposed to scrape data from companies that want to hide it?
by brikym 2mo ago
They have their problems but how else am I supposed to scrape data from companies that want to hide it?
- BLKNSLVR 2mo agoI'm a bit on the fence about this, but leaning towards the "bad luck". I'm sure there's a large swathe of nuance that I'm missing, but my simplistic view is: If they don't want to be scraped then they don't get included in "the thing" which, at minimum, is a data point for consumers to make consumer decisions about. It strongly depends on how scraped data is being used. If your 'cat' hasn't been able to catch their 'mouse' then the cat needs to get smarter, or look for alternative sources of mice, or the cat should be considered 'unviable'. Have you approached them to get access to their data? Have you explained to them how your service can benefit their business? (I generally come from a position of suspicion as to why someone wants to scrape data that the owner goes to certain lengths to protect, but then I'm also an 'information wants to be free' kinda person, but the Internet is increasingly an untrustworthy place, so security is overruling narrative).
- harimau777 2mo agoI know there've been efforts to scrape information that MAGA is trying to purge from government websites.
- BLKNSLVR 2mo agoGovernment website data should be openly available one way or another (at least within the country, and with reasonable security provision against obvious maliciousness). If it has to be scraped then there may be other problems (which include lack of resources to make the data API accessible).
- alightsoul 2mo agoTake reddit for example. If you're not big tech they ignore you.
- BLKNSLVR 2mo agoI'm kinda "fuck reddit". That's cutting the tether from a _lot_ (there's no way to overstate this) of useful information, but that's potentially one of the great things about the open AI/LLM models, is that all that info is baked in there. Disclaimer: As far as I understand it. Please educate me if I'm way off the mark. Also, how much of the content of reddit is in Common Crawl? (same disclaimer applies to this comment) I also understand 'the archival mindset', I hoard a bunch of data. But I also understand the logarithmic graph of futility.
- AnthonyMouse 2mo ago> I generally come from a position of suspicion as to why someone wants to scrape data that the owner goes to certain lengths to protect Let's try this for example. Suppose you want to create a price comparison site. A lot of the major retailers don't want these, because they want the customer going to their site when they want to buy something, not to the price comparison site that tells the customer which retailer has the best price and then might not always be them. If the comparison site doesn't have the prices from retailers who sometimes have the best price then it can't serve its function -- customers still have to check sites manually. And comparison sites are pro-consumer whereas blocking them is anti-competitive, so the ones doing the scraping are the good guys.
- BLKNSLVR 2mo agoSuppose someone wants to create a price comparison site, but is unable to access all the data that would serve their customers in finding the lowest price. The I guess the site is either useless, or can be used in combination with customer's own research that includes sites not included in the comparison. The site is still useful, it's just not exhaustive. There's nothing much in the world that's exhaustive, it's all a range of percentages. Also, maybe the source that isn't scrape-able becomes less popular as a result of not being included in the price comparison list. If they're always more expensive, then neither you nor they have any advantage in listing on your site. If they're always cheaper, then your site may have no purpose to serve. Everyone seems to want _everything_ despite _everything_ not being necessary to provide a service. The lack of a certain set of data may itself be a data point, and should be used as marketing material for or against the company resisting the scraping. Just because residential proxies are somewhat of an 'easy answer' doesn't mean they're right. Use more creativity and imagination! Or follow a new idea. It's not like price comparison sites are a revolutionary idea. (I'm a consumer who happily does significant research before a buying decision, going to various sites and checking prices, warranties, model numbers, reviews, availability, delivery times, and all the shit. And prefer doing the research myself than trusting a price comparison site, so I'm not the target market, thus my bias in my commentary and opinion on that topic). > so the ones doing the scraping are the good guys. Everyone thinks they're the good guys. If there's profit involved then there's always an element of delusion.
- indianmouse 2mo agoExactly! I like the comment!
- BLKNSLVR 2mo agoApologies for the double post, but I've got an alternate perspective: Do you allow the data you've scraped to be scraped? Do you share it as freely as you desire the 'companies that want to hide it' would? Or do you consider the scraped data is 'hard earned reward for effort' and therefore has value that others should subscribe to your service for?
- fnord77 2mo agodata wants to be free
- akoboldfrying 2mo agoWhat's your online banking password?
- fnord77 2mo agoOTP, baby