12 ms·
Show HN: Spider Pro – easy and cheap way to scrape the internet
- uptown 7y agoInteresting that their demo screen cast shows Zillow. Zillow is fairly aggressive in applying anti-scraping defensive measures.
- mahesh_rm 7y ago.. which may be the reason why their demo screen cast shows Zillow.
- mehhh 7y agoI wonder how Spider Pro does with Facebook, Linkedin, Whitepages and others that try their best to block scraping but still have an introductory free to view webpage...
- ianmobbs 7y agoSince this is designed for non-technical users and only scrapes content that's already been displayed to the user, I can't imagine many folks would use it in such a way that they could tell, unless they included a script to detect this scraper explicitly on their site
- uptown 7y agoYeah, the problem is that Zillow imposes IP bans on you when you've been found to be scraping their site.
- invisiblerobot 7y agoWhich is why any serious effort involves rotating pools of proxies.
- walrus01 7y agoNot just rotating pools of proxies but sometimes shady gray market residential proxies, so that you can appear to be coming from hundreds or thousands of unique geographically distributed end-user DOCSIS3/ADSL2+/VDSL2/GPON/whatever last mile end user customer netblocks. If you want to go down a rabbit hole of shady proxies run on compromised/trojaned end user SOHO routers or PCs, google "residential proxies for sale" https://www.google.com/search?client=ubuntu&channel=fs&q=residential+proxies+for+sale&ie=utf-8&oe=utf-8 https://www.google.com/search?client=ubuntu&channel=fs&q=res...
- core-questions 7y agoOnce worked for a place using this to scrape search engines. It's amazing how easy and comparatively cheap it is to get access to thousands of residential IPs. Is it via spyware running on people's machines? Shady people working at ISPs doing nefarious things for cash? We never knew.... The key thing to know is that if you want your traffic to come from an IP "in" some other country (according to geolocation databases anyway) it's really only a few bucks a month to get a proxy. Most of them have poor IP reputation so they suck to use on Google, but work very well for everything else out there...
- luckylion 7y ago> Is it via spyware running on people's machines? Shady people working at ISPs doing nefarious things for cash? Might be as simple as https://hola.org/ https://hola.org/ & https://luminati.io/ https://luminati.io/ - "unblock a website, download our VPN client", meaning you "unblock" by using somebody else's line. And the also sell access at luminati. Most users aren't aware of the implications.
- walrus01 7y agoIt's a combination of three general things: a) The type of "services" luckylion mentions where people have opted in to a shady gray market thing reselling proxies through their connection. b) compromised home routers/gateway devices/internet of shit devices c) compromised home PCs (mostly windows 7/10 trojans/botnets)
- not_a_cop75 7y agoStep 1) Invest money in non-Zillow real estate app Step 2) Hammer Zillow with all known ip addresses Step 3) Profit
- mycall 7y agoIP bans are simple to bypass.
- deleted 7y ago[deleted]
- bdcravens 7y agoWorth noting it really doesn't automate paging through results, and they go out of their way to make the behavior seem organic, and explicitly say they won't change that approach. "Automating things in this way could put load on servers in a way that a manual user couldn’t, and we don’t want to enable that behavior."
- hbcondo714 7y agoAnd their documentation shows them scraping HN: https://www.notion.so/Spider-Pro-Documentation-5d275abd49c64b3185a0bb53ebae4ae0#1f4a16dead214ae69bcdc36c1dd2989e https://www.notion.so/Spider-Pro-Documentation-5d275abd49c64...
- dang 7y agoPlease respect the robots.txt if you do. HN's application runs on a single core and we don't have much performance to spare.
- walrus01 7y agoSerious question, how is that possible? Somebody recently gave me a Dell R720 2RU server with 16 cores and 128GB of RAM for free. There's literally that much slightly used server gear showing up on the used market from companies that have migrated everything to aws/gcp/azure/whatever. If all of HN has only a single core then you're running it on less server resources than I could buy on ebay with $180 and a visa card? https://www.ebay.com/itm/DELL-R610-64GB-12-CORE-2X-HEX-CORE-2-26GHz-L5640-PERC-6I-2X-146GB-SAS-HD/352793823865?hash=item5224268a79%3Ag%3AmukAAOSwVARdcVfr&LH_BIN=1 https://www.ebay.com/itm/DELL-R610-64GB-12-CORE-2X-HEX-CORE-...
- meritt 7y agoIt's a single-threaded process running a single core.
- chucksmash 7y agoDon't know how much it is still true, but HN was originally implemented in arc. The language homepage[1] says "Arc is unfinished. It's missing things you'd need to solve some types of problems. [...] The first priority right now is the core language." Perhaps parallelism is still pending. A Ctrl-F on the tutorial doesn't turn up any hits for "process", "thread", "parallel", or "concurrency". [1]: http://www.arclanguage.org http://www.arclanguage.org
- 7y ago
- porker 7y agoThis has a nice UI. It reminds me of Kantu (now https://ui.vision/ https://ui.vision/) which I've used with varying degrees of success. That works by recording Selenium scripts; is Spider Pro entirely custom?
- johnwheeler 7y agoAs an aside, it looks like this was created by one person, which shows an amazing level of talents in design, UX, programming, marketing, and presumably devops. Kind of scary.
- shujito 7y agois that a bad or a good thing?
- meddlepal 7y agoWell it makes them a damn unicorn so... good thing?
- blunte 7y agoI suspect the "scary" comment may not have been understood because of ESL, perhaps. So I think the person who asked if it was a good thing wasn't being sarcastic.
- samstave 7y agoYeah but when was the last time you actually caught a unicorn stabbing the fuck out of someone with their horn... so I vote scary
- bob1029 7y agoDepends on the unicorn. Some can spin up the full stack, put it under source control, wrap a CI/CD pipeline around it, and have it deployed to a TLS-encrypted public website in under an hour. I think you can actually find the above level of developer - what many would classify a 'unicorn' - in approximately 1-2% of the workforce. That doesn't mean they understand your business or products, know how to work well with other individuals, etc. There are a lot of other factors beyond just the productivity angle. Now, the above developer also with the ability to fundamentally understand the business, as well as interact with all of the key people while performing said productivity stunts on a consistent basis... I think you are getting into the .1-.01% range.
- joshmn 7y agoThis is scary good and easy to use. Clever idea with the Chrome extension too.
- mmcwilliams 7y agoClever, but how long until sites learn to detect the extension and block the user?
- penagwin 7y agoDetection is always a cat and mouse game. Using an extension in a real browser (instead of electron/headless chrome) is probably one of the hardest to detect because it requires running a "real" browser. Of course somebody will find a way to detect it, then the extension maker will patch it, and the cycle will continue.
- mmcwilliams 7y agoCorrect me if I’m wrong, but is it not still as simple as knowing the “chrome-extension://“ unique id of the extension? I’m aware of the cat and mouse aspect of scraping and that was one of the pitfalls I’ve been wary of as a fingerprinting vector.
- SquareWheel 7y agoI'd be surprised if sites had permission to read a chrome-extension:// URL. That'd be a sizable privacy leak.
- mmcwilliams 7y agoI'm not sure about the chrome-extension protocol, but this API still seems to be present: https://developer.chrome.com/extensions/runtime#method-sendMessage https://developer.chrome.com/extensions/runtime#method-sendM...
- ct520 7y agogood-time to showcase given this was recently in the news. Looks awesome thanks for sharing https://news.bloomberglaw.com/privacy-and-data-security/insight-linkedin-data-scraping-case-9th-circuits-trigger-for-cfaa-liability https://news.bloomberglaw.com/privacy-and-data-security/insi...
- seanwilson 7y agoAny comments on using Gumroad as your payment processor? I went with Paddle for a Chrome extension - it seems more flexible (e.g. you could do team subscription plans with custom pricing) but I think Gumroad is a lot easier to integrate (e.g. they have a simple license check API).
- visualphoenix 7y agoThis is lacking clarity in the payment schedule... Is it One-time? Monthly? Yearly?
- ithkuil 7y agoEDIT: snippet from the website to help answer this question without requiring you to click. I didn't intend this to be a rude answer, I don't think it deserves downvotes. " + No server involved (zero subscription fee for you!) ... Unlike other web scraping softwares, it requires only one time payment to scrape for unlimited time and data. No more subscriptions or huge fee for your small data analysis projects! "
- gravypod 7y agoHow is this at all sustainable as a business model? My company is in charge of scraping a few TB/week. Would we be allowed to use your service?
- ithkuil 7y agofrom a cursory look it seems that this software runs on your client, i.e. it's not a service. Scrolling down: " Can I scrape password protected stuff with Spider? Yes! It’s a browser extension, so as long as you log in first, you can scrape whatever you like. "
- chrismarlow9 7y agoGo watch the demo video. From what I can see it's not automated enough to do tb/week. That said I bet you could hook it into scrapy and automate it. Probably will be slow though
- evandena 7y agoIt's just a browser extension
- thinkloop 7y agoIt's all client-side, it's not what you think, it's more a light scraping-on-the-go type product than a permanent enterprise solution. Good for what it does tho.
- bvm 7y agoAh it reminds me of Kimono Labs. I miss that product, it was fantastic.
- basch 7y agohttps://www.diffbot.com/ https://www.diffbot.com/ is still well regarded, no?
- soared 7y agoI was trying to remember the name! I set up an incredible amount of automation with Kimono Labs.. that was one of the best products I've ever used. I remember when they shut down, but only realized now they were acquired by... Palantir?
- MattyMc 7y agoMe too. Absolutely one of the best products I’ve ever used. It’s also the product that taught me that I can’t trust a SaaS business with important work. Wish they still existed.
- hbcondo714 7y agoAlso seems similar to Dashblock (YC S19): https://news.ycombinator.com/item?id=21006475 https://news.ycombinator.com/item?id=21006475
- amitport 7y agoalso similar to parts of yahoo! pipes and some of dapper features (acquired by Yahoo!) I wonder why it was shutdown
- tigerBL00D 7y agoIt's like http://80legs.com http://80legs.com which, by the way, was acquired.
- btbuildem 7y agoSeems like a sure-fire way to get bought out is to build a web-scraper service...
- chirau 7y agoHow is it the cheapest if it is not free? There are plenty of free scrapping tools out there
- bluetidepro 7y ago^ Came to say the same thing.
- tmikaeld 7y agoHow many are offline-only? Almost everyone one chrome web-store are online-based and "freemium" only.
- CobrastanJorji 7y agoSorry, maybe this is obvious, but what is an offline web scraper?
- dang 7y agoWe've unsuperlatived the title above.
- turtlebits 7y agoSeems to work well, but it appears that you can't re-scrape a site without starting from scratch, it's a one time deal.
- tmikaeld 7y agoYeah, it needs a lot of additional features. I hope this isn't one of these "abandon-ware" kind of deals.
- deleted 7y ago[deleted]
- impatientduck 7y agoNonsense - beautiful soup is free. Don't pay for something you can program yourself.
- Cyberdog 7y agoSure, everyone's time is free too, after all. (Alternate smart-ass reply: Surely you coded the browser, operating system, and network stack involved in posting this response by yourself, right?)
- sillydinosaur 7y agohttps://www.crummy.com/software/BeautifulSoup/bs4/doc/ https://www.crummy.com/software/BeautifulSoup/bs4/doc/ I mean... soup.find(id="link3") # <a class="sister" href="http://example.com/tillie" http://example.com/tillie" id="link3">Tillie</a> Is this hard?
- Cyberdog 7y agoFor HN's audience, no, that's not hard. But this product is clearly not only for us, and you've also cherry-picked a very simple example. There's also the aspect of downloading the page in the first place and dealing with things like authentication and bot detection which a product like this helps solve. I personally don't have a use for this product right now, but I won't be so bold as to say I'll never find a case where using it wouldn't be easier or more cost-effective than hacking up my own solution.
- penagwin 7y agoIf you show that to non-programmers they'll take one look at that and understand that about as well as egyptian hieroglyphics. "Soup? Python? What's an 'id'? class? href? I just want to select the title!"
- slig 7y agoGood, now do that on a page that uses JS and authentication. Trivial on this extension, not so much using BS.
- deleted 7y ago[deleted]
- wumms 7y agoError info: contact us at the bottom shows page not found.
- atarian 7y agoSince the extensions are being distributed via Gumroad, how would updates work?
- neovive 7y agoNice example of Tailwind CSS website as well.
- ikeboy 7y agowebscraper.io is free and has more functionality. I've used it for a quick and easy way to scrape data off multiple pages. Probably has a slightly higher learning curve but once you get past that it's easy.
- riantogo 7y agoI'm not familiar with web scraping but have been looking for feed(s) from shopping sites (e.g. deals.ebay.com, deals.amazon.com etc.). In good ol' days they used to publish RSS, but not any more. Can I use this for what I'm looking for? Will eBay and Amazon end up banning me? Alternately does anyone know of a good service that aggregates shopping feeds?
- seeekr 7y agoWhile looking at the tool I've realized that I've built something similar many years ago. I wonder if it's worth digging up the source code, polishing and publishing it. Does the market need more of these tools? Are there features in this type of tool on the market that seem to be completely missing or inadequately implemented? And it's likely that running these as a server-side application is ultimately more appealing than this simpler (to implement) automation inside the user's browser, right? Seems that many companies providing similar (server-side) scraping tools have been successfully sold off... Is that an indication of (still existing) high demand for these? EDIT to add: The tool and the accompanying website look fantastic, by the way. Congratulations on the 1.0 launch, Amie!
- meritt 7y agoDoes the market need yet another browser extension scraper like this? No, I don't think it does. [1] https://chrome.google.com/webstore/search/scraper?_category=extensions https://chrome.google.com/webstore/search/scraper?_category=... [2] https://addons.mozilla.org/en-US/firefox/search/?platform=windows&q=scraper&type=extension https://addons.mozilla.org/en-US/firefox/search/?platform=wi...
- deleted 7y ago[deleted]
- Mathnerd314 7y ago25 is not very many on AMO, there are only a few that seem relevant and most are unmaintained. The only one that really seems usable is the webscraper.io one. Maybe ScrapeMate but the author has left and is working on the Python/cloud ScrapingHub. Similarly Chrome has Web Scraper Plus w/ poor maintenance and the rest are wrappers/helpers for various websites. The extensions market isn't particularly lucrative, they're like mobile apps but with only 33% of users even knowing they exist. The SaaS companies don't have a growth limitation hence their success. But if you want an extension for a resume item or something getting a few thousand AMO users or Chrome reviews shouldn't be hard.
- xurias 7y agoUnfortunate that the contact us page doesn't work. I'm interested, but I wanted to ask if a FF extension is planned, since that's my primary browser.
- s3nnyy 7y agoHave a look at dashblock.com, they went through YC recently
- Sephr 7y agoThe DRM scheme could use some work. Here's a simple crack (run from the extension license page): chrome.storage.sync.set({ spider_valid_license: { key: 'No license (00000000-00000000-00000000-00000000)', lastChecked: new Date('Jan 1 3000') } })