4 ms·
I appreciate that you want access to their content, and I may personally agree they ought to build an API. But isn't it up to them to decide whether or not to d
by dewitt 13y ago
I appreciate that you want access to their content, and I may personally agree they ought to build an API. But isn't it up to them to decide whether or not to do so?
Otherwise isn't it just stealing?
(Fully expecting this to be a controversial comment, but I'm curious what the prevailing sentiments on scraping without permission are these days. Good, because all content should be free, especially crowd-sourced content? Or bad, because the contributors didn't agree to such third-party usage, and it could limit Rap Genius' ability to build a business?)
- timrogers 13y agoOP here. It's a fair point. I'd love to hear the thoughts of other HNers. I tend to think that this kind of thing is okay for personal use, whereas it wouldn't be en masse - for instance, if you were using it to power and advantage a competing music site. Others' opinions may differ. For instance, I've previously built scrapers for banking services (American Express [https://github.com/timrogers/amex https://github.com/timrogers/amex] and Lloyds TSB [https://github.com/timrogers/lloydstsb] https://github.com/timrogers/lloydstsb]) and the UK's university application system, UCAS (https://github.com/timrogers/ucas https://github.com/timrogers/ucas) and none of these organisations seem to have had any problem with it, but the library has been really gratefully received by users.
- xtrumanx 13y agoAny advice on how to build the API if a banking service uses 2-factor auth? Perhaps SMSing the secret number they send you back to your system (assuming you can't give your system direct access to phone number your the secret number will arrive at)?
- timrogers 13y agoInteresting question - I haven't had to try it yet. My scrapers are all read-only at the moment, and (afaik) none of the UK banks require two-factor in any form for that kind of operation. Potentially using Twilio (http://www.twilio.com http://www.twilio.com) you might be able to have an application which receives the two-factor request itself and handles it. It'd make the process async though, which wouldn't be ideal.
- dewitt 13y agoThoughtful reply, I really appreciate that. Thank you. And don't get me wrong—I love APIs, I build them for a living. And I know how much users love them. Unfortunately, not all companies love providing them, or have the economics of doing so fully worked out.
- timrogers 13y agoNo problem - great to engage with people's opinions. HN has such a wealth of knowledge. You might find my response below (https://news.ycombinator.com/item?id=6229617 https://news.ycombinator.com/item?id=6229617) interesting, on that front. My guess is that for Rap Genius, the only API that will make real sense is one offered on a commercial basis.
- xSwag 13y agoAs someone applying to university this year, I can't thank you enough for the ucas scraper!
- timrogers 13y agoHah - glad you like it. Did you find it before, or is this the first time you've spotted it? I built a cool frontend too that you might be able to cannibalise if you'd like to share your status online - see https://github.com/timrogers/ucas-frontend https://github.com/timrogers/ucas-frontend. If you wanna chat at all about any of this stuff, drop me a line at me@timrogers.co.uk.
- deleted 13y ago[deleted]
- simon_weber 13y agoI see non-sanctioned access like this as a huge boon to both the service and its users: it enables easy marketing, community-supported apps, easy research on what features users want, etc. I also think the fear of abuse is mostly unwarranted. If somebody wanted to steal your content or abuse your platform, they'd figure out a way to do so -- with or without an unofficial api (especially if all that's needed is http scraping). (disclaimer: I maintain https://github.com/simon-weber/Unofficial-Google-Music-API https://github.com/simon-weber/Unofficial-Google-Music-API)
- timrogers 13y agoOP here. I agree - for me, the only exception is where data is being abused for overtly commercial purposes. This obviously isn't one of those cases, and I'm doubtful as to whether unofficial APIs (i.e. screen scraping) can ever work properly for that kind of scale.
- simon_weber 13y agoRight, of course. It can get really hard to tell where the line is, though. For example, multiple for-pay third-party clients for Google Music now exist [0]. Personally, I think this is still a value-add for Google, but I could understand why some would think it's exploitative. Another thought experiment: say the unofficial use of the platform originally required a fair amount of reverse engineering. Can the person who figured it out be held responsible for other's abusive behavior? [0] https://itunes.apple.com/us/app/gomusic/id457883228 https://itunes.apple.com/us/app/gomusic/id457883228
- timrogers 13y agoVery interesting point - I'm not sure how best to judge what is exploitative and what isn't. I guess it becomes exploitative if it serious subverts in some way the business model of the original data/service provider. But then I'm not sure still...maybe that is okay if consumers massively benefit from the reverse engineering, and thus overall utility is increased...! What do you think?