7 ms·
I have looked into making a business like this before, there are quite a few of them and I do like scraping. But don't you have to break a lot of 'terms of use
by utnick 17y ago
I have looked into making a business like this before, there are quite a few of them and I do like scraping.
But don't you have to break a lot of 'terms of use' agreements to scrape this data? Could you get in legal trouble for that?
- ashishk 17y agoI think that's something to worry about if sales grow to significant levels. There are probably several solutions for that problem.
- cschneid 17y agoMaybe, but there's a decent legal argument that it's not copyrightable data (facts of where stuff is for example), and that it's publicly posted ("find a store near you" links). I suppose it gets more and more fuzzy if you move into non-fact data, and you open yourself up more to lawsuits.
- lonestar 17y agoBy the same legal argument, couldn't anyone buy one of AggData's datasets and then publish it for free on the internet?
- wallflower 17y agoI imagine that you have to sign something/checkbox an agreement when you purchase a dataset that makes you legally and/or financially at-fault if you are discovered to be the leak/source of the leak
- skolor 17y agoYou couldn't distribute their dataset, but you should be able to modify it and resell it. My guess would be that individual datasets like this would not be worth trying to resell, and if they were it would likely be cheaper to generate them yourself from the sites. For example, taking a look at their most recent dataset (gamestop store locations) it looks to be a simple matter of: * Find the Gamestop store locator (http://web.sa.mapquest.com/gamestop/?tempset=search http://web.sa.mapquest.com/gamestop/?tempset=search, linked to on www.gamestop.com). * Plug in list of US Zip codes (I suppose you would need one of these for any scraping project of this type, might as well buy it some place). * Scrape the page that is returned. * They also include Longitude and Latitude. They must have a seperate database they run the address against to generate that, or it may be encoded into the Mapquest applet. All in all, for someone who is skilled in data scraping, and already has a scraping tool set up, its a matter of ~20 minutes worth of coding, followed by (maybe) an hour of scraping. Unless you feel your time is worth >$120 per hour, you would be better off scraping it yourself than purchasing it to resell.
- kbrower 17y agoYou can get a zip code database for free: http://www.populardata.com/zipcode_database.html http://www.populardata.com/zipcode_database.html
- frig 17y agoYes + no. NO: The kicker is that if you did that they'd sue you for breach of contract. YES: once you passed it off to third parties (in violation of contract) it's not clear how strong of a remedy AggData actually has. Some of the notions + folk wisdom about copyrightability of facts is based on assumptions that increasingly don't hold. EG: you can't really copyright the factual contents of the phonebook; if I want to compete with the existing phonebook publisher by retyping the phone book they don't have a copyright claim against me (provided I only transcribe the facts, and organize the information in an obvious or mechanical fashion, like alphabetical ordering). If you were instead to compete with an existing phone book company by literally xeroxing their product they could probably take you to trial (on a theory that the underlying facts aren't under copyright but the specific page layouts and so on are; there's also the issue of the ads you'd be xeroxing but let's not muddle things overmuch). If this has already happened and been litigated I've never heard of it. What something like AggData is doing shows some of the conceptual limits of the existing framework: - existing physical instantiations of abstract "databases" (collections of fact) -- (1) couldn't be economically "xeroxed" (EG: if you do it cheaply it is visibly inferior-looking; if you do a very high fidelity reproduction it's about as costly as just re-doing it from scratch) -- (2) had enough "wiggle room" in how they might be represented in a human-friendly medium such that: --- (A) on the one hand there's the possibility of a viable copyright claim against a "xeroxer" (under the theory that the page layout is under copyright even if the facts themselves are not) --- (B) on the other hand allowing for the possibility of "retypers" to actually take advantage of the not-copyrighted status of the underlying facts and actually produce a different product (b/c it is possible to reproduce the same facts with a format sufficiently-different from the source you drew them from) - but as the "database" becomes increasingly digital -- (1) "xeroxing" is very economical (far more so than "retyping") -- (2) there's increasingly less meaningful "wiggle room" as to how the facts might be represented in a "database"; changes-of-format dont' do much, but once the data is shorn of its human-friendly formatting all the useful ways of storing it are essentially isomorphic, meaning (2.B) above is increasingly unlikely (it may no longer be possible to "clone" the abstract data without being too close to the exact format of the source for legal protection). If someone's actually seen these issues played out or "settled" I'd love to learn more about it.
- 17y ago
- MrMatt 17y agoI don't think that something being a fact makes it exempt from copyright. Map data is copyrighted, and often has minor inaccuracies in order to indicate copyright infringement.
- frig 17y agoCF my longer post. Basically in the USA at least (EU has "database directive") the situation is that "facts" aren't themselves copyrightable but the intuitions / guidelines non-experts (like me) have to go from are very much pre-internet and pre-computers. In the map case: if I make a map of the coast of Florida by: - looking at maps A, B, and C - drawing my own map based on what I learned ...then the publishers of A (B or C) can't come sue me for violating their copyright over the shape of the coast of Florida; they have copyright over their specific depiction of that shape but not the shape itself. Once you move into all-digital datasets a lot of the grounding assumptions are no longer there (perfect reproduction is easy; the data is more abstract and may only really be representable in one way).
- evgen 17y agoThe facts themselves are not subject to copyright, although if some business decides it does not like how its TOS is being violated it would probably stuff a few "fake" entries in there that would be returned to IPs that trigger a spider warning and then later go back and see if they were put into the db; if so you go after aggdata for violating a copyright on the collection as a whole. The data being provided by the businesses that are being scraped is not much different than that provided by a map service provider or phone book provider -- the individual facts are not subject to copyright, but the collection as a whole is and if aggdata is basically doing the web equivalent of photocopying a phone book and selling it to you then they are open to attacks from this vector.
- cschneid 17y ago"The collection as a whole" isn't necessarily copyrightable as far as I know (not that I'm anywhere near a lawyer...). I think curated collections (say, of short stories) are copyrightable since it requires artistic decisions to be made, but a website like McDonalds that has a big list of locations is just a statement of fact, with no artistic basis. Maps have artwork that is protected, and phone books have page layout, font choice, etc. I don't think that textual addresses of a set of facts would be protected, but I think that it's a gray area of law that hasn't been decided.
- evgen 17y agoIt is not just the "artwork" of a map that is protected, or else you would see people turning online map databases into collections of lines and points. The entire collection, as a whole, can be subject to a compilation copyright. This is why mapmakers put in little bits of fake data, so that if their fake data turns up in your map they can nail you for copyright infringement.