4 ms·
The Github they link to says it's under the CC BY-NC-SA 4.0.
by ekc 7y ago
The Github they link to says it's under the CC BY-NC-SA 4.0.
- bigiain 7y agoHmmm. I wonder how much fun lawyers will have arguing about whether that "NC" clause means a model trained on this data cannot be used commercially by the researcher who built it?
- parsimo2010 7y agoI would guess (I'm not a lawyer) that a commercial model would be in the clear as long as the company doesn't release anything including the data itself. Model weights that are derived from the data are not the data. I would make an analogy where the training data is like a textbook. If I read in a textbook about how to design/build a bridge, I don't have to give royalties to the textbook author when my civil engineering and construction firm gets paid to build a bridge. The copyright/license of the textbook can't prevent me from using the knowledge gained from the book to do a commercial job. In a similar vein, the knowledge gained from a public data set is probably fair game for whatever you want, just as long as you aren't repackaging the data itself. There's probably a good boundary requiring someone to stop at a point where the model weights can't be used to reconstruct the original data. Of course, other people can disagree. I would look forward to an actual legal opinion to clear this up.
- jahewson 7y agoI disagree. The license specifies that "using the material for commercial purposes" is prohibited. The act of training a commercial model is obviously a commercial purpose. Whether or not the data is somehow incorporated into the resulting model is irrelevant. You're confusing this CC license with open source licenses that do not restrict use but require derived works to be created/distributed under certain conditions. This CC license restricts use, in that you are not allowed to use the data for any commercial purpose - this is totally different from open source, and more like the "academic use only" licenses which used to be more common.
- parsimo2010 7y agoI think this is a pretty good argument, and more correct that my previous comment. It's too late for me to edit that one, but it's a pretty clear argument. I doubt that any company would be willing to take the risk of creating a commercial self-driving car using this data to get a head start. I don't think I'll be seeing this particular license fought over in court.
- yodon 7y agoCC is rarely a good license to choose other than to make yourself feel good. It does a terrible job of dealing with the actual matters of importance to a license outside of photo sharing (and even there it's not a good choice, as many photographers have found because it explicitly grants rights to others that the photographer may not be able to to grant to them, leaving the photographer open to lawsuits as has happened). In this case, imagine a student working on a class assignment. They use this data for purely academic purposes with no commercial intent in mind. After they train their system, they realize yow I could use this trained system and get rich. There was arguably no commercial use during the training. The use of the data was purely academic, like a person learning math or French. What you do after running the learning is a separate matter, just as using a CC licensed textbook to learn math doesn't prevent you from getting a job as a statistician. Again, the tl;dr is instead of trying to divine how a court will deal with a poorly specified problem, it's much better to just not license your stuff using a CC license. There are almost always much better licenses to choose from.
- bigiain 7y agoI'm curious about what you (or anyone else) would recommend as better license choices for datasets that might be used in machine learning model training? (With, I suppose, hints about what restrictions you might be wanting to grant or prohibit by particular license options?)
- yodon 7y agoIANAL but my sense is that it will be a number of years before we really know what works with regards to this kind of licensing of DNN/ML training data sets. It's almost certainly going to be decided on the basis of what's called "case law", which really just means "a bunch of random judges who don't know anything about ML made decisions on a bunch of random lawsuits that probably were not good samples to pick and which were probably taken to court by pairs of parties with wildly different abilities to pay lawyers and now we are stuck with those decisions as precedent for future cases." If that sounds like a crappy legal footing for the next 20 years of software development, yeah, it is. It's also why Mitch Kapor and a few others founded and funded the EFF to try to encourage better case law decisions around the early days of electronic privacy law. We definitely need an EFF like effort around ML/DNN/etc., but I'm not holding my breath. I wish I had a better answer, I'm mostly hoping someone else here does.