3 ms·
I find the first argument, that if you're project is in GitHub then they have the right to train in it, weak. Plenty of projects are hosted elsewhere but have b
by SethTro 5y ago
I find the first argument, that if you're project is in GitHub then they have the right to train in it, weak. Plenty of projects are hosted elsewhere but have been mirrored by random users (e.g. not the copyright holder) to GitHub
- kbenson 5y agoI think they have a right to train in it, but not to present portions verbatim. Do you have a right to look at a bunch of open source code and come to conclusions about good programming practices? Are you prevented from knowing that a specific library in a language is good/common for a specific task because you see others using it? That's analogous to training, where there are associations between things, in my mind. I don't think that means they can provide licensed code verbatim though, just as you should not copy GPL code directly out of a Github repo and paste into your own private commercial code base.
- heavyset_go 5y agoYou're taking the machine "learning" metaphor literally. A human being learning something is not analogous to training an ML model. Training models is more analogous to compilation or lossy encoding or compression.
- michaelpb 5y agoThe biggest mistake of the ML field is its metaphorical naming. So many people seem to be taking Artificial Intelligence, Machine Learning, Neural Networks etc literally. They don't do this for other concepts in coding (eg for an absurd example, no one is arguing we ride a CPU "bus" to work), but with ML algos its a free-for-all. Grandiose naming conventions might be good for extracting VC money but it's also seriously confusing people.
- kbenson 5y agoI'm thinking more "association" than "learning", and in both cases. If an algorithm of some sort scans a bunch of repos regarding video encoding and decoding and sees a lot of ffmpeg use, it might associate ffmpeg with video encoding and decoding, and decide to present some info about ffmpeg and a generic snippet to include ffmpeg as a library and initialize it if it associates the current project with that. If I have perused a few encoding or decoding repos at some point and I think of the current project as having to do with encoding or decoding of video, I might immediately think ffmpeg even if I've never used it in a project as a library because I remembered seeing it in projects that used it, and look for some initialization code. In what ways are these materially different? What makes the random conceptual associations in my head from what I've seen previously different than an algorithm that collects the same? > Training models is more analogous to compilation or lossy encoding or compression. And learning in people isn't? Isn't all knowledge transference in people analogous to lossy encoding and compression? I don't know about you, but in college I don't remember regurgitating sections of "Advanced Programming in the UNIX Environment" to complete assignments, I remember studying it, internalizing parts of it on a conceptual level (as well as remembering specific fairly small chunks almost exactly), and using that to solve problems or answer questions or make associations. I'm not saying ML and and learning in humans is the same. I do think for the very specific case presented here in how it's used, there are some parallels. Feel free to disabuse me of that notion if you have evidence that contradicts it though. I'm not wedded to that position, but I would want to see arguments to the contrary before abandoning it.
- lucideer 5y agoThere's a general false assumption at the very beginning of this article that users own the code they upload to Github (I can freely & legally upload some code I forked from a FOSS project elsewhere). Pretty concerning that this fairly obvious oversight was missed by an IP lawyer consulted by a platform apparently dedicated to managing open-source code. Github's own ToS makes explicit & careful distinction between code you upload and code you own.