3 ms·
The major problem is that the AI companies are redirecting the profit from artists to themselves. The creative industry will remain, but artists won't be able t
by PAMANOCH 4y ago
The major problem is that the AI companies are redirecting the profit from artists to themselves. The creative industry will remain, but artists won't be able to receive even one penny, companies could just 1) grab their work with absolute zero payment, 2) fine-tune a new model 3) profit on the style. Artist as a job will cease to exist very soon as they are becoming free suppliers for AI companies. Like, Uber with remote drivers controlling cars, but all drivers work for free as they claim all cars are capable of "self-driving". How is this acceptable and legal is beyond my understanding.
- throwaway290 4y agoThat's the exact concern. There is a way we can pit some corporations against the others to help with this though: Train models exclusively on content by megacorps like Disney, then claim it's fair use. Those guys lobbied to get copyright extended for ages for own profit; for once they could help protect the ordinary artist.
- ShamelessC 4y agoSimilar to Moore's law, "transformer" deep neural nets have been found to following scaling laws[0]. This means the faster and more VRAM your GPU's have, the better a model you can train "for free". Training models from scratch only works with massive (labeled) datasets covering a massive data distribution. With language models, the datasets being used are quickly approaching "all known written text" sizes. Training a model from scratch on Microsoft's internal code, with not only its precious intellectual property, but also its technical debt. Code at Microsoft is not going to get close to covering the broad range of styles that a coder could possibly use. The model will possibly diverge without enough data, as it needs to see a given "example usage" in multiple different contexts before it can learn it. My current understanding is that deep NN's are quite good at modeling an underlying distribution of data without needing any priors hard-coded about that dataset. But! They need to see a whole lot more of it than an adult human would. Several orders of magnitude more. - and they need to see accurate labels about 75-80% of the time. [0] https://www.lesswrong.com/tag/scaling-laws https://www.lesswrong.com/tag/scaling-laws