3 ms·
There's a large amount of liability in disclosing your training data.
by ru552 2y ago
There's a large amount of liability in disclosing your training data.
- imjonse 2y agoCalling the model 'truly open' without is not technically correct though.
- Lacerda69 2y agoIt's open enough for all practical purposes IMO.
- nicklecompte 2y agoIt's not "open enough" to do an honest evaluation of these systems by constructing adversarial benchmarks.
- imjonse 2y agoAs open as an executable binary that you are allowed to download and use for free.
- bevekspldnw 2y agoIn the age of SaaS I’ll take it. It’s not like I have a few million dollars to pay for training even if I had all the code and data.
- londons_explore 2y agoI expect we'll see some AI companies in the future throwing away the training dataset. Maybe some have already. During a court case, the other side can demand discovery over your training dataset, for example to see if it contains a particular copyrighted work. But if you've already deleted the dataset, you're far more likely to win any case against you that hinges on what was in the dataset if the plaintiff can't even prove their work was included. And you can argue that the dataset was very expensive to store (which is true), and therefore deleted shortly after training was complete. You have no obligation to keep something for the benefit of potential future plaintiffs you aren't even aware of yet.