3 ms·
> But more importantly, under your definition, there will never exist in any form a useful open source set of weights. Because almost all data is proprietary. A
by guerrilla 2y ago
> But more importantly, under your definition, there will never exist in any form a useful open source set of weights. Because almost all data is proprietary. Anybody can train on large quantities of proprietary data without permission using fair use protections, but no matter what you can't redistribute it without permission. Any weights derived from training a model on data that can be redistributed by a single entity would inherently be so tiny that it would be almost useless. You could create a model with a few billion parameters that could memorize it all verbatim.
That may very well be so. We'll see what the future holds for us.