3 ms·
making sure that a dataset is clean and not full of material that's improperly sourced, copyrighted, unfit for use due to licensing or ethics, is not nearly har
by pxoe 3y ago
making sure that a dataset is clean and not full of material that's improperly sourced, copyrighted, unfit for use due to licensing or ethics, is not nearly hard enough nor "impossible" for it to be a situation where people should just "give up".
and yes, while open source models might be harder to regulate, those big corporations that currently use those things without distinction, exist as pretty established entities, and profit from services they offer in millions of dollars. there's more of a substantial existence, and more of a substantial scale of money they actually move. and they don't just "make a tool available", or have users do unambiguous actions where it would be the users that are infringing on anything, but do indeed use questionably sourced data and turn that into a model and offer that as a service. dirty data is very much a part of the deal with those.