5 ms·
This is so backwards. If you limit scraping and fair use, it further entrenches the power of the tech and content giants. If you're upset about big tech's pow
by celestialcheese 4y ago
This is so backwards.
If you limit scraping and fair use, it further entrenches the power of the tech and content giants.
If you're upset about big tech's power and influence, you should be cheering on the free use of publicly available information, as it has the best shot at unseating that power.
- rout39574 4y agoI'm not sure it's "unseating". Moving from one lobe of enclosed corporate power to another doesn't help. I think this serves more to enclose that which was free, than to liberate what was closed.
- celestialcheese 4y ago>I think this serves more to enclose that which was free, than to liberate what was closed. What has been closed? Every day there's new foundational models being released. It's exhausting trying to keep up with the pace of change - Alpaca, Vicuna, LLaMA, etc. I'd be shocked if there wasn't a truly open source foundational model with performance equivalent to OpenAI's GPT-4 by the end of the year. Any move to limit fair-use and scraping of publicly available information through copyright laws, no matter how good of a reason it is, gives more power to the biggest companies.
- digging 4y ago> If you limit scraping and fair use, it further entrenches the power of the tech and content giants. Explain?
- celestialcheese 4y agoFair use has been a thorn in the side of content and publishing giants for a long time as it allows for upstarts to create derivative works and databases without violating the original copyright. If you block fair-use and scraping, those that hold the most data and copyrights (the largest companies) will be the only ones able to create quality foundational models, thus holding the keys and further entrenching their power.
- Angostura 4y agoIf I'm a person and I limit the way you scrape my private blog, It's not clear how I'm entrenching anyone
- zdragnar 4y agoIf it's on the internet and not behind a paywall, it's not private.
- orev 4y agoIn the US, anything you create is automatically copyrighted and you have full rights to decide who does what with it, unless you explicitly waive those rights, even if you post it publicly on the Internet. Too many people seem to think that just because something’s public it means they’re allowed to do whatever they want with it. That’s incorrect. The only barrier is whether someone you copy from is willing to bring legal action.
- zdragnar 4y agoThere are a number of exceptions to copyright, and at least in the US, the supreme court's understanding of "transformativeness" applies to chatgpt and other similar tech. The relevant quote, originally written in 1990 and cited in 94, is thus: " [If] the secondary use adds value to the original--if the quoted matter is used as raw material, transformed in the creation of new information, new aesthetics, new insights and understandings--this is the very type of activity that the fair use doctrine intends to protect for the enrichment of society. " We've wandered into uncharted legal territory with only a few light posts guiding the way. Nothing about this is settled or obvious.
- deleted 4y ago[deleted]
- olalonde 4y agoAFAIK, using copyrighted works for training machine learning models does not infringe on copyrights and falls under fair use. https://www.thefashionlaw.com/ai-trained-on-copyrighted-works-when-is-it-fair-use/ https://www.thefashionlaw.com/ai-trained-on-copyrighted-work...
- meroes 4y agoDidn’t Google become entrenched by scraping? Scraping seems to create a few dozen winners and 99% losers from the data we have.
- celestialcheese 4y agoIf anything, it's the opposite. It's the legal restrictions around scraping that have entrenched the power of the walled-gardens today. They built their businesses on scraping, then when they had their lock-in and monopolies, they turned around and fought against scraping to keep up-starts from eroding their business. This blog is a great overview of the current landscape around scraping laws. https://blog.ericgoldman.org/archives/2022/12/hello-youve-been-referred-here-because-youre-wrong-about-web-scraping-laws-guest-blog-post-part-2-of-2.htm https://blog.ericgoldman.org/archives/2022/12/hello-youve-be...
- basisword 4y agoIt doesn’t have to be all or nothing. You can regulate the big companies and leave the smaller companies unregulated. I’m not sure what regulations, if any, are necessary here but pausing to think about it when the consequences are potentially world changing seems like a good idea.
- raverbashing 4y agoThe problem is not what was scraped, but add conversation data to the training set (without obvious disclosure) Samsung got bit by it https://twitter.com/GergelyOrosz/status/1643903974536347649 https://twitter.com/GergelyOrosz/status/1643903974536347649