5 ms·
And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?
by RandomLensman 3y ago
And the cost of creating the information you want should be borne by whom? Even if it isn't monetizeable IP, how to share the costs?
- Swizec 3y ago> Even if it isn't monetizeable IP, how to share the costs? The internet started thanks to ample government funding for research. So have many other technologies, including AI. I wonder if there's a way we could all somehow pool our resources and use that to pay for common goods that we all use. What would we call such a scheme?
- RandomLensman 3y agoWe could definitely fund things. Or tax LLM firms in some broad way to redistribute to the respective societies for using their creations.
- dbmikus 3y ago> I wonder if there's a way we could all somehow pool our resources and use that to pay for common goods that we all use. What would we call such a scheme? Is this tongue in cheek? I think it's called the government and taxes! :)
- TeMPOraL 3y agoI don't have good answers. I have some high-level intuitions. One of them is that creation costs of information are fixed, while its usefulness is unbounded, so it doesn't make sense to try and reward creators for each access/view/use, in perpetuity. Secondly, there's a lot of information laundering going on - any random book I read carries between a few to few hundred references to prior written work. What I pay for the book goes to the author and the publishers, but AFAIK it doesn't go to any of the authors and publishers of works referenced in the book. Wikipedia takes this one step further, effectively turning all that information free. Thirdly, AFAIK copyright explicitly does not cover information/knowledge - it covers specific works. So Google showing me an info box with a recipe scrapped from some site could technically fall afoul of the law - but an LLM generating me a recipe based on associations created from being trained on millions of recipes, this feels like it should be in the clear, at least from user's POV.
- RandomLensman 3y agoI think that is a somewhat narrow view. Maybe to make the contrast sharper: Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? The new situation isn't the same as search as that wasn't there to hide information sources or to immediately convert information into useful things (texts, guides, etc.).
- TeMPOraL 3y ago> Why should I contribute any information just so that it immediately gets monetized by a handful of LLM firms? If this matters to you, then you shouldn't. But to flip this around: why should you care? Unless you're doing some unique work targeting a global audience, the point when LLM gets trained on what you created is way outside space you'd normally care about. Trying to capture all the value your work generates does not lead to a good world. Or maybe it's me who isn't profit-minded enough, but e.g. a lot of what I wrote on-line, including blog articles and commentary on Reddit and HN, has been used by search engines for free for a long time (over a decade, in some cases), and now is (most likely) part of the training corpora for LLMs. But I never believed, and still don't believe, that I'm entitled to some share of the gains LLMs (or search engines) make.
- RandomLensman 3y agoThis isn't so much about compensation, but why should I help enrich a large, even more direct rent seeker? Valuable information in a way is becoming more valuable for the LLM provider, so I would expect a drop in high value information in the public domain.
- TeMPOraL 3y agoPerhaps there will be a drop in high value information in the public domain, but right now, I can't exactly see LLMs impacting the incentives for creation and sharing of that information. I don't see how LLMs would make someone go "oh well, AI is here, I might as well stop providing people with no-strings-attached high quality information", if the existence of search engines didn't make them stop already.
- detourdog 3y agoI'm organizing the publishing of my thoughts of technology development and design. This has been something I have been mulling for decades. Completely uncompensated. I'm not doing any of this for reasons I can understand it is just what I think about and do. Originally I saw a website as a way to hang out a shingle. Until recently I was thinking I could just publish away and maybe someone would hire me based off the website. Currently I don't feel the same way regarding publishing on the web. I will be me more guarded in what I share.
- detourdog 3y agoI want attribution if I inspire a thought in AI. I'm surprised the nerve the issue has struck in me.
- martingalex2 3y agoIt isn’t clear to me (other than it’s an open engineering problem) why LLs couldn’t also include attribution as part of training. Also tracking attribution could lead to some insights on how its internal representations in vector space are created.
- detourdog 3y agoThat is true I see no reason obvious reason why the LL companies take pride in not being able to document ideation process. I have no justification but I feel it is deceitful not technical reasoning.
- TeMPOraL 3y agoThe issue here is that memorization of any distinguishable part of IP is an incidental aspect - those models aren't memorizing stuff, they're learning it. We don't expect people to keep track of the source of every single piece of information they encounter. It would arguably make learning impossible - as much for humans as for LLMs. As an intuition pump, when I write "2+2 = " and you mentally complete it with "4", should I chastise you for not completing it with "4, as per ${your elementary class math textbook} and ${that other book you read as a kid}, corroborated by ${your first math teacher} and ${your parent} quoting ${some other work}"?
- martingalex2 3y agoWhat is the hard technical barrier that makes the tracking of attribution for input sequences for LLM training impossible? I don't see any.
- TeMPOraL 3y ago