3 ms·
You mentioned the corpus prediction being the core of the LLM. Because of this, my prediction is we will see way more data withholding to prevent LLM learning j
by mrcode007 3y ago
You mentioned the corpus prediction being the core of the LLM. Because of this, my prediction is we will see way more data withholding to prevent LLM learning just like we’ve seen with stack overflow, Reddit, X.
I myself have started doing this. For example, I don’t publish code on GitHub anymore to prevent copilot training on my own code. Normally I like to get paid for work, instead of paying for GitHub and doing work for them:) I’d much rather take a page out of an artists playbook and require royalties.
I have also built a coding test harness to test for AGI and let me tell you, that the LLMs are completely unable to reason on the most intermediate and advanced algorithmic problems you present to them the moment they are not in the corpus. This involves a nearly complete lack of the ability to generalize and infer new concepts needed to solve the problem even if the concept is explained in the problem description.
Now the issue is that it is impossible to present the test because once it ends up in the corpus, then someone will put the solution in the corpus as well and claim the ability to reason. You can believe me or not; I won’t be making my test available either way, but it is not too hard to make one of your own to convince yourself.
I wish it was possible to make the tests like that available but alas we will likely end up with a medical-style double blind kind of tests to get rid of AI FUD.
- visarga 3y agoYou are not fundamentally wrong, but there are degrees. LLMs can do some symbolic manipulation and a few reasoning steps. Maybe they are just the result of pattern matching, but I have a hunch that most of the time humans do the same. We fake our understanding as much as we can, we use all shortcuts.
- vidarh 3y agoWithholding like that won't achieve anything other than your own obscurity.
- LikelyABurner 3y agoDon’t kid yourself, posting your code on GitHub won’t lead to your being “discovered”, either. Online portfolios are the worst lie we tell young developers. It’s the software industry’s version of “exposure” gigs: they get a few more bits they can shove into their models, you get jack shit. Every, and I mean, every job I got I got because I either knew a guy or I knew a guy that knew a guy. You are far better off going to conferences and networking.
- vidarh 3y agoI have a 30 year career in this business, and I know full well what my Github and other online exposure has done for my career, and my experience does not mirror yours. It has mattered. It's not "made" my career, but it's smoothed paths and gotten me respect from people I'd otherwise not have gotten anything out of. I do accept it won't matter for most people. But that was not my point. My assumption to start with is that most of us do not put much on Github that more than a few other people will care about. But most individual accounts are also of minimal value to Github or other people training LLMs. So however small recognition you get from having code accessible publicly (for me it's the occasional recognition when doing job interviews and a few e-mails from people now and again), you lose that to take away far less value to people training LLMs. For the vast majority of us, it takes a vastly inflated ego to think withholding our code from a public repository will even be noticed by more than a handful of people. There are exceptions, to be sure, but they are vanishingly small proportion of us.