12 ms·
That's an interesting hypothesis : that LLM are fundamentally unable to produce original code. Do you have papers to back this up ? That was also my reaction w
by bsaul 10mo ago
That's an interesting hypothesis : that LLM are fundamentally unable to produce original code.
Do you have papers to back this up ? That was also my reaction when i saw some really crazy accurate comments on some vibe coded piece of code, but i couldn't prove it, and thinking about it now i think my intuition was wrong (ie : LLMs do produce original complex code).
- jacquesm 10mo agoWe can solve that question in an intuitive way: if human input is not what is driving the output then it would be sufficient to present it with a fraction of the current inputs, say everything up to 1970 and have it generate all of the input data from 1970 onwards as output. If that does not work then the moment you introduce AI you cap their capabilities unless humans continue to create original works to feed the AI. The conclusion - to me, at least - is that these pieces of software regurgitate their inputs, they are effectively whitewashing plagiarism, or, alternatively, their ability to generate new content is capped by some arbitrary limit relative to the inputs.
- andrepd 10mo agoExcellent observation.
- bfffbgfdcb 10mo ago[flagged]
- jacquesm 10mo agoI think my track record belies your very low value and frankly cowardly comment. If you have something to say at least do it under your real username instead of a throwaway.
- andsoitis 10mo agoI like your test. Should we also apply to specific humans? We all stand on the shoulders of giants and learn by looking at others’ solutions.
- jacquesm 10mo agoThat's true. But if we take your implied rebuttal then current level AI would be able to learn from current AI as well as it would learn from humans, just like humans learn from other humans. But so far that does not seem to be the case, in fact, AI companies do everything they can to avoid eating their own tail. They'd love eating their own tail if it was worth it. To me that's proof positive they know their output is mangled inputs, they need that originality otherwise they will sooner or later drown in nonsense and noise. It's essentially a very complex game of Chinese whispers.
- andsoitis 10mo agoI share that perspective.
- handoflixue 10mo agoEqually, of course, all six year olds need to be trained by other six year olds; we must stop this crutch of using adult teachers
- subscribed 10mo agoBeautiful, thank you.
- deleted 10mo ago[deleted]
- measurablefunc 10mo agoThis is known as the data processing inequality. Non-invertible functions can not create more information than what is available in their inputs: https://blog.blackhc.net/2023/08/sdpi_fsvi/ https://blog.blackhc.net/2023/08/sdpi_fsvi/. Whatever arithmetic operations are involved in laundering the inputs by stripping original sources & references can not lead to novelty that wasn't already available in some combination of the inputs. Neural networks can at best uncover latent correlations that were already available in the inputs. Expecting anything more is basically just wishful thinking.
- xyzzy123 10mo agoUsing this reasoning, would you argue that a new proof of a theorem adds no new information that was not present in the axioms, rules of inference and so on? If so, I'm not sure it's a useful framing. For novel writing, sure, I would not expect much truly interesting progress from LLMs without human input because fundamentally they are unable to have human experiences, and novels are a shadow or projection of that. But in math – and a lot of programming – the "world" is chiefly symbolic. The whole game is searching the space for new and useful arrangements. You don’t need to create new information in an information-theoretic sense for that. Even for the non-symbolic side (say diagnosing a network issue) of computing, AIs can interact with things almost as directly as we can by running commands so they are not fundamentally disadvantaged in terms of "closing the loop" with reality or conducting experiments.
- measurablefunc 10mo agoSound deductive rules of logic can not create novelty that exceeds the inherent limits of their foundational axiomatic assumptions. You can not expect novel results from neural networks that exceed the inherent information capacity of their training corpus & the inherent biases of the neural network (encoded by its architecture). So if the training corpus is semantically unsound & inconsistent then there is no reason to expect that it will produce logically sound & semantically coherent outputs (i.e. garbage inputs → garbage outputs).
- 10mo ago
- ninetyninenine 10mo ago[dead]
- fpoling 10mo agoPick up a book about programming from seventies or eighties that was unlikely to be scanned and feed into LLM. Take a task from it and ask LLM to write a program from it that even a student can solve within 10 minutes. If the problem was not really published before, LLM fails spectacularly.
- crawshaw 10mo agoThis does not appear to be true. Six months ago I created a small programming language. I had LLMs write hundreds of small programs in the language, using the parser, interpreter, and my spec as a guide for the language. The vast majority of these programs were either very close or exactly what I wanted. No prior source existed for the programming language because I created it whole cloth days earlier.
- jazzyjackson 10mo agoObviously you accidentally recreated a language from the 70s :P (I created a template language for JSON and added branching and conditionals and realized I had a whole programming language. Really proud of my originality until i was reading Ted Nelson's Computer Lib/Dream Machines and found out I reinvented TRAC, and to some extent, XSLT. Anyway LLMs are very good at reasoning about it because it can be constrained by a JSON schema. People who think LLMs only regurgitate haven't given it a fair shot)
- zahlman 10mo agoFWIW, I think a JSON-based XSLT-like thing sounds far more enjoyable to use than actual XSLT, so I'd encourage you to show it off.
- fpoling 10mo agoLanguages with reasonable semantics are rather similar and LLMs are good at detecting that and adapting from other languages.
- 10mo ago
- _heimdall 10mo agoI have a very anecdotal, but interesting, counterexample. I recently asked Gemini 3 Pro to create an RSS feed reader type of experience by using XSLT to style and layout an OPML file. I specifically wanted it to use a server-side proxy for CORS, pass through caching headers in the proxy to leverage standard HTTP caching, and I needed all feed entries for any feed in the OPML to be combined into a single chronological feed. It initially told multiple times that it wasn't possible (it also reminded me that Google is getting rid of XSLT). Regardless, after reiterating that it is possible multiple times it finally decided to make a temporary POC. That POC worked on the first try, with only one follow up to standardize date formatting with support for Atom and RSS. I obviously can't say the code was novel, though I would be a bit surprised if it trained on that task enough for it to remember roughly the full implementation and still claimed it was impossible.
- jacquesm 10mo agoWhy do you believe that to be a counterexample? In fragmentary form all of these elements must have been present in the input, the question is really how large the largest re-usable fragment was and whether or not barring some transformations you could trace it back to the original. I've done some experiments along the same lines to see what it spits out and what I noticed is that from example to example the programming style changed drastically, to the point that I suspect that it was mimicking even the style and not just the substance of the input data, and this over chunks of code long enough that it would definitely clear the bar for plagiarism.
- handoflixue 10mo ago> In fragmentary form all of these elements must have been present in the input Yes, and Shakespeare merely copied the existing 26 letters of the English alphabet. What magical process do you think students are using when they read and re-combine learned examples to solve assignments?
- jacquesm 10mo agoThis same argument has now been made a couple of times in this thread (in different guises) and does absolutely nothing to move the conversation forward. Words and letters are not copyrightable patterns in and of themselves. It is the composition of words and letters that we consider to be original creations and 'the bard' put them in a meaningful and original order not seen before, which established his reputation as a playwright.
- martin-t 10mo agoThe whole "reproduces training data vebatim" is a red herring. It reproduces _patterns from the training data_, sometimes including verbatim phrases. The work (to discover those patterns, to figure out what works and what does not, to debug some obscure heisenbug and write a blog post about it, ...) was done by humans. Those humans should be compensated for their work, not owners of mega-corporations who found a loophole in copyright.
- moron4hire 10mo agoNo, the thing needing proof is the novel idea: that LLMs can produce original code.
- marcus_holmes 10mo agoLLMs can definitely produce original other stuff: ask it to create an original poem and on an extremely specific niche subject and it will do so. You can specify the niche subject to the point where it is incredibly unlikely that there is a poem on that subject in its training data, and it will still produce an original poem on that subject [0]. The well-known "otter using wifi on a plane" series of images [1] is another example: this is not in the training data (well, it is now, because well-known, but you get the idea). Is there something unique about code, that is different from language (or images), that would make it impossible for an LLM to produce original code? I don't believe so, but I'm willing to be convinced. I think this switches the burden of proof: we know LLMs can produce original content in other contexts. Why would they not be able to create original code? [0] Ever curious, I tested this assumption. I got Claude to write an original limerick about goats oiling their beards with olive oil, which was the first reasonable thing I could think of as a suitably niche subject. I googled the result and could not find anything close to it. I then asked it to produce another limerick on the same subject, and it produced a different limerick, so obviously not just repeating training data. [1] https://www.oneusefulthing.org/p/the-recent-history-of-ai-in-32-otters https://www.oneusefulthing.org/p/the-recent-history-of-ai-in...
- jacquesm 10mo agoNo, it transformed your prompt. Another person giving it the same prompt will get the same result when starting from the same state. f('your prompt here') is a transformation of your prompt based on hidden state.
- marcus_holmes 10mo agoThis is also true of humans, see every debate on free will ever. The trick, of course, is getting to the exact same starting state.
- checker659 10mo agoI think the burden of proof is on the people making the original claim (that LLMs are indeed spitting out original code).
- checkmatez 10mo ago> that LLM are fundamentally unable to produce original code. What about humans? Are humans capable of producing completely original code or ideas or thoughts? As the saying goes, if you want to create something from scratch, you have to start by inventing the universe. Human mind works by noticing patterns and applying them in different contexts.