9 ms·
I call these "embarrassingly solved problems". There are plenty of examples of emulators on GitHub, therefore emulators exist in the latent spaces of LLMs. You
by zjp 8mo ago
I call these "embarrassingly solved problems". There are plenty of examples of emulators on GitHub, therefore emulators exist in the latent spaces of LLMs. You can have them spit one out whenever you want. It's embarrassingly solved.
There are no examples of what you tried to do.
- AuthAuth 8mo agoIts license washing. The code is great because its already a problem solved by someone else. The AI can spit out the solution with no license and no attribution and somehow its legal. I hope American tech legislation holds that same energy once others start taking American IP and spitting it back out with no license or attribution.
- irishcoffee 8mo agoThe models need to get burned down and retrained with these considerations baked in.
- blackqueeriroh 8mo agoNo. We need to light all IP law on fire. You shouldn’t able to license or patent software.
- reverius42 8mo agoWhat about novels? Nonfiction books? Scientific papers? Poems? Those things are all in the training data too.
- userbinator 8mo agoDo you give attribution to all the books, articles, etc. you've read? Everything is a derivative work.
- YeGoblynQueenne 8mo agoYou mean there are no new ideas? I think that's a big claim. As a for instance, how is mergesort "derivative work" of bubblesort?
- wredcoll 8mo agoNo but for a while we were required to pay amazon when we implemented a way to save payment details on a website.
- Guvante 8mo agoActually you might need to depending on how similar your implementation is. Copyright law here is quite nuanced. See the Google vs Oracle case about Java.
- ThunderSizzle 8mo agoI've seen many discussions stating patent hoarding has gone too far, and also that copyright for companies have gone way too far (even so much that Amazon can remove items from your purchase library if they lose their license to it). Then AI begins to offer a method around this over litigious system, and this becomes a core anti-AI argument. I do think it's silly to think public code (as in, code published to the public) won't be re-used by someone in a way your license dictates. I'd you didn't want that to happen, don't publish your code. Having said that, I do think there's a legitimate concern here.
- Retric 8mo agoA great deal of code on GitHub was not posted there by the original authors. So any argument that posting stuff online provides an implicit license is severely flawed.
- ibeckermayer 8mo ago1. Equality under the law is important in its own right. Even if a law is wrong, it isn’t right to allow particular corporations to flaunt it in a way that individuals would go to prison for. 2. GPL does not allow you to take the code, compress it in your latent space, and then sell that to consumers without open sourcing your code.
- creato 8mo ago> Even if a law is wrong, it isn’t right to allow particular corporations to flaunt it in a way that individuals would go to prison for. No one goes to prison for this. They might get sued, but even that is doubtful.
- ibeckermayer 8mo agoJust flat out false, and embarrassingly so, but spoken with the unearned authority of an LLM. See: The Pirate Bay.
- degamad 8mo agoAaron Swartz would probably disagree. https://en.wikipedia.org/wiki/Aaron_Swartz https://en.wikipedia.org/wiki/Aaron_Swartz
- candiddevmike 8mo agoIf I include licensed code in a prompt and have a LLM include it in the output, is it still licensed?
- phpnode 8mo agoThe other day I had an agent write a parser for a niche query language which I will not name. There are a few open source implementations of this language on github, but none of them are in my target language and none of them are PEGs. The agent wrote a near perfect implementation of this query language in a PEG. I know that it looked at the implementations that were on github, because I told it to, yet the result is nothing like them. It just used them as a reference. Would and should this be a licensing issue (if they weren't MIT)?
- Guvante 8mo agoNo one knows until a law about it is written. You could postulate based on judicial rulings but unless those are binding you are effectively hypothesizing.
- fsmv 8mo agoIt would be nice to give them some kind of attribution in the readme or something since you know which projects you referenced
- r-w 8mo agoExactly. If you have the decency to ask, you probably have the capacity to be courteous beyond the minimum required by law.
- phpnode 8mo agoI'm more interested in the general question rather than the specifics of this situation, which I'm sure is now incredibly common. I know it looked at those implementations because I asked it to, and therefore I will credit those projects when I release this library. In general though, people do not know what other material the agents looked at in order to derive their results, therefore they can't give credit, or even be sure that they are technically complying with the relevant licenses.
- tty456 8mo agoAt the end of the day it's up to the publisher of the work to attribute the sources that might end up in some commercial or public software derivative.
- _zagj 8mo ago> The AI can spit out the solution with no license and no attribution and somehow its legal Note that even MIT requires attribution.
- _zagj 8mo agoI'm not sure why this was downvoted. The MIT license, which many devs (and every LLM) treat as if it were public domain, still requires inclusion of the license and its copyright notice verbatim in derivative works.
- 20k 8mo agoThis is why its astonishing to me that AI has passed any legal department. I regularly see AI output large chunks of code that are 100% plagiarised from a project - its often not hard to find the original source by just looking up snippets of it. 100s of lines of code just completely stolen Ai doesn't actually wash licenses, it literally can't. Companies are just assuming they're above the law
- thedevilslawyer 8mo agoThis is oft-repeated but never backed up by evidence. Can you share the snippet that was plagiarized?
- vohk 8mo agoI can't offer an example of code, but considering researchers were able to cause models to reproduce literary works verbatim, it seems unlikely that a git repository would be materially different. https://www.theatlantic.com/technology/2026/01/ai-memorization-research/685552/ https://www.theatlantic.com/technology/2026/01/ai-memorizati...
- thedevilslawyer 8mo agoAssuming that even works from a researcher's perspective, it's working back from a specific goal. There's 0 actual instances (and I've been looking) where verbatim code has been spat out. It's a convenient criticism of LLMs, but a wrong one. We need to do better.
- latexr 8mo ago> There's 0 actual instances (and I've been looking) where verbatim code has been spat out. That’s not true. I’ve seen it happen and remember reports where it was obvious it happened (and trivial to verify) because the LLM reproduced the comments with source information. Either way, plagiarism doesn’t require one to copy 100% verbatim (otherwise every plagiarist would easily be off the hook). It still counts as plagiarism if you move a space or rename a variable. https://xcancel.com/DocSparse/status/1581461734665367554 https://xcancel.com/DocSparse/status/1581461734665367554 https://xcancel.com/mitsuhiko/status/1410886329924194309 https://xcancel.com/mitsuhiko/status/1410886329924194309 > We need to do better. I agree. We have to start by not dismissing valid criticisms by appealing to irrelevant technicalities which don’t excuse anything.
- palmotea 8mo ago> The AI can spit out the solution with no license and no attribution and somehow its legal. Has that been properly adjudicated? That's what the AI companies and their fans wish, but wishing for something doesn't make it true.
- sharperguy 8mo agoTo me, it's just further evidence that trying to assert ownership over a specific sequence of 1s and 0s is an entirely futile and meaningless endeavor.
- Hendrikto 8mo agoRegardless of your opinion on that (I largely agree with you), that is not the current law, and people went to prison for FAR less. Remember Aaron Swartz, for example.
- Andrex 8mo agoI did have the thought that the SCOTUS ruling against Oracle slightly opened the door to code not being copyrightable (they deliberately tap-danced around the issue). Maybe that's the future: all code is plumbing; no art, no creative intent.
- deleted 8mo ago[deleted]
- Nition 8mo agoIn a way it shows how poorly we have done over the years in general as programmers in making solved problems easily accessible instead of constantly reinventing the wheel. I don't know if AI is coming up with anything really novel (yet) but it's certainly a nice database of solved problems. I just hope we don't all start relying on current[1] AI so much that we lose the ability to solve novel problems ourselves. [1] (I say "current" AI because some new paradigm may well surpass us completely, but that's a whole different future to contemplate)
- BobbyJo 8mo ago> In a way it shows how poorly we have done over the years in general as programmers in making solved problems easily accessible instead of constantly reinventing the wheel. I just don't think there was a great way to make solved problems accessible before LLMs. I mean, these things were on github already, and still got reimplemented over and over again. Even high traffic libraries that solve some super common problem often have rough edges, or do something that breaks it for your specific use case. So even when the code is accessible, it doesn't always get used as much as it could. With LLMs, you can find it, learn it, and tailor it to your needs with one tool.
- kranner 8mo ago> I just don't think there was a great way to make solved problems accessible before LLMs. I mean, these things were on github already, and still got reimplemented over and over again. I'm not sure people wrote emulators, of all things, because they were trying to solve a problem in the commercial sense, or that they weren't aware of existing github projects and couldn't remember to search for them. It seems much more a labour of love kind of thing to work on. For something that holds that kind of appeal to you, you don't always want to take the shortcut. It's like solving a puzzle game by reading all the hints on the internet; you got through it but also ruined it for yourself.
- thaumasiotes 8mo ago> I just don't think there was a great way to make solved problems accessible before LLMs. I mean, these things were on github already, and still got reimplemented over and over again. What kranner said. There was never an accessibility problem for emulators. The reason there are a lot of emulators on github is that a lot of people wanted to write an emulator, not that a lot of people wanted to run an emulator and just couldn't find it.
- albert_e 8mo agoI tried writing a plain text wordle loop as a python exercise in loops and lists along with my kid. I saved the blank file as wordle.py to start the coding while explaining ideas. That was enough context for github copilot to suggest the entire `for` loop body after I just typed "for" Not much learning by doing happened in that instance. Before this `for` loop there were just two lines of code hardcoding some words ..that too were heavily autocompleted by copilot including string constants. ``` answer="cigar" guess="cigar" ```
- zjp 8mo agoI hate aggressive autocomplete like that. One thing to try would be using claude code in your directory but telling it that you want it to answer questions about design and direction when you get stuck, but otherwise never to touch the code itself, then in an editor that doesn't do that you can hack at the problem.
- cess11 8mo agoThis makes it really hard for juniors to learn, in my experience. When I pair with them I have them turn off that functionality so that we are forced to figure out the problems on our own and get to step through a few solutions that are gradually refined into something palatable.
- threethirtytwo 8mo agoStop repeating this trope. It can spit out something you've never built before this is utterly clear and demonstrated and no longer really up for debate. Claude code has never been built before claude code. Yet all of claude is being built by claude code. Why are people clinging to these useless trivial examples and using it to degrade AI? Like literally in front of our very eyes it can build things that aren't just "embarrassingly solved" I'm a SWE. I wish this stuff wasn't real. But it is. I'm not going off hype. I'm going what I do with AI day to day.
- zjp 8mo agoI think we are in violent agreement and I hope that after reading this you think so too. I don't disagree that LLMs can produce novel products, but let's decompose Claude Code into its subproblems. Since (IIRC) Claude Code's own author admits he built it entirely with Claude, I imagine the initial prompt was something like "I need a terminal based program that takes in user input, posts it to a webserver, and receives text responses from the webserver. On the backend, we're going to feed their input to a chatbot, which will determine what commands to run on that user's machine to get itself more context, and output code, so we need to take in strings (and they'll be pretty long ones), sanitize them, feed them to the chatbot, and send its response back over the wire." Everything here except the LLM has been done a thousand times before. It composed those building blocks in novel ways, that's what makes it so good. But I would argue that it's not going to generate new building blocks, and I really mean for my term to sit at the level of these subproblems, not at the level of a shipped product. I didn't mean to denigrate LLMs or minimize their usefulness in my original message, I just think my proposed term is a nice way to say "a problem that is so well represented in the training data that it is trivial for LLMs". And, if every subproblem is an embarrassingly solved problem, as in the case of an emulator, then the superproblem is also an ESP (but, for emulators, only for repeatedly emulated machines, like GameBoy -- A PS5 emulator is certainly not an ESP). Take this example: I wanted CC to add Flying Edges to my codebase. It knew where to integrate its solution. It adapted it to my codebase beautifully. But it didn't write Flying Edges because it fundamentally doesn't know what Flying Edges is. It wrote an implementation of Marching Cubes that was only shaped like Flying Edges. Novel algorithms aren't ESPs. I had to give it access to a copy of VTK's implementation (BSD license) for it to really get it, then it worked. Generating isosurfaces specifically with Flying Edges is not an ESP yet. But you could probably get Claude to one shot a toy graphics engine that displays Suzanne right now, so setting up a window, loading some gltf data, and displaying it definitely are ESPs.
- Aerroon 8mo ago>I call these "embarrassingly solved problems". When LLMs first appeared this was what I thought they were going to be useful for. We have open source software that's given away freely with no strings attached, but actually discovering and using it is hard. LLMs can help with that and I think that's pretty great. Leftpad wouldn't exist in an LLM world. (Or at least problems more complicated than leftpad, but still simple enough that an LLM could help wouldn't.)
- MichaelRo 8mo agoStrange that noone noticed the article saying "Nobody said 'Google did it for me' or 'it was the top result so it must be true.'" Because they did. They were the quintessential "Can I haz teh codez" Stack Overflow "programmer". Most of them, third world. Because that's where surviving tomorrow trumps everything today. Now, the "West" has caught up. Like they did with importing third world into everything. Which makes me optimistic. Only takes keeping composure a few more years until the house of cards disintegrates. Third world and our world is filled to the brim with people who would take any shortcut to avoid work. Shitting where they eat. Littering the streets, rivers, everywhere they live with crap that you throw out today because tomorrow it's another's problem. Welcome to third world in software engineering! Only it's not gonna last. Either will turn back to engineering or turn to third world as seemingly everything lately in the Western world. There's still hope though, not everybody is a woke indoctrinated imbecile.