8 ms·
I don't think you get it. The game has already changed. And nobody trained the AI to write unit tests or refactor code. It could do it as emergent behavior. It
by byby 3y ago
I don't think you get it. The game has already changed. And nobody trained the AI to write unit tests or refactor code. It could do it as emergent behavior.
It happened so fast that so many people are in denial and many people aren't even asking the obvious questions. Simply ASK chatGPT to refactor your code to be more unit testable and IT Can do it.
Of course it clearly has your typical LLM problems but it is roughly 50% there. The game has changed. If AI gets to 100%, not only will the game be changed, but humans will no longer be part of the game.
See what I asked chatGPT:
ME:
def (x: int):
for i in range(x):
print(i)
Can you refactor the code so it can be more unit testable?
CHATGPT:
Certainly! To make the code more unit testable, we can refactor it into a function that returns a list of integers instead of directly printing them. Here's the refactored code:
python
def generate_number_list(x: int) -> List[int]:
number_list = []
for i in range(x):
number_list.append(i)
return number_list
In this refactored code, we've created a function called generate_number_list that takes an integer x as input. Instead of printing each number, we use a number_list variable to store the generated numbers by appending them in the loop. Finally, we return the number_list.
By returning the list instead of printing the numbers directly, we make the code more testable because we can now write unit tests that assert the correctness of the generated list.
If boggles my mind that I see threads of people saying "what if we train LLMs to do this or to do that?" or whatever. Guys, just ask the LLM to do what you want in plain English. It can do it.
- mjburgess 3y agoThat isn't the game. The game is designing software to requirements. It's writing literature for a new era. It's creating X for A audience with N vauge unspecified needs -- where X is a complex product made of many parts, involving many people, with shifting and changing problems/solutions/requirements. The game was never writing the stack overflow answer -- that was already written.
- byby 3y ago>The game was never writing the stack overflow answer -- that was already written. The problem is this was never a stackoverflow question and there was never an answer for it. Try finding it. The LLM is already playing the game because it came up with that answer which is Fully Correct, Out of Thin Air. Look, clearly the LLM can't play the game as well as a trained adept human, but it's definitely playing the game. >The game is designing software to requirements. It's writing literature for a new era. It's creating X for A audience with N vauge unspecified needs -- where X is a complex product made of many parts, involving many people, with shifting and changing problems/solutions/requirements. It can do all of this. It can talk like you and parrot exactly what your saying and also go into more detail and re-frame your words more eloquently. What you're not getting is that all the things you mentioned the LLM can do in actuality to varying degrees to the point where it is in the "game." and at times it does better than us. Likely, you haven't even tried asking it yet.
- mjburgess 3y ago> Fully Correct, Out of Thin Air I think if you're an expert in an area, this effect is easier to see through. You know where the github repo is, where the library example is, which ebooks there area -- etc. and you're mostly at-ease not using them and just writing the solution yourself. These systems are not "fully correct" and not "out of thin area". They are trained on everything ever digitised, including the entire internet. They, in effect, find similar historical cases to your query and merge them. In many cases, for specific enough queries, the text is verbatim from an original source. This is less revolutionary than the spreadsheet; it's less than google search. It's a speed boost to what was always the most wrote element to what we do. Yes, that often took us the longest -- and so some might be afraid that's what labour is -- but it isnt. We never "added value" to products via what may be automated. Value is always a matter of the desire of the buyer of the products of our labour (vs. the supply) -- and making those products for those buyers was always what they wanted. This will be clear to everyone pretty quickly, as with all tech, it's "magic" on the first encounter -- until the limitations are exposed. I actually work in an area where what took 3mo last year, I can now do in maybe 3 days due to ChatGPT. But when it comes to providing my customers with that content, the value was always in how I provided it and what it did for them. I think this makes my skills more valuable, not less. Since the quality of products will be even more stratified by experts who can quickly assemble what the customer needs from non-experts who have to fight through AI dialogue to get something generic.
- mchaver 3y agoI agree. LLMs are very impressive, but it isn't helpful to think of them of magic. LLMs are a great tool to explore and remix the body of human knowledge on the internet (limited to what it has been trained on). The user needs to keep in mind that it can give plenty of false information. To make good use of it, the user needs to be able to verify if the returned information is useful, makes sense, compare with first hand sources, etc. In the hands of expert that is really powerful. In the hands of a layman (on the subject in question), they can generate a lot of crap and misunderstand what it is saying. It is similar to the idea that Democracy can be a great tool, but it needs an educated and participatory populous or it may generate a lot of headaches.
- moffkalast 3y agoSo? Those requirements can be specified, holes inferred, and probably stuck to much more closely by a machine than man. If history's shown anything it's that if something takes a lot of mental effort for people it's probably an easy target for automation. The best developer is the one that doesn't get depressed when the requirements change for the 15th time in a month and just rewrites everything again at 2000x the speed of a human dev while costing basically nothing in comparison. People say, "oh but clients will have to get good at listing specs, that'll never happen". Like bruh the clients will obviously be using LLMs to make the specs too. Eventually the whole B2B workflow will just be LLMs talking to each other or something of the sort.
- Turskarama 3y agoYeah, it does great on little toy examples. What I would like to do is feed in my entire 50k line program and get something out.
- moffkalast 3y agoI really wonder how Claude 100k does on larger workspaces, has anyone tried that? (I don't feel like paying another $20 to Anthropic too) Allegedly it's only marginally better than 3.5-turbo on average so it'll probably spit out nonsensical code but maybe the huge context can help.
- cookieperson 3y agoOr you can just write the code yourself instead of praying something else can do your job for you better than you can...
- moffkalast 3y agoBut that would require me to not be a lazy ass, which we all know is impossible. Also writing code, in 2023? With your hands? Pshh, that's so 2010s.
- byby 3y agoSo I said it's like 50 percent of the way there implying that it gets things right at a rate of 50 percent. That's a fuzzy estimation as well, obviously so don't get pedantic on me with that number. When you ask for large output or give it large input you are increasing the sample size. Which means more likely that part of the answer are wrong. That's it. Simple statistic that are inline with my initial point. With AI we are roughly half way there at producing answers. If you keep the answers and questions short you will have a much higher probability of being correct. So that 50k line program? My claim is roughly 25k of those lines are usable. But that's a fuzzy claim because I LLMs can do much better than 25k. Maybe 75% is more realistic but I'll leave it at 50% so there's a lower bar for the nay sayers to attack.
- blibble 3y agoif it actually understood what it was doing it would tell you that that logic doesn't need a test as the python has the range(x) functionality built-in instead it generates a load of redundant boilerplate if I saw a developer check that in I'd think they were incompetent
- gjadi 3y agoThis. I'm not good at prompting (if I believe what others say they can do with ChatGPT), but that's one thing that bother me with this system. They will do anything you ask them to without questioning it (in the limit given by their creators). Is it possible to set it up in a way that they will challenge you instead of blindly doing what you ask? In this particular case, is it possible to ask it to do a code review in addition to performing the task? I've tried various time (with the v3.5) to "tune" it so that each answer will follow a specific format, with links and recommended resources, with several alternatives, etc. The goal is to have it to broad my perspectives as opposed to focus too much on what I'm asking. But it never worked for more than a couple of questions. Are there ways to do that? What am I doing wrong?
- pixl97 3y agoIn some cases you can ask about a particular function "Is this the best way to do this" or "Is there a better way to do this"
- byby 3y agoSort of. There's an input variable that adjusts the "creativity" of the LLM. If you adjust the variable the answers become more and more "creative" approaching the point where it can challenge you. But of course this comes at a cost. As it stands right now, chatGPT can actually challenge you.
- byby 3y agoI simply asked it to make it unit testable and it did the task 100 percent. I'm not sure where your side track is coming from. Who in their right mind would ever check in code that prints a range of numbers from 0 to x? The example wasn't about writing good code or realistic code. It's about an LLM knowing and understanding what I asked it to do. It did this by literally creating a correct answer that doesn't exist. Sorry it doesn't satisfy your code quality standards but that's not part of the task is it? Why don't you ask it to make the code quality better? It can likely do it Maybe that will stop the subtle insults (please don't subtly imply I'm incompetent that's fucking rude) Like why even get into code quality about some toy example? What's the objective? To fulfill some agenda against AI? I think that's literally a lot of what's going on in this thread.
- deleted 3y ago[deleted]
- Marazan 3y agoYou do get that that code is garbage and if any dev tried to check that code in they would get laughed out of the room right? It is pure pablum. It is almost the perfect example of LLMs producing fluent vapid bilge.
- byby 3y agoThe code is not garbage, it's just your highfalutin python opinion makes it so you only ever use list comprehensions or return generators. For loops in python that return non lazy evaluated lists are fine. Python was never suppose to be an efficient language anyways, grading python based off of this criteria is pointless. It doesn't matter how snobbish you are on language syntax though. I fed it code and regardless of whether you think it's garbage it did what I asked it to do and nothing else. Would you prefer the AI say, "this code is garbage, here's not only how to make it unit testable but how to improve your garbage code." Actually we can make the output more unpredictable as LLMs do have a non deterministic seed that can increase the creativity of the answer.
- Marazan 3y agoIt has wrapped range() with useless code. It has added no functionality, it has not improved testability in any way. . Please, take the code it has produced and integrate it into the original function. All it does is replace the range call. That's it. It has absolutely and totally failed at the given task whilst outputting plausible garbage about why it has succeeded.
- moomoo3000 3y agoIt removed the print and returns a list instead.
- Marazan 3y agoWhich is useless because _it has changed the semantics of the function_
- toolslive 3y ago> ME: > f(x) = exp(cos(x)) > > what is f(0) ? > > GPT-3.5: > f(0) = 1 (note, it most often generates the correct result, but I've seen it do the above too)
- glenneroo 3y agoThanks for at least admitting you used GPT 3.5, which is very out of date and hence no longer useful when discussing AI capabilities. If you want to test current tech (which is moving fast), at least use GPT-4 (which also gets updated regularly).
- toolslive 3y ago(latest greatest, a different function) > g(x) = sin(x)/x ; what is g(exp(-200)) ? > ChatGPT > > To find the value of g(x) = sin(x)/x at the point g(exp(-200)), we > substitute x = exp(-200) into the function: > g(exp(-200)) = sin(exp(-200))/exp(-200) > > Now, let's calculate this value using numerical methods: > > sin(exp(-200)) ≈ > 0.0000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000 > (there it breaks off, running out of tokens )
- byby 3y agoI mean. I said 50 percent of the way there right? So your example here is inline with what I said. It's not perfect, but it's halfway there.
- cookieperson 3y agoSo your current stance is, LLMs can't do everything yet, but don't bother thinking about extending it's capabilities just ask it it can do everything? Fascinating...
- byby 3y agoIt's not a stance. I'm stating a fact of reality. Huge difference. I didn't say don't bother extending it's capabilities either. You're just projecting that from your imagination. An hallucination so to speak not so far off from what LLMs do. I find your similarity to LLMs quite fascinating. What I said is, the capability of doing the "extension" you want is already in the LLM. Sure go extend it but what you're not getting is that we've already arrived at the destination.
- barrkel 3y agoThis is a fine, absolutely trivial, example. But LLMs are simply not all that. IME GPT-4 can't write a bug-free 10 line shell script. It's particularly poor at inferring unstated requirements - or the need to elicit the same. There's a general problem with LLMs: they're too eager to please. It shows up as confirmation bias. Embed a perspective in your prompt, and LLMs continue in the same vein. You can, with careful prompting, try to provoke and prod the text generation into a more correct shape, but often it feels to me more like a game than productivity. I have to know the answer already to know how to ask the right questions and make the right corrections. So it feels like I'm supervising a child, and that I should be amazed it can do anything at all. And it is amazing; but for productivity outside tightly constrained environments (e.g. converting freeform dialogue into filling out a bureaucratic form - I think this is a close to ideal use case), I struggle to see it scaling up much, from what I've seen so far. For creativity - e.g. making up a story for a child - it's not bad. One of my favourite use cases, after discovering how bad it is at writing code.
- byby 3y agoIf you read my post I said 50 percent of the way there. So that means a 10 line bash script at best would have 5 lines of bugs. But the AI is actually better then that. You've just increased the bar.