46 ms·
Meta's new LLM-based test generator
- holoduke 3y agoI am using copilots now since a few months and it really makes me a 2x more productive developer. Its like you become an orchestrator of a dev team. You still need to look into details, but things just flow much much faster. I can only imagine how it will be like if I also have access to a AI debugger / end to end tester. Completing the loop and making it super efficient. Also think that this is not only the case for developers. Even for lawyer it could be the same thing. Expected production output of a worker is going to rise. The ones who do not embrace AI assistants in the near future will have a hard time in the future.
- zdragnar 3y ago> Even for lawyer it could be the same thing Funny you mention that, since lawyers have already gotten in trouble for citing fictional cases when submitting work performed by chatgpt. It's useful for rote things, but for anything that you depend on, you still need to give it just as much attention as if you'd done it yourself.
- qup 3y agoI agree, but it does change the shape of the work you would be doing. Validating work isn't the same as generating new work. Whether that's better, or useful, probably differs by situation.
- blibble 3y ago> Validating work isn't the same as generating new work. in many cases it's way harder
- cloverich 3y ago> you still need to give it just as much attention as if you'd done it yourself. It's different than a Lawyers case, where the facts require manually cross referencing. In our case, that verification can come via directly and immediately running the code. How good the actual code is varies, but the fundamental difference remains.
- skissane 3y ago> Funny you mention that, since lawyers have already gotten in trouble for citing fictional cases when submitting work performed by chatgpt. For something like legal work, you don’t want a raw LLM. You want an agent integrated with a legal database, in such a way that it can’t cite cases which don’t exist in the database. And it can’t generate a direct quote from a case unless that text actually occurs in the case. You still need a human lawyer familiar with the law to pick up on subtler errors, but better technology can prevent the grosser ones. And the subtler errors (e.g. misrepresenting what a case says through partial or out of context quoting) is the kind of error human lawyers sometimes make too - like every other profession, lawyers vary greatly in their competence, and often only the grosser cases of incompetence incur sanctions
- zmgsabst 3y agoTried Lexis+ AI, and… it’s just not very good yet. Much like ChatGPT, it can only handle recall and summation — and even then, I don’t fully trust it because it often misses key ideas. And much like ChatGPT, it can’t do anything coherent at length without a lot of help and working around its faults. And seems entirely unaware of similar sounding words having distinct legal meanings. Which is not good.
- skissane 3y ago> Tried Lexis+ AI, and… it’s just not very good yet. I wonder to what extent that’s due to current inherent limitations of the technology, and to what extent it is due to quality of implementation issues. It is hard to say because (I presume) there is only limited public information on how it is actually implemented. e.g. which LLM is it using? How much fine-tuning has been done? Are they using other potentially helpful techniques such as guided sampling? Or breaking down the task into parts and having multiple agents each specialised to handle one particular part? > Much like ChatGPT When you compare it to ChatGPT, do you mean GPT3.5 or GPT4? I also would guess that certain areas of law (especially criminal law) may be prone to triggering “safeguards” which result in poorer performance than if those safeguards were absent. Arguing that what your client did was legal (even be it ethically unsavoury) is an essential part of a lawyer’s job
- unshavedyak 3y agoI still haven't figured out how people use this that much. I use LLMs almost daily via ChatGPT, Phind, Kagi Ultimate (my current one), etc. However i spend too much time pushing the LLM towards my goal that i can't imagine it speeding my coding up. I clearly find value in LLMs to some degree, but speeding up my coding is not yet one.. i'd love to, but i just don't understand. I can only imagine typing more in explanation for the LLM than it would take me to write it to begin with.
- oblio 3y agoIMO a chunk of it is like Peel developers: people that prioritize their own speed above all else: readability, team reviews, etc. I'd love to be proven wrong.
- 010101010101 3y agoI rarely find interfacing with an external chat interface useful, but integration with the coding environment (e.g. Copilot) is an immediate productivity boost.
- swatcoder 3y agoSome revered novelists spend all day to write a single page and others are able to get in a flow and produce a chapter in that same time. Without getting into touchy questions of quality and talent and the whole 10x trope, there are a lot of working engineers out there that produce code a lot more slowly than others and that work on more common problems than others. My sense is copilot-like products provide the biggest boon to those people and are a lot harder for people who are more naturally prolific or who work on more esoteric things.
- holoduke 3y agoHave you ever tried it?. It sounds like you are a bit against the idea of an AI supporting your work.
- freedomben 3y agoIt all seems to depend on what you are doing. The more niche and technical the task, the less the AI can help. If you are just generating crud and points for a common web language and framework, it can be an 80% boost. I think the real problem with this is that people aren't differentiating these different types of work when they give these numbers.
- rglover 3y ago> The ones who do not embrace AI assistants in the near future will have a hard time in the future. The exact opposite will be true and the funny (sad?) part is that they will lack the skills necessary (because they got lazy and over-trusted the AI) to fix mistakes/incompatible solutions.
- fhd2 3y agoMy thinking as well. If coding assistants become so good that I'm at a competitive disadvantage, I'll just start using them. It's not rocket science. So far, everything I've tried largely slowed me down. As a Google/SO replacement for some types of questions, they sure save me maybe an hour per week, but that's really all I could extract so far. Maybe my work is not too typical though, I spend only a fraction of my time actually typing in code. And I do eliminate the need for boilerplate through other means (picking frameworks/libraries that are a good fit for the problem, refactoring, meta programming, scripts, suitable tool chain etc).
- zmgsabst 3y ago1 hour per week of increased coding is a 5-7% boost in productivity, using the Amazon guidelines for how SDEs use their time — 50%/20hrs for SDE1 and 33%/13hrs for SDE3. Is that enough to be competitive? I’m not sure — but at scale that would be a 5% reduction in headcount for the same work, or ~$12M/yr for every 1,000 engineers. If you can figure out how to get 2-3 hours more coding done a week, we’re talking real gains.
- fhd2 3y agoDepends on where the time is saved. If it's in figuring out how to do something in Django where StackOverflow is flooded with outdated answers, sure. I see those kinds of savings. But the tragic beauty of programming is that a little time saved today can very well mean lots of time lost later. The former you can measure, the latter is a tougher nut.
- skwirl 3y agoPeople said the same thing about garbage collectors.
- taude 3y agoin the future, it seems like we might just become PR reviewers
- rco8786 3y agoI’ve had it enabled for months across both Javascript and Kotlin codebases and it’s…fine? Good enough that I leave it enabled. But only barely. I’m certainly not orchestrating a dev team. It has probably the same productivity boost that intellisense gave back when it came out. Which is good, but still marginal. Certainly not replacing anyone’s job.
- peter_l_downs 3y ago[flagged]
- dvaun 3y ago[flagged]
- engineercodex 3y agoI inserted these because I personally like reading related discussions and articles of topics at hand. Not sure how this is a negative :/
- dvaun 3y agoYou’re right. That was undeserved, my apologies. Edit: I’d like to note that your writing was fun to read—my comment was instead leaking a bad mood I had at the time.
- engineercodex 3y agoOuch - half of the article (the "Actionable Takeaways" section) was my own commentary. The summary was for those who didn't want to necessarily parse through the entire paper's PDF. Happy to listen to any constructive feedback if you have any, though!
- samsk 3y agoFinally some AI Codegen, that makes sense to me.
- ShamelessC 3y agoFinally?
- refulgentis 3y agoThere's a persistent rather large minority that has a nuanced take: it can't write code they like (don't want to edit), but it's great for weekend projects (where they're trying new things without established personal preferences). Forest for the trees if you ask me, but, to each their own.
- skissane 3y agoSometimes, I find writing pseudocode easier than code. And then I ask an AI to turn it into code. Sometimes the results aren’t too bad, and just need a few tweaks for me to use it-overall I’ve saved mental effort compared to translating the pseudocode into code by hand. And if the results aren’t useful, I’ve only wasted a few seconds, and then I just have to do it manually.
- huytersd 3y ago[flagged]
- plufz 3y agoI don’t know if I’m slow but I try to use ML for dev work but I have a hard time using it to be productive. Maybe it is because I don’t want to use GPT and copilot but I rather use local codellama, mistral, dolphin, etc. But I have found very limited use for it. I use ML for other things like transcriptions with whisper, translations with m2m100 and summerizing articles with custom python scripts and different models. Not even the simplest things like convert this Apache conf to nginx works in a satisfactory way. It can’t solve bugs. It can’t help me with like edge cases for weird libraries etc. It’s like it only can help me with very obvious things that is quicker to just write myself. How do you use it?
- LASR 3y agoEveryone by now should be writing unit tests using ChatGPT4. I paste in functions / classes I want to write unit tests for. Paste in a sample unit test, and it does a solid job of writing tests for it in the same manner as in the sample. For unit tests, you don't even need the multi-step coverage optimization in this article. You just manually inspect, adjust it etc.
- swatcoder 3y agoIf I have a choice between a clever, practiced antagonist to plan and write my tests and an automated system that can fit common testing patterns to my code... I'm always going to get more robust results from the former. But yeah, if you're just working solo on basic stuff and need to protect against off-by-one errors and accidental mutations in later refactoring, it's a great tool. You'd never write duly antagonistic tests for your own code anyway.
- interroboink 3y agoWhat about people who don't trust sending their code to a 3rd party for processing? (edit: I didn't downvote you, but I do think your claim is over-broad)
- deleted 3y ago[deleted]
- biot 3y agoI suspect a lot of people overestimate how special their code is. Also, if your code exists in a private repo on GitHub, then you're already trusting the same third party when using GitHub Copilot.
- dns_snek 3y agoThis isn't just about IP. In the not so distant future we'll find out about 3 letter agencies using companies like OpenAI to deploy corporate backdoors with ease.
- 3y ago
- nicklecompte 3y agoI don't want to review this whole thing but one part in particular seems way off. [Caveat: I sorta-read the original paper shortly after it was posted, my memory is fuzzy and I am only skimming it now.] From the blog: > Most of the test cases created by Meta’s TestGen-LLM only covered an extra 2.5 lines. However, one test case covered 1326 lines! The value of that one test case is exponentially more valuable than most of the previous test cases and exponentially improves the value of TestGen-LLM. LLMs can vigorously “think outside the box” and the value of catching unexpected edge cases is very high here. Of course "exponentially more valuable" should set off your BS detector. But to verify, from the paper: > However, this result arose due to a single test case, which achieved 1,326 lines covered. This test case managed to ‘hit the jackpot’ in terms of unit test coverage for a single test case. Because TestGen-LLM is typically adding to the existing coverage, and seeking to cover corner cases, the typical expected number of lines of code covered per test case is much lower....The median number of lines of code added by a TestGen-LLM test in the test-a-thon was 2.5. This is a more realistic assessment of the expected additional line coverage from a single test generated by TestGen-LLM. Nowhere do the authors mention "unexpected edge cases" or "thinking outside the box." They clearly present this 1,326 lines of coverage test as a fluke, e.g. maybe the test case checked one branch of a horrible switch statement, or perhaps it was even a fluke in how code coverage is counted. It is noteworthy that the authors do not seem to have looked into it any further, even in the "qualitative results" section. Inaccurate editorializing really doesn't help anyone. The internet is too damn full of people pretending to understand things they pretended to read.
- engineercodex 3y agoHey! Thanks for your comment - I'm the one who wrote this article. I wasn't trying to say that the paper authors talked about "unexpected edge cases" or "thinking outside the box." I edited the post to be more clear that some of these takeaways are my own opinions. This article is less of a summary of a paper and rather commentary on what the results of the paper entails. After all, Hacker News is meant for discussion :) I will say though that I do believe that I still stand by the "exponentially more valuable" portion. I think the fact that LLMs can fluke their way into "hitting a jackpot" in terms of test coverage is exactly why they're so valuable. When you have something constantly trying out different combinations, if it hits even one jackpot, like in the paper, it's extremely valuable to the team. It's a case that could have been either non-obvious or simply too tedious to write a test for manually. I think there's tremendous value in that, especially speaking as someone who has spend way too much time simply figuring out how to test something within a Big Tech codebase (F/G) when I already knew what to test.
- romwell 3y agoYeah, after working in semiconductor industry (computational lithography) where test-driven design is the norm... I'm not convinced. I'm not saying that writing tests before the production code is something that should always be done. But tests are just as much a part of the codebase as anything else, and absolutely must be written alongside the code being tested. The most important part of the test is that it showcases intent of the developer. A test suite demonstrates the following: * How the code should be used * What the code does * What the code doesn't do * What it was written for Then when that code is used or modified by another developer, they don't have to hunt for clues in the codebase like they're Sherlock Holmes. If the tests aren't telling a story, you're writing tests wrong. And until the computers gain the ability to read your mind and do a better job at understanding what you want to do, AI/LLM-based generators can't do this job for you. Of course, if the only goal of your test suite is getting a green checkmark on a pre-commit check (and being able to show great coverage numbers), then yeah, you can double your productivity with AI. Automatic code generators will surely help you write more bad code at lightning speed. And if others complain that tons of boilerplate make the code bloated and hard to understand — just tell them to use AI to deal with it. Worked for you! That really does seem to be the future of development. But not the future I'm looking forward to.
- azeirah 3y agoI agree with almost everything you said, although I do think this type of testing has a place. There are different types of testing, what you're describing sounds to me like testing the "core" of your code, part documentation, part validation, part stability, etc. Other types of testing like fuzzing provide an entirely kind of value. I believe this AI- driven testing can inherit a space to target tests at the tail end of the distribution, many tests with little value. Providing extra coverage where human energy and time is lacking. That is how I see the current state of AI tooling regardless, as a cognitive assistant. I'd be surprised if this line of research doesn't end up being very fruitful in the coming years.
- romwell 3y agoThat I can fully agree with (particularly, comparison with fuzzing). Your comment presents a way more grounded perspective on the future of LLMs in programming than the article does.
- gxt 3y agoElementary tests, like unit tests, should be mecanically generated by walking the AST, differences ack`d and snapshoted when commiting. Every language should come with this built-in.
- superb_dev 3y agoWhat exactly are we testing at that point?
- Groxx 3y agoEnsuring that `if x == 1` works when x == 1. Very important. Very valuable. Imagine if `if err != nil { return err }` just stopped working tomorrow. Your tests would detect it! Outage prevented!
- Cthulhu_ 3y agoYou're being sarcastic but honestly, I've never found a regression because of a unit test. Only past few days though, I did find two bugs that would've been prevented if the original code was covered by a decent unit test.
- sangnoir 3y ago> You're being sarcastic but honestly, I've never found a regression because of a unit test So you've never made a change caused a unit test to fail? If not, how large is your codebase, and is ownership shared across multiple teams? I caught dozens of latent or unreported bugs by writing unit tests for a 6kloc JS app which had 0% coverage before.
- cgdub 3y agoI don't write Perl or Ruby anymore, but this would have been immensely helpful back then.
- bluefishinit 3y agoThis is called compiling with a type system.
- kissgyorgy 3y agoWhat future? LLMs got into our tech stack faster than a JavaScript framework was created! If you are not using some kind of Copilot TODAY, you are missing out a lot.
- josefresco 3y agoMy success rate for writing code with ChatGTP or Copilot is about 5%. Started much better but now I can’t get either to fix any mistakes or generate useful code. Anecdotal but it hasn’t changed my life as a coder.
- qwertox 3y agoYes, but for me ChatGPT today was more useless than a rubber duck. Only when I said "thanks for nothing" it tried to turn all the blabla into code, which was unusable. GitHub copilot instead, as an intelligent Intellisense, is absolutely great; a real blessing and gift to coders.
- nozzlegear 3y agoYMMV but trying to use any code that ChatGPT or Copilot generates for F# (my main language) just leads to a lot of compilation errors or worse, subtly incorrect code.
- Williams77 3y ago[dead]
- elzbardico 3y agoI feel for the future maintainers of all this crappy LLM legacy code in the future. It’s gonna be ugly.
- idle_zealot 3y agoSurely we will get LLMs to maintain it.
- duderific 3y agoSo, I guess LLMs are actually creating jobs rather than destroying them. Not exactly fun jobs though.
- steve_adams_86 3y agoNot exactly well paid either, I suspect.
- bigfudge 3y agoI suspect it will be no worse than enterprisey code. It might even look quite similar, although the comments and docs will be more thorough and less likely to be actively wrong.
- bongodongobob 3y agoAgreed. LLMs will never get any better than they are right now and haven't improved at all in 2 years. Just fancy Markov chains. The only way they can be used to write code is by people who don't know how to code blindly commiting code to prod without any review whatsoever. People who do know how to code couldn't possibly have a use case and it won't make them any more productive. I'm just going to ignore all this LLM nonsense that isn't changing the world at all and you definitely should too.
- siliconc0w 3y agoGood testing is hard to do - coverage is not a categorical good. You can easily write too many tests that calcify programs and basically just creates a change-detector program. Oh it looks like you changed something, oh no - all the tests are broken, but it's okay we can now ask the LLM to regenerate them! 100% Coverage! Amazing! What progress!
- webdood90 3y ago> ... basically just creates a change-detector program interesting perspective - why do you think this is a bad thing? to me, it's an opportunity to verify that the change is intended. without it, how do you know that the program does what it is supposed to do?
- whoisjuan 3y agoNo op, but I don’t think test-driven development resounds with everyone who writes code. I don’t want to write tests for everything. I just want to write the ones that matter.
- nyrikki 3y agoThat is a common misconception about TDD. TDD is _about_ writing tests that matter, but most people think it is about writing all unit tests first. If you are following TDD anywhere close to the way it is described, you will only be writing tests that relate to domain functionality first. Note how it is described here, although it is turse. https://martinfowler.com/bliki/TestDrivenDevelopment.html https://martinfowler.com/bliki/TestDrivenDevelopment.html The coverage metric as a goal writing style doesn't work for TDD, sorry you were exposed to that. You are correct that model doesn't work.
- randomdata 3y ago> The coverage metric as a goal writing style doesn't work for TDD Coverage is not a goal of TDD, but in practice you will have 100% coverage by following TDD as you would never have reason to write code that isn't covered by test. Ultimately, the purpose of coverage tools is to let you know what you might have forgotten to clean up during a refactor, to help you remove what you missed.
- aussieguy1234 3y agoAlready done it with GPT-4. I showed it a TypeScript module, asked it to generate a unit test and it made a working test not only covering the happy paths but a few edge cases as well.
- ramoz 3y agoYea… agree. I’m not resonating with the downvotes here on similar comments. ChatGPT goes above and beyond for me in many ways. Tests seem… easy in terms of gpt capabilities. Last week I had it write python that traversed an AST and construct a react flow graph as well as the component. I made no edits, went through a few iterations of prompt feedback, and it worked great. Many similar interesting abilities I’ve observed from gpt.
- yes_man 3y agoI think the future of development is the other way around. Devs and PMs define the goalposts with tests, AI will do the implementation
- yes_man 3y agoI think the future of development is the other way around. Devs and PMs define the goalposts with tests, AI will handle the implementation
- ajmurmann 3y agoI find it interesting that generally the first instinct seems to be to use LLMs for writing test code rather than the implementation. Maybe I've done too much TDD, but to me the tests describe how the system is supposed to behave. This is very much what I want the human to define and the code should fit within the guardrails set by the tests. I could see it as very helpful though for an LLM to point out underspecified areas. Maybe having it propose unit tests for underspecified areas is a way to do look at that and what's happening here? Edit: Even before LLMs were a thing, I sometimes wondered if monkeys on type writers could write my application once I've written all the tests.
- deleted 3y ago[deleted]
- ralusek 3y agoI basically agree with this but some caveats. I often find there are maybe 5% of the tests I should write that only I could write, because they deal with the specifics of the application that actually give it its primary purpose/defining features. As in, it's not that there is any test I believe AI eventually wouldn't be able to write, it's more that there are certain tests that define the "keyframes" of the application, that without defining explicitly, you'd be failing to describe your application properly. For the remaining 95% of uninteresting surfaces I'd be perfectly happy to let an AI interpolate between my key cases and write the tests that I was mostly not going to bother writing anyway.
- ajmurmann 3y agoYou are probably right and the percentages change with the language and framework being used. When I write Ruby I write enormous amounts of tests and many of these could probably be derived from the higher-level integration tests I stared with. In Rust on the other hand, I write very few tests. I wonder if this also shows which code could be entirely generated based on the high-level tests.
- xboxnolifes 3y ago> I find it interesting that generally the first instinct seems to be to use LLMs for writing test code rather than the implementation. Maybe I've done too much TDD, but to me the tests describe how the system is supposed to behave. This is very much what I want the human to define and the code should fit within the guardrails set by the tests. I feel the same way about how test code is viewed even outside of AI. A lot of the time the test code is treated as a lower priority code given to more junior engineers, which seems like the opposite of what you would want.
- acituan 3y agoUnless well separated, this will easily turn developer-hostile by some clueless management demanding high coverage and enthusiastic juniors smuggling in massive amounts of AI tests so that at the end of the day you will need get a rubberstamp from an hard-to-maintain llm-gen test code each time you want to submit your work. Yes authoring some tests might be sped up but not necessarily maintaining them - or maintaining the code under test because you are not necessarily generating good ones. Not to mention sweating over tests usually help developers with checking the design of the code early on too; if not very testable, usually not a good design either, e.g not sufficiently abstracted component contracts which suck in a context where you need to coauthor code with others. What some people miss is that tests are supposed to be sacrifical code, that most of which will not catch anything during their lifetime - and that is OK because it gives an automated peace of mind and saves from potential false clues when things fail. But that also means max investment into a probabilistic safeguard is not gonna pan out at all times; you will always have diminishing marginal utility as the coverage tops. Unless you're writing some high traffic part of the execution path - e.g. a standard library - touting high coverage is not gonna pay off. Not to mention almost always an ecology of tests need be there - not just unittests but integration, system etc - to make the thing keep chugging at the end of the day. Will llm's sit at the design meetings and understand the architecture to write tests for them too? Or what they can do will be oversold at the expense of what should be done. A sense of "what is relevant" is needed while investing effort in tests - not just at write-time but also at design-time and maintain-time - which is what humans are pretty OK at, and AI tools are not. What llms can save time with is keystrokes of an experienced developer who already has a sense of what is a good thing to test and what is not. It can also be - and has been - a hinderance with making the developers smuggle not-so-relevant things into the code. I don't want an economy of producing keystrokes, I want an appropriately thought set of highly relevant out keystrokes, and I want the latter well separated from the former so that their objective utility - or lack thereof - can be demonstrated in time.
- Jtsummers 3y agoQuoting myself (lightly edited) from when the paper itself came up. They misrepresent the stats in their writeup. https://news.ycombinator.com/item?id=39406726 https://news.ycombinator.com/item?id=39406726 Their abstract doesn't match their actual paper contents. That's unfortunate. Their summary indicates rates in terms of test cases: > 75% of test cases built correctly, 57% passed reliably [implying test cases by context], and 25% increased coverage [same implication] The actual report talks about test classes, where each class has one or more test cases. > (1) 75% of test classes had at least one new test case that builds correctly. > (2) 57% of test classes had at least one test case that builds cor- rectly and passes reliably. > (3) 25% of test classes had at least one test case that builds cor- rectly, passes and increases line coverage compared to all other test classes that share the same build target. Those are two very different statements. They even have a footnote acknowledging this: > For a given attempt to extend a test class, there can be many attempts to generate a test case, so the success rate per test case is typically considerably lower than that per test class. But then in their conclusion they misrepresent their findings again, like the abstract: > When we use TestGen-LLM in its experimental mode (free from the confounding factors inherent in deployment), we found that the success rate per test case was 25% (See Section 3.3). However, line coverage is a stringent requirement for success. Were we to relax the requirement to require only that test cases build and pass, then the success rate rises to 57%.
- jimbob45 3y agoFor greenfield projects, these LLM coders would be invaluable. For my old codebase with observed requirements and magic numbers? Lol it’s going to be just as confused as I am.
- cavisne 3y agoDoesn’t meta famously not do much testing at all? Ie they use experiments to “test in prod”.
- anoopelias 3y agoI thought that unit tests are a balance. A balance of not too much, not too little. "Too little" means you are not covered on the edges. "Too much" means the tests are too rigid its scary to change the code. Ideally, "one change" (Whatever that might be) in production code should cause exactly 1 test to fail. How does TestGen-LLM address this problem?
- galaxyLogic 3y agoHow does the AI know what tests it should write? I think this is an interesting experiment but somewhat dubious. The way I see AI would best help software development is that I the programmer have a question about my or somebody else's code, which the AI then answers, sometimes with a code-proposal but not always. It should be able to answer questions like "Is there a way to simplify this code? What are some inputs that would cause an error?" etc. AI should help us understand the code, and understand how to improve it. Not write all of it on its own because if we don't tell it what to do, it cannot know what we want it to do. Tests is a good example. What do we want it to test?
- avereveard 3y agoDevelopers will do anything not to write tests
- mdaniel 3y agoMy life experience has been that is often the intersection of two very hard problems: test-thinking is a learned skillset that often consumes a lot more active-thought than implementation-thinking and, as I repeatedly and loudly say to my team: testing is always AGAINST REQUIREMENTS. No requirements means no accurate tests, only busywork/metric-gaming. And, as I also always point out: no, your fever-dream one sentence statement of outcome is not a "requirement" The bad news is that often the business folks don't know what they want, either, which is how "agile" became a thing. I recognize the ship has sailed on that, but it's "cake and eat it too" to think one can have good tests and ship "PoCs that do something valuable" in 2 week increments
- TeeWEE 3y agoThe proof is in the pudding, show me the code! In my experience LLM are smart but sometimes inconsistent and over a long chat it might say things that are logically self contradictions… when you tell it that it confirms it. It just seems like it lacks a consistent world view. I don’t trust them yet. Maybe with even more scale they become better. They act a little bit like young children, with a lot of domain knowledge.
- Fricken 3y agoMeta likes to release positive news about itself in the wake of it's competitors misfortunes.
- Temporary_31337 3y agoAll this to write another CRUD app ;)
- MASNeo 3y agoOk, so test case generation has been around a while and now that it is working, where is the GitHub Action?
- sandGorgon 3y ago>using private, internal LLMs that are probably fine-tuned with Meta’s codebase. what does this mean ? i would have thought they would simply use codellama. is there any research around privately finetuned code llms ? why would they be better ?
- mdaniel 3y agoI am not privy to Meta's situation, and to be honest don't have any hand-to-hand experience with finetuned code LLMs, but my mental model is that any corpus of rules will always produce better outcomes when taking local norms into consideration. It's a silly one, but code formatting styles is a perfect example: a hypothetical Google one that has been finetuned on the Google codebase will more easily produce code that already follows their documented code style merely because it has seen more "already correct" examples. Variable nomenclature, method ordering, what things are versus are not documented, any nullability annotations (where appropriate to the language), etc are more that spring to mind More germane to this discussion, I would guess a locally tuned model will also recognize the kinds of things they care about testing, up to and including spotting any bug fix tests that were hard won and can carry forward in any such generated tests for future code
- paradoxyl 3y agoJust another way to censor the free speech of those who opppose the technocracy, or "private-public partnership" or whatever weasel words they use to take away freedom from the masses.
- bjackman 3y agoThese papers are interesting but I think it's impossible to have a valuable opinion without practical experience using the tool and reviewing its output on a codebase you know well. Everyone seems to feel one way or the other about AI code, it's a very political topic. But I would just wanna try it and see. This is pretty interesting, because a lot of these technologies are staggeringly expensive to develop. The AI tooling I've used so far has been somewhat useful, but if it doesn't get much better it won't have been worth the cost that was paid to create it. I'm pretty optimistic about what will be achieved but even with my optimism it's far from clear that it's actually gonna pay for itself.
- haliskerbas 3y agoNice this will make people 15% more effective so we can do another 10% company wide layoff at least!
- adi4213 3y agoAudiobook summary of the paper : https://player.oration.app/ec4770f4-3c2e-47a5-8257-492c25369519 https://player.oration.app/ec4770f4-3c2e-47a5-8257-492c25369...
- test1120 3y ago[flagged]
- test1120 3y ago[flagged]
- test1120 3y ago[flagged]