4 ms·
I think they've done it backwards in regards to it writing tests. Tests are the check to make sure the A.I is in check. If A.I is writing tests, you have to dou
by UK-AL 4y ago
I think they've done it backwards in regards to it writing tests. Tests are the check to make sure the A.I is in check. If A.I is writing tests, you have to double check the tests.
You should write tests, then the A.I writes the code. It almost doesn't matter what the code is, as long the AI can regenerate the code from tests.
- layer8 4y agoTests don't (can't) prove tthat code is correct. They are merely a rough plausibility check that the code isn't completely wrong and didn't regress. You generally can't derive the right code just from tests.
- UK-AL 4y agoYou can write tests about properties you care about which may not be everything. Generally in some of the more financial applications i've written I would be ok with people rewriting the app as long as it passes the tests. I've even written tests that say this set of input goes to this output, for various different subsets of input. Anything outside of the of the defined input sets fail validation. Than it randomly picks a couple of thousand inputs from the input sets I've defined and runs them. More confidence you need, the more exhaustive setting you put it on. It's a bit like QuickCheck.
- layer8 4y agoYou can approximate it, but to represent really all properties, in the end it becomes a mirror picture of the actual code you are testing, which then begs the question. A random sample of inputs that is hidden from the AI also won’t allow it to derive a corresponding implementation. And if the set of sample inputs is not hidden, then the AI is still free to produce an implementations that only works for those sample inputs.
- UK-AL 4y agoYou'd probably separate example tests and validation tests. Also test descriptions should fed into the prompt to help guide it, like BDD style tests. On test failure, the data is fed back into the prompt about what failed for another iteration. This will help avoid over-fitting, and generate another generation on test failure. I mean you can't guarantee correctness, but you could probably get it pretty close. Humans also have the same problem.