4 ms·
> 3. The "Trust but verify" pattern To add on to this point, there's a huge role of validation tools in the workflow. If AI written rust code compiles and the
by pcwelder 2y ago
> 3. The "Trust but verify" pattern
To add on to this point, there's a huge role of validation tools in the workflow.
If AI written rust code compiles and the test cases pass, it's a huge positive signal for me, because of how strict rust compiler is.
One example I can share is
https://github.com/rusiaaman/color-parser-py https://github.com/rusiaaman/color-parser-py
which is a python binding of rust's csscolorparser created by Claude without me touching editor or terminal. I haven't reviewed the code yet, I just ensured that test cases really passed (on github actions), installed the package and started using it directly.
- aiono 2y agoI checked the code. The code validates that values are within the range https://github.com/rusiaaman/color-parser-py/blob/ec739c80ba73edbf8d01a322ac3d38814f95de39/src/lib.rs#L14 https://github.com/rusiaaman/color-parser-py/blob/ec739c80ba... but the library it wraps already does validation and a lot more checks (see range check here: https://docs.rs/csscolorparser/latest/src/csscolorparser/parser/mod.rs.html#449 https://docs.rs/csscolorparser/latest/src/csscolorparser/par...). So the code generated by AI has unnecessary checks that will never be visited.
- fmbb 2y agoDid you write the test cases?
- User23 2y agoTDD is, in my opinion and based on some tentative forays, one of the areas where LLM assisted coding really shines.
- pcwelder 2y agoNo everything in the repo is AI generated.
- dkdbejwi383 2y agoHow can you trust the tests? Sure, they may pass, but they could be testing for the incorrect outcome.
- fmbb 2y agoThem you really have no idea what you are running. Why include test cases at all?
- eesmith 2y agoYou might look into why rgba_255 return a fixed-length tuple while rgba_float returns a fixed-length list. If it's so important to test isinstance(r, int) then you should also have tests for g and b, and likely similar tests for the floats. Is is really worthwhile to 'Convert back to ints and compare' when you know the expected rgba floats already?
- tuetuopay 2y agoThe readme even confuses itself, as the example shows rgba_255 returning a list and not a tuple. Oh well, I guess Claude was confused by the conventions between Rust and Python. Also, all the checks of "if u8 < 255" will make me not want to use this library with a 10-foot pole. It screams "ai" or "I don't know what I'm doing" so much.
- pcwelder 2y agoFirst one is due to me asking it to return a 4 tuple instead of a list for the rgba_255 specifically, I guess it didn't update Readme or other return values. The second is an artefact of a test case failing, which it tried to fix it using this check. Thankfully not a correctness failure, only optimisation issue. You're right though it's not worth publishing for general public.
- tuetuopay 2y agoWell this is a good experiment. I don't find your idea bad at all: use AI to autogenerate bindings to expose a library in another language. This would be a good usecase for AI as it's not complex (well, most of the times) and is a lot of boilerplate. Publishing the repo is worth it, because it showcases what the AI can (and cannot) do, and it is not there yet. But as a real package to pypy, indeed less.
- eesmith 2y agoWhat gets me is that tools like SWIG exist, in that case a tool which started in the 1990s to read a C header file and autogenerate bindings for Python and other languages. Or, JPype uses reflection to generate run-time bindings from Python to Java. Why does it require AI and all of this infrastructure?
- LiamPowell 2y agoI've been saying for a while that formal verification seems like the obvious way forward here. Test cases are good but only cover a limited number of cases, whereas formal verification can cover all cases. The biggest reason that we don't see formal verification being used more is that it takes some time to formally define the output of an algorithm and significantly more time to convince a theorem prover that your code is correct according to the pre- and post-conditions. If we can get LLMs to work well alongside formal verification tools then I think that would greatly improve their usefulness. Here's a made up and trivial example of how I imagine this might look: User: Write me a binary search over "type A_Type is array (Integer range <>) of Integer;" LLM: Implementing the following: function Search (A : A_Type; Target : Integer) return Integer with Pre => Sorted (A) and Target in A, Post => A(Search'Result) = Target; First implementation: function Search (A : A_Type; Target : Integer) return Integer is Left : Integer := A'First; Right : Integer := A'Last; begin while Left <= Right loop declare Mid : Integer := (Left + Right) / 2; begin if A (Mid) = Target then return Mid; elsif A (Mid) < Target then Left := Mid + 1; elsif A (Mid) > Target then Right := Mid - 1; end if; end; end loop; end Search; Compiler: Phase 1 of 2: generation of Global contracts ... Phase 2 of 2: flow analysis and proof ... wrapper.adb:12:36: medium: overflow check might fail, cannot prove lower bound for Left + Right 12 | Mid : Integer := (Left + Right) / 2; | ~~~~~~^~~~~~~~ reason for check: result of addition must fit in a 32-bits machine integer wrapper.adb:12:45: info: division check proved wrapper.adb:14:19: medium: array index check might fail 14 | if A (Mid) = Target then | ^~~ reason for check: value must be a valid index into the array LLM: Looks like there's a possible overflow for very large arrays, attempting new solution... Etc. until a valid solution is found and the result is returned to the user. As an aside, this bug was present in Java for 9 years and would have been caught by formal verification (which admittedly was more difficult at the time): https://research.google/blog/extra-extra-read-all-about-it-nearly-all-binary-searches-and-mergesorts-are-broken/ https://research.google/blog/extra-extra-read-all-about-it-n...
- eru 2y ago