14 ms·
GPT-Migrate converts repos from one lang/framework to another
- DaiPlusPlus 3y agoI'm tempted to point this at the Linux kernel repo and tell it to convert it "to work on Windows" and see what happens...
- jayd16 3y agoWill it spit out WSL1 or WSL2?
- PrimeMcFly 3y agoIt will be rewritten in .NET and rely on HyperV.
- olalonde 3y agoRewrite in Rust for guaranteed 1st spot on HN's front page.
- transitivebs 3y agothis would be legendary. prolly going to fall on its face for something of this complexity, but "the next big thing always starts out looking like a toy"
- EGreg 3y agoOr in Brainfuck
- deleted 3y ago[deleted]
- gailees 3y agoHas anyone tried it yet? Would be crazy if it worked.
- cwalv 3y agoI remember when Siri was first demoed, a dev I worked with was convinced it was just like having human personal assistant. I hadn't seen the demo, but I was confident he was wrong. Haven't tried this either, but I'm equally confident it wouldn't work .. or even kind of come close to working
- DaiPlusPlus 3y agoSiri was objectively better in the past - I remember using it when I got my iPhone 4S and being impressed with how it returned very relevant data from WolframAlpha and other sources - then it all dried-up and all Siri is useful for now is voice-control for my bedroom lights and phone alarm...
- TT-392 3y agoAs far as I have tried, at least chatgpt basically cannot do brainfuck. I think it might just be too far from natural language. It basically just gives you hello worlds with randomly injected characters.
- tylercrompton 3y agoDoesn't ChatGPT excel at providing step-by-step instructions? That's kind of what Brainfuck is with only eight options for each step. I'd imagine that it could reach the end result using a combination of those eight options. Can it not?
- TT-392 3y agoIdk, I'd get nonfunctioning code. With explanations of what every piece of code does. Just that the code it is explaining, is actually gibbrish.
- YeGoblynQueenne 3y agoI wonder why. GPT-3 could do brainfuck just fine: https://papers.nips.cc/paper/2021/hash/0cd6a652ed1f7811192db1f700c8f0e7-Abstract.html https://papers.nips.cc/paper/2021/hash/0cd6a652ed1f7811192db... (though worse, for complex cases, to one of the compared systems - see figure 2) (oops, fully disclosure: the system in question is my PhD work).
- 0xpayne 3y agoWill use GPT-Migrate to migrate to rust ;)
- gailees 3y agoDid you try it yet? Curious how far it could get.
- sigotirandolas 3y agoFunnily enough, it seems like someone already attempted to do this around 20 years ago - search for "Umlwin32"
- isoprophlex 3y agoVery interesting concept! I'm left wondering if you could also use this to document or clean up machine generated code. Eg, some process generates a huge wad of bytecode, or autogenerated Java; a GPT tool cleans it up so you can actually do some things with it as a regularly skilled human.
- acover 3y agoHow well does this work? Could it generate tests to confirm behavior in the target and source?
- xeromal 3y agoLooks like it does. (Optional) If you'd like GPT-Migrate to validate the unit tests it creates against your app before it tests the migrated app with them, please have your existing app exposed and use the --sourceport flag.
- gailees 3y agoSeems like this actually works by generating tests and continuing to try different things until the tests run successfully on both the source and the target codebase.
- seanthemon 3y agoYeah.. same.. :<
- transitivebs 3y agoyep; it's generating unit tests and making sure that things are at least statically correct in terms of syntax. early days, but the ability to have loosely-defined acceptance criteria would go a long way here (like you would have for a well-groomed dev task issue).
- aardvarkr 3y agoI have a community website that I built a while ago in angular that I wish I had built in react instead. Mayne then I would at least get some help maintaining the open source repo on GitHub. I’ll give this a try and report back
- gailees 3y agoDid it work? Link the repo.
- bdhcuidbebe 3y agoSuch brave
- cfn 3y agoThis would really useful if it worked with legacy code. For example, you could migrate all that COBOL code into Java or Python, or all the Fortran scientific code into C++ or Python. I tried to migrate a twenty year old Visual Basic 6.0 project to c# by doing it piecemeal with GPT4 and it failed completely. Both in the UI and the backend. I am keeping my fingers crossed for a GPT n+1 that actually can do this. Incidentally, I found out that GPT4 (as in chatGPT) is very useful if you need to program in VB6 which is nearly absent from search results these days.
- codeonline 3y agoWhen you say it's failed completely what do you mean? I've been translating between c# and python and having a great deal of success at the function and class level. I even ported unit tests easily between xunit and pythons unittest library. I've got close to 100% test coverage so I'm fairly confident it's done well
- throwawayadvsec 3y agotranslating between popular mainstream languages is not usually an issue. translating a 20yo project written in a language that isn't used much anymore is a lot different
- fomine3 3y agoWe should upload VB6 projects to GitHub.
- lionkor 3y agoaren't the tests also written by chatgpt then?
- still_grokking 3y agoBut they are confident everything is done well! The code looks good, and that's a great success. What else would you like?
- cpursley 3y agoInteresting and potentially great use case. How does it figure out the minutia selecting adequate dependencies?
- lionkor 3y agoIf you think this does anything more than serve files to chatgpt with a custom prompt, youre not living in this reality
- bdhcuidbebe 3y agoITT most are wishfully duped
- simion314 3y agoUnfortunately I can't trust GPT will not hallucinate something, I tested it to code review my code and it hallucinated issues. It would be great if you could give it old , ugly code and you could get something better. Maybe Intellj guys can use this tools to increase their productivity and we can get 100% correct tools that work with AST not with tokens, and can do advanced transformations and review that you can trust without having to double check it.
- dumdumchan 3y agoThey will have to train their own model. ChatGPT has been trained on textual data.
- villgax 3y agoI just want a small model which will ingest a language spec & be capable of understanding external repo code
- gailees 3y agouseadrenaline.com
- still_grokking 3y agoI just want a model which will be capable of understanding actually anything…
- _odey 3y agoWho/what is the: 1. Author 2. Copyright holder 3. Copyright license ...of the code generated by this tool? Unless the answer to all of these is unambiguously "the original", then you shouldn't be using such tool on any code, especially on your employer's intelectual property. Sorry to be so negative about it but this is something that I see skipped over in all discussions related to AI. Just because it's AI does not make it immune to copyright law. You're giving away your code to a 3rd party company under their terms and conditions, and receivig some new code back, again under their terms and conditions. The fact that it uses AI under the hood is irrelevant, you're dealing with a business that produces you an output and you should know the terms before submitting anything to them, especially if you don't own that thing.
- transitivebs 3y agoGreat feedback. I'm not the author of this project, but in my understanding, it's the same as if you were to write the code yourself. The project doesn't publish anything and it works entirely locally aside from LLM calls (which could in the future be 100% local as well). So you remain the author and have complete control over the license of the generated code.
- jamil7 3y ago> and it works entirely locally aside from LLM calls So not entirely locally. Yes these could eventually also run locally but OP’s point still stands.
- _odey 3y agoThis is great to hear but I don't fully know how far reaching are openai's claws (since it does require an openai api key). If it runs 100% locally then yes, it would be safe to use.
- transitivebs 3y agototally agreed that diff projects need to be careful w/ sharing proprietary code w/ third parties. openai's official stance is that it will never use API calls as training data, and that in my understanding it may retain API call data for up to 30 days for compliance purposes, but that it legally won't store it beyond that (whereas chatgpt convos are meant to be stored and used for training purposes). as a next step, they could provide a swappable version of the LLM provider using something like https://github.com/imartinez/privateGPT https://github.com/imartinez/privateGPT, https://github.com/alexanderatallah/window.ai https://github.com/alexanderatallah/window.ai, etc. would love to have a standard develop here as the community matures around LLM usage
- revskill 3y agoI expect to see how you use your gpt-migrate to migrate gpt-migrate to JS. Then from JS to Python again. Run the test and compare. Once done, good job !
- rurban 3y agoWhow, I always wanted to move away from my old perl web app to typescript. Will report the results
- gailees 3y agoPlease do. I’m curious to see what this looks like in production repos.
- hkt 3y agoAh, finally a way to turn the output of nodejs programmers into something less hateful! All it took was the early days of the singularity ;)
- gailees 3y agoI can’t believe the author doesn’t allow conversion to JavaScript It’s the most popular language for heavens sakes. Developers really need to stop bringing their religion into open source.
- transitivebs 3y agowould be cool to see a loop of: js ⇒ python ⇒ js and then compare the output JS w/ the input JS. could get wilder too like: js ⇒ rust ⇒ typescript ⇒ java ⇒ js
- jamaicahest 3y ago> Developers really need to stop bringing their religion into open source. Says the person complaining "why isn't my language supported"
- gailees 3y ago[flagged]
- tylercrompton 3y agoThat's a take.
- jahsome 3y agoDelusion and cognitive dissonance run deep for a large chunk of the JS world.
- discordance 3y agoSupported languages are in config.py: Python, JavaScript, Java, Ruby, PHP, C#, Go, Rust, C++, C++, C++, C, Swift, Objective-C, Kotlin, Scala, Perl, Perl, R, Lua, Groovy, TypeScript, TypeScript, JavaScript, Dart, Elm, Erlang, Elixir, F#, Haskell, Julia, Nim, PHP
- still_grokking 3y ago| sort | uniq C, C#, C++, Dart, Elixir, Elm, Erlang, F#, Go, Groovy, Haskell, Java, JavaScript, Julia, Kotlin, Lua, Nim, Objective-C, PHP, Perl, Python, R, Ruby, Rust, Scala, Swift, TypeScript
- stuaxo 3y agoI've done using chatgpt on some demo code I got it to write between languages and Frameworks, it works to an "ok" level. Edit: this was demo code I asked chatgpt to come up with in the first place, so the output had no problems license wise that the input didn't already have.
- complex_exp 3y agoMuch less impressive, though still useful: ChatGPT is an awesome movie subtitle translator. Only very unusual phrases need to be corrected, often there are no such cases. There are projects on GitHub that automate the translation. Short SRT files can be just pasted into the chat with appropriate instructions.
- Makhini 3y agoHow can ChatGPT translate a comedy show without knowing what's going on on screen, and various context that contribute to the humor?
- djmips 3y agoThat's a valid point! In similar vein, I was always impressed with the translation of Asterix and Obelix into English. The puns were done quite well!
- vorticalbox 3y agoFeed the scene into whisper to extract the audio and then feed that into got 3.5/4 for context?
- H8crilA 3y ago> Only very unusual phrases need to be corrected, often there are no such cases.
- mariuz 3y agoWe need migration tool between frameworks : Rails -> Django Fast API - Express
- kingrolo 3y agoIt feels to me as though LLMs should (eventually?) really shine at these kinds of tasks where the intent is already defined in code of some sort and the challenge of the task is lots of detailed legwork that humans find hard, more because it's time consuming and not interesting so hard to focus on, rather than because it's technically challenging. So swapping languages, yeah maybe, but I expect of more practical use would be the situation where you inherit a legacy codebase in an ancient version of a language or framework that hasn't been loved in a long time. I saw this so many times when doing dev team for hire work. Obviously you'd want to do boat loads of testing and there may well be manual work left to do afterwards, but I think it would be the kind of manual work that felt like you were polishing something new and clean and beautiful rather than trying to apply bits of sticky tape to something unmaintainable. I also wonder about eventually being able to say to an LLM "take this codebase and make it look like my code", or maybe one of your favourite open source developer's code. Maybe everyone could end up with their own code style vector attached to their github profile describing their style. You could find devs with styles close to yours to work on your team, or maybe find devs with styles different to yours so you could go and argue about tabs vs spaces or something.
- tinco 3y agoThat's it precisely. I'm working on the exact same thing GPT-migrate is doing, but I'm approaching it from the other direction first. My project is trying to generate a test suite that aims to cover the full original functionality (bug for bug as they say). That way a tool like GPT-migrate has a much better chance of generating the translation without errors and whoever uses it can have more confidence in that the output will be correct. I'm a bit intimidated that Josh came so close in just a week of work but it's also inspiring confidence that this is the right track and it's actually going to work when all the puzzle pieces fall into place. edit: damn, this project actually creates rudimentary tests as well. It's such a lean approach, makes me feel like I'm still coding in 2022 when Josh is firmly in 2023.
- arach 3y agoI agree! This is a pretty elegant approach and while like you I haven't fully internalized the 2023 way of building AI native product, I'm inspired and increasingly confident in how much can be accomplished in a lean way.
- dns_snek 3y agoThese examples always look so interesting and promising until you try them out with anything more than a "Hello world" application. It would be very interesting if it worked beyond trivial examples, but I'm not holding my breath.
- kohlerm 3y agoI bet it doesn't really work for reliably for anything other than a toy project
- ic4l 3y agoHow many people even have access to the gpt-4-32k, or gpt-4-32k-0613 models. I think they give everyone access to the gpt-3.5-turbo-16k, but I have not found a way to request access for the 32k model. There does seem to be an option through azures openai service: https://azure.microsoft.com/en-us/products/cognitive-services/openai-service https://azure.microsoft.com/en-us/products/cognitive-service...
- jumpCastle 3y agoIs Azure as simple to use as OpenAI? Their documentation are much less simple.
- 00117 3y agoIt would be nice if it could estimate GPT usage costs with a dry-run.
- vessenes 3y agoSome good comments in here; I have done a fair amount of work with GPT-4 as a writer of go code, and there are two categories of difficulties for a "Tier 2" (Tier 1.5?) language with GPT-4. The first is API hallucination, which hits as soon as you drop down into non "major" repository packages. Even GPT-4 acts like 3.5, and will cheerfully make up / use old API interfaces, pretend it knows newer versions that it does not know, and generally loop you around in very, very convincing-looking code that just does not work. The second is style related. In particular, Go is picky with its error return semantics, and GPT-4 doesn't worry too much about this; I'm remembering a particularly subtle and annoying deadlock where it didn't defer closing a database connection inside a go routine, or alternately check for an error, and close the handle. On balance, both of these seem super, super solvable, either by a custom LLM, or a next version with updated training. I think of GPT-4 as a reasonable mid-to-senior engineer in terms of output right now, and I think it's reasonable to start trying to port frameworks. That said, I think I'd want it to do an excellent job at porting tests over first, and I'd inspect those heavily, and then I'd consider how to deliver a style guide for the target language in the prompts. By default, GPT-4 doesn't know exactly how you want things coded. One last comment, Claude seems appealing to me here, with its longer context window. That said, I haven't been successful at fully using the context window -- e.g. "here's a tarball of a repo, please do x/y/z". I think word on the street is that the Claude folks use ALiBi, regardless, the 100k attention window from Claude feels more like one that can choose to alight on key areas of the input, not one that can take the entire 100k tokens into context.
- itsTyrion 3y ago> API hallucination This term is better than anything I was able to come up with. The ability of LLMs to make up convincing looking bullshit is remarkable. I'll share a funny (because it's just so dead wrong) thing I had: I was asking about a problem SnakeYAML (popular JVM YAML lib) and it suddenly started adding Jackson (JSON Object Mapper for JVM) annotations, insisting those would work. (Spoiler alert: no)
- YeGoblynQueenne 3y ago>> GPT-Migrate is currently in development alpha and is not yet ready for production use. For instance, on the relatively simple benchmarks, it gets through "easy" languages like python or javascript without a hitch ~50% of the time, and cannot get through more complex languages like C++ or Rust without some human assistance. In other words, it doesn't really work. The current wave of LLM applications still seems to me like someone just invented homeopathy and a whole bunch of people are convinced it's real and are trying to use it to create a cure for cancer. It's just people waving their hands about and intoning magick formulae, that don't work and don't produce anything useful at all. I am curious to see where all this is going to end up. Is someone going figure out a way to make LLMs work for real-world er work? Are we all waiting patiently the next big LLM version to see if it can do the things that the current best-of-the-best can't?
- golergka 3y agoA prototype that works 50% of the time is not homeopathy. A new cure for some types of cancer that would succeed 50% of time would be an achievement.
- groestl 3y agoUnless it straight up kills the other 50%.
- golergka 3y agoThan it's certainly not homeopathy. It doesn't kill anyone.
- YeGoblynQueenne 3y agoThe readme says "~50%" of the time, which seems to be a number pulled out of a hat (to be nice) rather than any serious attempt at quantifying performance (there are no metrics of any kind and the authors are asking for relevant benchmarks). It's more like the authors' feeling, than anything they have ascertained in any reliable way. That's pretty much the same way homeopaths test their "remedies" (i.e. water).
- bdhcuidbebe 3y agosure it does. good riddance kids these days
- shipscode 3y agoI really doubt this works reliably given my own attempts at doing this for extremely small use cases.
- dumdumchan 3y agoIs there an upper limit on the size of code base it can handle? (LLM context size)
- swader999 3y agoSomeday we'll all be writing code that injects the programming language as a dependency. This migrate magic seems bass-ackward.
- enjoylife 3y agoOne could argue we do already via code generation when we define protobuf definitions or other idl’s. But yeah a chat oriented idl which then code generates c, Go, python, just based on the required problem domain is an interesting vision.
- forgingahead 3y agoNeed all Python ML repos to be migrated to Ruby equivalents ASAP.
- aitchnyu 3y agoAnybody who noticed semantics correctly translated? a [1,2]==[1,2] will be different in Python and JS.