2 ms·
sorry i am using dictation software. I have found this research paper[0] that describes using a machine translation with a CNN and a r n n. it seems one of the
by raidicy 6y ago
sorry i am using dictation software.
I have found this research paper[0] that describes using a machine translation with a CNN and a r n n. it seems one of the main problems that come up with is a lack of training data so to progress this further it seems like we would need a big corpus of parallel text between natural English describing what the code does in the actual source code. I only have some novice experience with training machine learning models through transfer learning with fast AI but I have found a repository of tutorials[1] that teach you how to do machine translation between two languages and it seems maybe this approach might be applicable. There is a small corpus of text[2] with python code with annotations but and the paper itself at States. It probably needs more to get a higher than a 74% accuracy.
the only thing I can really think is to scrape GitHub make a website where you can get people to crowdsource annotations for code samples. I'm not sure on how to deal with how much and how little code would be for a sampling. I eat how much of the code would you need for a good training sample versus just using isolated functions. Perhaps that something like this might be used in correspondence with GPT.
[0]https://www.groundai.com/project/machine-translation-from-natural-language-to-code-using-long-short-term-memory/1 https://www.groundai.com/project/machine-translation-from-na...
[1]https://github.com/bentrevett/pytorch-seq2seq https://github.com/bentrevett/pytorch-seq2seq
[2]https://ahclab.naist.jp/pseudogen/ https://ahclab.naist.jp/pseudogen/