4 ms·
Okay so I came here expecring something different. Turns out that this is a stack overflow search engine. So ever since I read a book on AI I have been ponderi
by TBF-RnD 7y ago
Okay so I came here expecring something different. Turns out that this is a stack overflow search engine.
So ever since I read a book on AI I have been pondering on when and how computers will steal our jobs.
I do have an idea on the subject though if anybody is interested in working with me on this please comment:
So with github and loads and loads of other open source repos don't we have enough training data.
Couldn't throwing loads and loads of processing power on training an AI on commits create an intelligence with an "actual" understanding of day to day programming problems.
Love to discuss this please comment! Crazy idea but huge if it works!
- YeGoblynQueenne 7y ago>> Couldn't throwing loads and loads of processing power on training an AI on commits create an intelligence with an "actual" understanding of day to day programming problems. I'm curious to know why you think it's possible to create "an intelligence" with "actual understanding" thanks to lots of processing power and data. What makes you think this is actually doable? Do you think that this has been done before, for example?
- TBF-RnD 7y agoOkay so seeing how a brain is made up from 86 billion neurons. 16 of those in the most recently developed cerebral cortex. Randomly creating a set of random configuration of billions of neurons and then subjecting them to simulated evolutionary evolutionary preasure on the basis of how close they come to A) the correct solution B) compiling code that runs C) tests if available* Kill of the worst performing networks. Mate the most succesful and have the offspring replace those who where not fit enough for survival. How would it work exactly, I wouldn't know which is sort of the point. But sooner or later representations of abstractions within the code would appear within the graphs. I get that this is not either trivial and cheap. Let alone hosting the wast amount of source code and parse it would require computing power weighing in at quite a hefty sum. Human beings probably have the advantage of having built in functionallity that we are born with encoded in our DNA. I suppose that a fair bit of "intelligent design" would be required. That is having coded parses that breaks down the source code into input signals that would help speed up the evolution. Even reaching halfway there could be useful. Say by providing useful analysis, maybe warn if something seem stupid. To some extent I believe that for a natural language programming language to be viable some understanding of abstract concepts have to be present in the algorithm. That's why I don't think we'll ever program via speech recognition for example. Whenever Siri is smart enough to parse what we are trying to do, why bother with programming. Why not cut the middle man and plug her directly into Jira. but I'd love to be proven wrong regarding this though. I am by no means saying that it would be easy or even viable within the forseable future. Thank God for that because that would means unemployment for quite a few of us. * As I am writing this it struck me that test driven development just happens to make this easier.
- YeGoblynQueenne 7y ago>> How would it work exactly, I wouldn't know which is sort of the point. But sooner or later representations of abstractions within the code would appear within the graphs. Sorry to be mean but all this sounds awfully confused and terribly hand-wavy. It sounds like you think that stuff would just magically work like you want it, if only you mix up enough ingredients (data, computation, neural networks, evolution, abstractions, representations, graphs...). Well, stuff doesn't just magickally work. Lots of people (like, literally thousands) are training machine learning models and not a single model has ever come close to "understanding" anything. I'd start from trying to figure out why that is the case, and then see if something can be done about it. Otherwise -and, sorry again, but- all these are mere time-wasting fantasies.
- tracer4201 7y agoWhen op used the word “understanding”, I interpreted it as some program that can provide code snippets, skeletons, or whatever abstraction in some contextual way so it’s not just a dumb table lookup. You’re post is unnecessarily pessimistic and generally bad faith. And maybe ops vision can’t be realized today, but if someone is thinking beyond the current state of the art, it’s short sighted and incredibly ignorant to dismiss them as time wasting fantasies.
- YeGoblynQueenne 7y agoAssuming the OP expects to gain something from this conversation, the most benefit that can be gained is from telling them why their idea makes no sense. That is what "good faith" means and that is what I did. Your comment, on the other hand, seems to be confusing good faith with indulgence. Just saying "yes" to people and being nice to them does them no favours.
- tracer4201 7y agoA discussion that makes us think about the boundaries of what is or isn’t possible today is indulgence? You failed spectacularly by your own definition of a good faith discussion by dismissing ops post as “fantasy”. You give no explanation of why it is unachievable today or why it would continue to be impossible in the future. My previous team had a member who outright dismissed ideas and criticized things for the sake of sounding intelligent without proposing solutions to further the discussion or solve the problem. We let him go. Good luck.
- tylerhou 7y agohttps://arxiv.org/abs/1904.02818 https://arxiv.org/abs/1904.02818 Disclaimer: I work at Google.
- TBF-RnD 7y agoThank you and see my response to thomasahle!
- thomasahle 7y agoGithub basically did this themselves using (comment, code) pairs as training data: https://github.blog/2018-09-18-towards-natural-language-semantic-code-search/ https://github.blog/2018-09-18-towards-natural-language-sema...
- TBF-RnD 7y agoI love it when I find that someone has done something that I've been thinking about. Proves that I am not crazy and saves me from the burden of putting in the hard work myself! :D At a closer look I don't think that mapping up correlations in a vector space is good enough for the task of fixing bugs based on an issue tracker description though. The way I see it word2vec is basically statistical connections. So if you compare it to google translate which also works via statistical correlation. Which explains the errors it might do. At some point to get a good translation you'd need human intelligence. To continue with the google translate analogy. Imagine a translator trying to translate Fyodor Dostoyevski's The Karamazov Brothers to English. That person have to have a tremendous understanding of Russian and English culture not only now but also when it was written along with ideoms and so forth. Along with a deep understanding of politics and religion not to botch the job completely. Or imagine entering a greek epos written in hexameter into google translate and publishing the book. So you see just as translation and translation might be two totally different things, search and making changes to source code might be different beasts. Then again you are totally right in that they are doing it. Getting access to this datasource for minining is probably a huge factor in microsoft's purchase of github, so I'd think it would be a safe bet to say that they are working on it... At least originally, I am sure it's an incredible complicated piece of software by now applying multiple heuresstics.
- machiaweliczny 7y agoThe possible models where change C fits are unlimited. Code describes only behaviour but it's hard to guess intent. So unless languages won't be just specification languages (as close to describing intent as possible) then I doubt it's possible. Anyway already part od programming work is deciding what you really want. Maybe that's where machines might help - with discoverability. I actually think that more likely is some CRUD maker as an expert system? (you put what you vaguely need, answer questions and it comes with solution), doesn't need ML.
- TBF-RnD 7y agoThe same thing applies for a human to try to solve a problem as well doesn't it? And we have bugs and inefficent code don't we? I remember seing a talk on youtube about a lisp program finding every program that would fit a certain solution. So reasoning about things like this is what a lot of lisp's gurus are doing all day long basically. You are probably right about AI being used in an assitive part at first. I imagine it'll be a while until people dare putting their business logic in the hands of an AI. Probably CRUD applications would be the first to be automated. Boring tedious work that has been done a million times. And unhappy coders are probably more prone to create bugs make bad decisions. Also there are litteraly thousands and thousands of implementations doing the same thing with variantions that are more or less just flavour. Then again if you are at analyzing all open source software ever written. Why not plow through all open source issue trackers at the same time. If you found the top 20% bugs that cause 80% of the issues and have get an AI that solves 20% you'll be a very very rich man if you feel like it. So don't expect to tell a computer to figure out the best way to create top off the line lossless video codec anytime soon or anything else. Then again it's hard to find a coder that you could tell that and get back the results you wanted either...