34 ms·
Devin: AI Software Engineer
- xzfyes 3y ago[dead]
- xzfyes2 3y ago中国的酒吧有多少个
- cdeutsch 3y agoI was really hoping Devin was an actual human and this was a meme.
- ridruejo 3y agoGiven the excitement on X right now about this, I don't understand how this is not in the front page already :)
- nycdatasci 3y ago@dang: did this get flagged as spam by chance?
- sohzm 3y agoThere's another post both are not really getting attention. I really wanted to see what hacker news had to say on this :(
- mpg33 3y agoI could take a guess...
- cal85 3y agoSay more?
- pzo 3y agoit was shortly on front page and somehow got buried
- Bjorkbat 3y agoI mean, this might just be existential cope, but my first thought when looking at the Upwork demo posted on Twitter (https://x.com/cognition_labs/status/1767548768734294113?s=20 https://x.com/cognition_labs/status/1767548768734294113?s=20) was that it seemed a little bit suspicious. Namely, the client was asking for an unusually specific (for Upwork) ask. It was an almost perfect example of a job to be given to an AI agent for testing purposes.
- supafastcoder 3y agoTo be fair, a lot of those Upwork job requests are now being written with AI…
- swalsh 3y agoYou're looking at the worst version that will ever be released. If these guys don't get better over time, someone else will.
- ukuina 3y agoCue Microsoft AutoDev.
- HanClinto 3y agoI tend to agree. I really want to see what this does with an open-source backlog. The more mundane (yet critical) of a project, the better. What would be good projects to feed it? I suggested Mozilla and llama.cpp, but there's got to be something better as a use-case for it.
- dustyharddrive 3y agoUsing this on memory-unsafe languages is just asking for trouble.
- zeroonetwothree 3y agoIt doesn't seem like Devin actually did the work requested, which was to provide instructions for using AWS?
- datavirtue 3y agoThis awesome. Until Devin steals your startup idea.
- swalsh 3y agoSoftware is a commodity, find something more valuable to differentiate your startup.
- shombaboor 3y agojust need to get an A list celeb like ryan reynolds involved and people will follow
- gowld 3y agoYou'll need an AI list celeb. A list is dead.
- mattlondon 3y agoThere is no way this is going to make it so that "engineers can focus on more interesting problems and engineering teams can strive for more ambitious goals." Instead it will mean that bosses can fire 75-90% of the (very expensive) engineers, with the ones who remain left to prompt the AI and clean up any mistakes/misunderstandings. I guess this is the future. We've coded ourselves out of a job. People are smiling and celebrating this all - personally I find it kinda sad that we've basically put an end to software engineering as a career and put loads of people out of work. it is not just SWEs - it is impacting a lot of careers... I hope these researchers can sleep well at night because they're dooming huge swathes of people to unemployment. Are we about to enter a software engineering winter? People will find new careers, no kids will learn to code since AI can do it all. We'll end up with a load of AI researchers being "the new SWEs", but relying on AI to implement everything? Maybe that will work and we'll have a virtuous circle of AIs making AI improvements and we'll never need engineers again? Or maybe we'll hit a wall and progress in comp sci will essentially stop?
- krainboltgreene 3y agoThis post has the same energy as when a junior programmer talks to be about the "dead language called java".
- optimalsolver 3y ago>can focus on more interesting problems Kind of like how art was supposed to be what humans would be doing while AI does the jobs we don't want, but looks like that will be the first thing to fall to the machines, while humans fight for carpet installation and plumbing jobs (for a while).
- CipherThrowaway 3y agoAI still can't do art. Tacky AI generated imagery is mid-2020s clip art, already recognizable to consumers and signalling negative brand associations like "cheap", "scam", "low quality."
- Almondsetat 3y ago
- hackerlight 3y agoThis is where inference speed starts to matter. H100 might be cheaper per inference than Groq but cutting down the wait time from 1 minute to 10 seconds could be a big deal.
- anonzzzies 3y agoHave you tried Groq? We did a few days testing on replacing gpt4-turbo with it and, while incredibly fast, the results were horrible, even after a lot of specific prompt engineering. So many hallucinations and such. Our products all have to do with strict generation and software quality; it basically has to fill in the blanks but it was incredibly hit or miss. Some results came in within a second so even a few iterations beat gpt4 when correct, but some needed so many (that we quit) iterations that gpt4 beat it hands down.
- simonvc 3y agothey just run other open models, so you're complaint isn't about Groq, it's about GPT-4 vs mixtral 8x7b accelerated
- anonzzzies 3y agoSure, so when openai moves to groq it might be something. Groq with the current models is impressive but doesn’t work for us is what I am saying. As I don’t actually have access to other models on groq, this is groq as it stands.
- hackerlight 3y agoGroq is hardware not software... It's like saying the H100 hallucinated.
- throwaway11460 3y agoWell, why not. It's like saying that Windows crashed - while it actually was some driver or app that caused it. The hardware (or OS) is useless if it doesn't give good results, even if it theoretically could.
- deleted 3y ago[deleted]
- rafadc 3y agoPrepare for a lot of copycat companies. Hey Devin, copy this company's software.
- _factor 3y agoSo basically taking money away as an obstacle? It’s not illegal to copy an application. Look at Instagram’s stories compared to Snapchat. The only difference is that there might be a little less profit incentive now. Perhaps we’ll get some interesting ideas from people who never would have been able to create them.
- anonzzzies 3y agoIt will try and you end up with nothing but a bill from using Devin.
- shombaboor 3y agothe most successful software will be a result of marketing and going viral (e.g. featured on HN) if the ais just constantly clone each other's apps.
- deleted 3y ago[deleted]
- ukuina 3y agoHey Devin, copy Devin?
- gazelle21 3y ago[dead]
- cvhashim04 3y agoWell, pack it in. It was a great run boys. Onto better things.
- hiddencost 3y ago"first" lol. Making false, grandiose claims like that burns a lot of trust. Focus on execution and quality.
- steve_adams_86 3y agoAlthough the demos are impressive, they seem short and limited in scope which makes me wonder how well this will work outside of these planned cases. Can it do software architecture at all? Is it still essentially just regurgitating solutions? How often will the solution only be 90% correct, which is 100% not good enough? Even so, I realize the demos are still broad in scope and the results are incredible. Imagine seeing this even 2 years ago. It would seem like magic; you wouldn't be able to believe it. Today, this was inevitable and entirely believable. There will be even better versions of this soon.
- andoando 3y agoThere is a similar product called Sweep AI thst I tried. For extremely simple things "like add a button to the page that prints hello" it was very good. I then asked it to do something more complex, which was to render my d3.js graph vertically rather than horizontally, and it tried to redefine constant variables (it just added a new modified code block without deleting the old one), put function clauses in places that were not synctactic. After I manually fixed those, the functionality just didnt work.
- swalsh 3y agoAh yes, you've entered the first stage of grief. Denial. Next you'll start bargaining, you'll get angry, and you'll become depressed, eventually you'll just accept that AI is taking over software. In my mind, I've concluded that I have less than 3 years to find an off ramp.
- 4star3star 3y agoWhat kind of work do you have in mind?
- swalsh 3y agoIn terms of my "off ramp"? I have a multi-part plan. Immediately, i'm working to get closer to the business. To be closer to the position of defining requirements, not implementing requirements. Secondary, i'm experimenting with ideas I hope can become a business. as a final fallback, I have a hobby woodshop in my basement, and I love making furniture.
- devinprater 3y agoHey, they named it after me!
- Buttons840 3y agoMe too. It sucks. At least it's a small team who will probably be shown up by bigger players in the market and go out of business.
- lacoolj 3y agoThis is just nice packaging on top of current models. Very nicely done but still not a giant leap forward from what is already here
- steve_adams_86 3y agoI'm not suggesting great work didn't go into this, but I was able to build a very crude version of this on GPT 3.5. It was evident then that the real power of these models isn't in chat, but in a sandboxed environment where they can recursively iterate on solutions and feedback from their sandboxed environment. I was able to feed mine small applications with bugs and have it comb through and find the bugs, write and apply solutions, write tests for solutions, etc. Adding features was too hard to implement in my limited spare time, but was clearly possible. I would have needed some form of test running for UIs or CLIs, and I wasn't prepared to go that deep on a project I wasn't going to get much out of. It was crude and overly specific to what I was trying to get it to do, but it worked well enough to convince me that someone smarter than me could make a capable and truly useful version of it that could actually impact the industry meaningfully.
- the_newest 3y agoWhile impressive, the demo on UpWork didn’t even come close to fulfilling the job requirements. The job asked for instructions on how to set it up on an EC2 machine. It didn’t ask to run the model, or do anything that was depicted. It makes me question the truthfulness of the other claims.
- pjmorris 3y agoFrom the graph at the end: 13.8% of issues resolved. Devin may need some additional help for awhile.
- jasfi 3y agoIs Devin a new LLM? Perhaps equiped with code and deploy plug-ins? The comparisons against other LLMs would suggest so. The real world eval benchmark puts Claude 2 way ahead of GPT-4, which doesn't sound right.
- nrub 3y agoI've seen a few suspect benchmarks for recent announcements of LLM releases. I'm sure they made an attempt at an honest benchmark, but until there's an independent assessment and benchmark (preferably multiple) you have to assume that there's bias in anyone's self published benchmarks like this. I'm guessing it's a fine-tuning of some existing LLM model or API, but this largely seems to be an "agent" and UI that includes some SWE like workflow coding to allow more complex requests to be asked than just an LLM could provide.
- MSFT_Edging 3y agoHumans seek work that provides satisfaction and meaning in their life. For every technological advancement, artisans are the first to be made obsolete. Sure we have landfills full of unworn textiles, the market says its good, but overall, we keep destroying what allows humans to seek meaning. Our governments and society have made it clear, if you don't produce value, you don't deserve dignity. We have outsourced art to computers, so people who don't understand art can have it at their fingertips. Now we're outsourcing engineering so those who don't understand it can have it done for cheap. We hear stories of those who don't understand therapy suggesting AI can be a therapist, of those who don't understand medicine suggesting AI can replace a doctor. What will be left? Where will we be? Just stranded without dignity or purpose, left to rot when we no longer produce value. I ask this question often, with multiple contexts, but to what end? Who benefits from these advancements? The CEO and shareholders, sure, but just because something can be found for cheaper, doesn't mean it improves lives. Our clothes barely last a year, our shoes fall apart. Our devices come with pre-destined expiration dates. Where will we be in the future? Those born into money can continue passing it around, a cargo cult for the numbers going up. But what about everyone else?
- ardaoweo 3y agoIf we got universal basic income, people could do whatever they want. I for one would be content spending my time gardening and trekking in the nature. I despise office work and do it only for money. It's forcing the rich to give us UBI that is the problem.
- MSFT_Edging 3y agoSure but UBI would never afford a garden. It would be bare minimum to survive, if that. We'd need a complete restructuring of society to even approach dignity via UBI.
- AndrewKemendo 3y agoRight! So lets get after it Start a local cooperative and build value from the bottom up
- dukeyukey 3y agoTechnological unemployment and doomerism aside, I think there's a big difference here - in the past, you've needed lots of capital to invest in those labour-saving devices. A farm labourer couldn't buy a tractor, a dockerworker can't buy a crane. But a software engineer absolutely can buy access to AI services. I have no idea how this will end up, but it'll be different to before.
- ThalesX 3y agoAs a developer but also product person, I keep trying to use AI to code for me. I keep failing, because of context length, because of shit output from the model, because of lack of any kind of architecture etc etc etc. I'm probably dumb as hell, because I just can't get it to do anything remotely useful, more than helping me with leetcode. Just yesterday I tried to feed it a simple HTML page to extract a selector, I tried it with GPT-4-turbo, I tried it with Claude, I tried it with Groq, I tried it with a local LLama2 model with 128k context window. None of them worked. This is a task that while annoying, I do in about 10 seconds. Sure, I'm open to the possibility that in the next 2 - 3 days up to a couple of years, I'll no longer do manual coding. But honestly. After so much hype, I'm starting to grow a bit irritated with the hype. Just give me a product that works as advertised and I'll throw money your way because I have a lot more ideas than I have code throughoutput!
- CipherThrowaway 3y agoDitto. I started out excited about LLMs and eager to use them everywhere, but have become steadily disillusioned as I have tried to apply them to daily tasks, and seen others try and fail in the same way. Honestly, LLMs can't even get language right. They produce generic, amateurish copy that reads like it's written by committee. GPT can't perform to the level of a middle market copywriter or content marketer. I am convinced that people who think LLMs can write have simply not understood what professional writers do. For me the "plateau of productivity" after the disillusionment has been using LLMs a bit like search engines. Quick standalone summaries, snippets or thoughts. A nice day-to-day productivity boost, but nothing that's going to allow me to work less hard.
- famouswaffles 3y ago>GPT can't perform to the level of a middle market copywriter or content marketer. I am convinced that people who think LLMs can write have simply not understood what professional writers do. GPT's rigid "robot butler" style is not "just how LLMs write". OpenAI deliberately tuned it to sound that way. Even much weaker models that aren't tuned to write in a particular way can easily pass for human writing.
- nsypteras 3y agoClearly an extremely impressive demo and congrats on the launch. I do wonder how often the bugs Devin encounters will be solvable from the simple fixes that were demonstrated. For instance, I notice in the first demo Devin hits a KeyError and decides to resolve it by wrapping the code in a try-catch. While this will get the code to run, I immediately imagined cases where it's not actually an ideal solution (maybe it's a KeyError because the blog post Devin read is incorrect or out of date and Devin should actually be referencing a different key altogether or a different API). Can Devin "back up" at this point and implement a fix further back in its "decision tree" (e.g. use a different API endpoint) or can it only come up with fixes for the specific problem it's encountering at this moment (catch the KeyError and return None)?
- mikebelanger 3y agoYeah that was my question too. Its one thing to know the most simple fix for a KeyError issue, its another to understand that its the result of not assigning the proper key in some other part of the code, or like you said, maybe it called the wrong API endpoint and passing that into the dictionary. Somewhat related: is anyone else not really impressed by Devin fixing the errors that are very preventable with a stricter language like Rust? The demo shows Devin coding in both Python in Rust, but I consider the latter being way less energy intensive in terms of maintenance. Then again, exhaustive pattern matching and strict typing won't get you lots of VC dollars these days.
- mellosouls 3y agoLooks interesting but claiming "first" seems pretty off, there have been others like Sweep featured here before. https://news.ycombinator.com/item?id=36987454 https://news.ycombinator.com/item?id=36987454 Sweep is an open-source AI-powered junior developer https://sweep.dev/ https://sweep.dev/
- singularity2001 3y agoVery interesting and honest description of the difficulties and solutions sweep.dev encountered: https://docs.sweep.dev/blogs/gpt-4-modification https://docs.sweep.dev/blogs/gpt-4-modification
- MichaelRazum 3y agoThis is awesome to bootstrap some ideas. The question is can it work with (large) existing code bases or modify it's own code. Guess a good test would be, can it reproduce Devin;)
- cp9 3y agosorry, but no automated bullshit machine is going to do my job.
- senko 3y agoAs someone who works in this space (https://pythagora.ai https://pythagora.ai), I welcome new entrants to this niche. Currently, mainstream AI usage in coding is at the level of assistants and glorified autocomplete. Which is great (I use GitHub Copilot daily), but for us working in the space it's obvious that the impact will be much larger. Besides us (Pythagora), there's also Sweep (mentioned by others in the comments) and GPT Engineer who are tackling the same problem, I believe each with a slightly different angle. Our thesis is that human in the loop is key. In coding, you can think of LLMs as a very eager junior developer who can easily read StackOverflow but doesn't really think twice before jumping to implementation. With guidance (a LOT in terms of internal prompts, and some by human) it can achieve spectacular results.
- jprete 3y agoWhile you sound reasonable, I can't tell the difference between an honest opinion and a sales pitch here.
- senko 3y agoI am making a sales pitch on behalf of all the projects I mentioned (not just the one I'm involved with). I see LLMs failing at coding daily (one of the "perks" of working in the space), and I'm incredibly bullish on this approach. And I don't think it'll replace humans or junior engineers. As programmers, we've been "replacing" ourselves since the days of assembler that replaced direct machine coding. This is just another iteration of it. (if you do want a sales pitch, here's one: https://twitter.com/senkorasic/status/1765769482985722267 https://twitter.com/senkorasic/status/1765769482985722267 )
- starbugs 3y ago> And I don't think it'll replace humans or junior engineers. As programmers, we've been "replacing" ourselves since the days of assembler that replaced direct machine coding. This is just another iteration of it. Assembler replacing direct machine coding or C replacing assembler was a new higher level ruleset replacing an existing lower level ruleset. But that change was largely static, meaning that you could learn a concept of a mapping of C syntax -> generated assembler and rely on it being predictive of the outcome of your program with relative ease. This is what enabled you to "forget about" assembler and move on with C only. With libraries and frameworks, we introduced another level of abstraction that introduced an element of dynamism since the underlying framework could change at any time. This caused a lot of frustration with many developers already. Now, with AI, another layer of dynamism and an additional layer of "lack of precision" is introduced. What the machine figures out as a solution may change in each run, the AI itself may change, the libraries, frameworks and programming language that it uses may change - all at the same time. This raises the question how any human being should afford the time to train themselves in all the underlying components and the AI so as to be a meaningful expert in the field. Given the speed of development in the AI field, this seems like a losing proposition for anyone to invest their time in. Much different from the direct machine coding to assembler analogy that you mentioned initially.
- goat_whisperer 3y agoPeople who try to draw historical analogies to AI replacing humans say things like: "cars replaced horse drawn carriages. But we managed to adapt to that, the carriage drivers got new jobs." My dudes. We are the HORSES being replaced in this scenario.
- wetmore 3y agoI don't get your pessimism, after we are replaced we can all work at the glue factory :)
- zeroonetwothree 3y agoHow about tractors replaced humans plowing fields? Or literally any other example of automation...
- ij09j901023123 3y agoProgrammers will be worse than fast food at this point. Good luck future CS grads, you're gonna need it
- xyst 3y agoNow I can farm out scut work to Devin lol.
- gnarcoregrizz 3y agoYet again, bad time to be on the labor side of the equation, great time to be a capitalist. For us laborers, if I had to choose from a list of fields to go into, anything creative would be low on the list. 'Prompt Engineer' will be the only one left. UBI is a pipe dream... it's not happening. The wealth and means of production won't be shared in any meaningful capacity. Wealth inequality can get a whole lot worse.
- nprateem 3y agoThe measure of bullshit in this field is promoting the term 'prompt engineering'. It's prompt futzing or prompt fiddling. There's no engineering involved.
- gerash 3y agowe still don't have agents that can do simple things like: find a funny photo of my dog in my phone and post it as a story on instagram with 100% reliability. I would wait for that to happen first before thinking there can be an autonomous software engineer
- ukuina 3y agoThis is where Large Action Models will shine.
- devinegan 3y agoHave I been replaced? AI coming for my job and now my name!
- Buttons840 3y agoMe too, let's find an Alexa support group or something. I always had some sympathy for people whose name becomes a product, but it was surreal to see the headline and realize it had happened to me. At least I'm not named Karen. I'll think twice about how I use people's names in the future. Maybe a silver lining is my name was attached to a clean and upstanding product. For the rest of you, maybe your name will be associated with the hottest new erotic fiction AI sometime soon.
- pushedx 3y agoScott Wu! I met Scott at a competitive programming event a few years back. He is one of a very small group of people (going back to 1989) to get a perfect raw score at the IoI, the olympiad for competitive programming. https://stats.ioinformatics.org/people/2686 https://stats.ioinformatics.org/people/2686 Glad to see that he's putting his (unbelievable) talents to use. To give you a sense, at the event where I met him, he solved 6 problems equivalent to Leetcode medium-to-hard problems in under 15 minutes (total), including reading the problems, implementing input parsing, debugging, and submitting the solutions.
- gardenhedge 3y agoSounds like he's talented. Isn't Devin "just" a AI wrapper tool? Devin's play is that it will be the first comprehensive option available but it will soon be eaten by OpenAI, Microsoft, Google and countless others.
- gitfan86 3y agoYes, but AGI will first emerge from keeping state between calls to multiple models and assessing how closely they resemble humans intelligence, and using a loop to keep it going and updating the state. Which is what they are basically doing here
- zeroonetwothree 3y agoI've participated in programming olympiads and worked with people of varying skill and I would say that overall the correlation between competitive programming and software engineering in a business environment is probably like 0.2 or less.
- asteroidz 3y ago> 0.2 or less I find that questionable. What does "software engineering in a business environment" require that a competent competitive programmer couldn't also learn?
- 3y ago
- aster0id 3y agoI have a few years of experience in backend development, and I have realized that LLMs are incredible productivity boosts for generating code only if you know the underlying libraries/frameworks/languages very well. You can then prompt it with very specific instructions and it can go do that. Helps with the typing, but that's pretty much all. I still have to know everything and it can definitely not do everything on autopilot. I would be surprised if this product can do any real work.
- smith7018 3y agoI dunno, I have an extreme command of my platform's framework and I'd guess that 85% of the time I've asked GPT-4 for help has been a waste of time. I think it's been most helpful in regards to writing regexes but beyond that, it hallucinates correct-sounding methods all the time which leads to _a_lot_ of wasted debug time before eventually getting to the right answer by Googling what it meant or by manually rewriting large portions of what it meant to do. It's funny how a year ago I was really excited for how AI can help my everyday coding while fearful that it would replace me. Now I'm not really sure either will happen in the short term.
- syedmsawaid 3y agoIs it built with pre-existing LLMs or did they created one from the ground up? With 21 million Seed A funding, an LLM powerful than GPT4 seems impossible. What am I missing?
- martinesko36 3y agoYeah seems like this could be replicated by another AI dev fairly easily.
- epolanski 3y agoI really don't like these announcements with invitation lists. Just let me try the goddamn product. By the time you let me in, I don't care anymore or another competitor catched my attention already. Neon, the Postgres as a service put me in such a long wait list that by the time they invited me in, I was already on a completely different solution (and was happy).
- huimang 3y agoWhen you have software in prod failing because it was built by shoddy "AI" and people who copy/paste because they don't know any better, and you need a fix, give me a ring. I have tried using GPT4 & gemini extensively, and the amount of bullshit generated makes it unreliable if you don't already know the domain. These tools lack the critical stuff (being context-aware), and just make up libraries and APIs. Yet you can't be sure when it's bullshitting or not, making it an exercise in frustration for anything that's not trivial. Save your money and buy an o'reilly subscription.
- crucialfelix 3y agoI have in my codebase several really long django views files (3k lines!). They were written in a poor fashion with many nested if statements for parsing and error handling. On a one by one basis I can use VSCode github copilot to rewrite each one the way I want it. What I want to do is iterate through all functions in the files and do each one of them. I know we are getting there, but does anybody know how that can be done right now?
- elietoubi 3y agoHave you tried cursor.sh Not affiliated with them but it's actually pretty incredible for long context
- Eager 3y agoGive Claude 3 Opus a shot maybe. One of the reasons I have stayed well clear of the IDE tools is they force me to use their own model. While they might be convenient it means I can't switch to whatever the SOTA model of the day is at the drop of a hat. Opus is awesome and well worth a shot.
- dakiol 3y agoDon't get it. If we have this amazing AI why don't we make good use of it? 90% of my job is not to write code (as a senior software engineer), is to: - deobfuscate complex requirements into well divided chunks - find gaps or holes in requirements so that I have to write the minimal amount of code - understand codebases so that the implementation fits nicely I don't need an "AI software engineer", I need an "AI people person who gives me well defined tasks". Now sure, if you combine those two kinds of AIs I could perhaps become irrelevant.
- gardenhedge 3y agoThe problem is getting enough information on requirements to even break them down :)
- random_cynic 3y agoYes, the horse carriage drivers had similar lines of thoughts when they saw first gen automobiles.
- earwin 3y agoYour point being? The horses were replaced alright, carriage drivers are thriving ever since.
- random_cynic 3y agoLmao what are you talking/coping about? What happened to carriage drivers is pretty well-documented. Maybe ask one of these AI chatbots, they will summarize it for you.
- zeroonetwothree 3y agoThey drive for Uber now
- Draiken 3y ago
- HarHarVeryFunny 3y agoLet's get realistic here - I just beat GPT-4 at tic tac toe, since it failed to block my 2/3 complete winning line ... Sure, one day we'll have AGI, and one day AGI will replace many jobs that can be done in front of a computer. In the meantime, SOTA AI appears to be an airline chatbot that gets the company sued for lying to the customer. This is just basic question answering, and it can't even get that right. Would you trust it to write the autopilot code to fly the airplane? Maybe to write a tiny bit of it - just code up one function, perhaps? I sure as hell wouldn't, and when it can be trusted to write one function that meets requirements and has no bugs, it's still going to be a LONG way before it can replace the job of the developers who were given a task of "write us an autopilot".
- sohzm 3y agoWell, I'll bring the perspective of a recent college graduate from India. Most students who started jobs in software aren't writing software for aeroplanes. They're writing crud apps, doing third-party API integrations or doing basic debugging. I worry that although not some specialized software devs but a lot will still have problems due to stuff like this. I'm not talking today, but say 2-3 years down the line, who will have an intern when you can get an AI intern that can perform as the top percentile and comes at $20 per month subscription?
- HarHarVeryFunny 3y agoCertainly some software jobs are easier than others, and anything that is more just coding (e.g. crud apps) than designing would be easier to automate. It's interesting though why this hasn't been done a long time ago? I remember a piece of software called "The Last One" from the 1980's that was meant to automate creation of simple business apps like these, and here we are 40 years later with people still creating "no/low code" solutions, and people still manually writing crud apps! Why? I don't think AI will replace jobs until it can do the entire job (full lifecycle from requirements to bug fixes, etc) and be interacted with in the same way a boss or team lead could interact with a developer. If you still need a person in the loop then it's not a person replacement - it's a productivity tool.
- 3y ago
- preommr 3y agoWe're not that far from a major turning point. Currently these models don't provide an adequate enough confidence measure that prevents them from maximizing their potential. In the next few years we're going to reach a point where models will be able to tell if something is possible and avoid hallucinating, guaranteeing much better correctness. Something like that would be absolutely killer. If you add on a top-down approach using a framework, such that it can architect a system down into small individual components, then that's a recipe for a really great workflow. The models we have now really shine in doing automated unit tests, and small bits of code to avoid limits with context size. Making the interfaces obvious enough, and being able to glue things together using obvious connections seems very possible. I really do think that in the next few years we're going to see one of these tools really do well.
- soacm 3y agoNo one is taking in consideration output decay? What data are these tools going to use 10 years from now? I am not an ML engineer but these models are not able to reason and are just as good as the training set. Please debate on these points.
- Oras 3y agoFrom their twitter: > When evaluated on the SWE-Bench benchmark, which asks an AI to resolve GitHub issues found in real-world open-source projects, Devin correctly resolves 13.86% of the issues unassisted, far exceeding the previous state-of-the-art model performance of 1.96% unassisted and 4.80% assisted. While it is a progress, its far away from being useful to be a software engineer.
- typon 3y ago13% unassisted is crazy. That's probably half the performance of an intern that costs ~100k/year.
- pstorm 3y agoWhat percent does an average junior engineer solve? If it is even close, these models can be run all day and night for cheaper than one yearly SWE salary.
- swatcoder 3y agoJuniors already have bad and sometimes even negative ROI but today's working junior is the trusted engineer of tomorrow and the senior of the day after that. The problems they work on impart the knowledge and instincts that advance them through towards mastery and real value. Budget-myopic executives already tried transfering that work to cheaper labor markets, but it worked much less than they expected and most ended up with unmaintainable software and loss of any hope for an actual engineering advantage against competitors. There's nothing new here. There will be organizations that find a good and smart use for fully automated code generation, just like there is for outsourcing/offshoring, but it's not a universal win to just go with what's "cheaper" and organizations that don't look at the big picture are (as usual) trading short-term accounting gains for long-term value erosion.
- paradite 3y agoFor something that you can download and try right now, and actually works for daily coding tasks, you can try my desktop app 16x Prompt. https://prompt.16x.engineer/ https://prompt.16x.engineer/ It's not 100% automated but saves a lot of time spent on writing code. It works by composing prompts from tasks instructions, source code context and formatting instructions, resulting in high quality prompts that can be fed into LLMs to generate high quality code.
- bachittle 3y agoI recommend looking at swe-bench to get an idea as to what breakthroughs this product accomplishes: https://www.swebench.com/ https://www.swebench.com/. They claim to have tested SOTA models like GPT-4 and Claude 2 (I would like to see it tested on Claude 3 Opus) and their score is 13.86% as opposed to 4.80% for Claude 2. This benchmark is for solving real-world GitHub issues. So for those claiming that they tried models in the past and it didn't work for their use case, maybe this one will be better?
- singularity2001 3y agoInteresting: The last demo on the blog took 2.5h to complete: https://www.cognition-labs.com/blog https://www.cognition-labs.com/blog https://www.youtube.com/watch?v=UTS2Hz96HYQ https://www.youtube.com/watch?v=UTS2Hz96HYQ "Devin's Upwork Side Hustle" I wonder how much time of this was consumed by manually directing Devin into the right direction, manually fixing and undoing the mess Devin produced and watching Devin burn through $$$. As others said, being completely non-transparent about this burns a bit of trust, but I'd really like to know where we are right now. Since Devin is currently "invite only demos", a more realistic peek into the state of the art can be seen here: https://docs.sweep.dev/blogs/gpt-4-modification https://docs.sweep.dev/blogs/gpt-4-modification My gut feeling (and limited experience): gpt-4 and other models are not quite there yet, but whoever prepares for the next generation of models now will eventually win big times. Or be replaced by simpler approaches.
- joshuahutt 3y agoWe're solving the wrong problem. People trying to use cars to pull horse carts are doomed to fail. Trying to use AI to build the software of yesterday is a waste of time.
- singularity2001 3y agoSo what should the AI build instead? Specifically with regard to UI. I don't want my banking app to run on a bunch of non-deterministic prompts.
- sergiotapia 3y agoAkin to alchemy, spring up the UI to solve the user's problem. When timelines are shortened from weeks to hours, what can we build? Can a user just talk to a computer and solve their problem regardless of the platform?
- joshuahutt 3y agoExactly my point. Why do I need a custom UI for every LOB task under the sun? Just let me use a common interface to address all manner of uninteresting data problems. The UI goes away, or fades into the background, and the focus rests solely on the information I need, and the decisions I make, which I can dive deeper into with a focused AI companion. Seems like a no-brainer. Maybe folks LIKE clicking on buttons and going through 10-step procedures to get tasks done. Some mice like the maze more than the cheese, I guess.
- m3kw9 3y agoUntil you can point out via video what is the issue(“see? Here it flickers a bit and here needs centering” or when you talk to the “swe agent” and say we need this feature taken out for now, and later you ask it to put the feature back in and it remembers it had code implemented at GitHub commit id xxyyzz, you really can’t call this a software engineer
- globular-toast 3y agoI guess one good thing is proprietary software is dead. When are we getting the 100% compatible free version of Windows?
- meindnoch 3y agoAI replacing one of the last well-paid jobs on the planet is a good thing. Large-scale societal changes are triggered when a critical number of haves turn into have-nots. I would recommend junior engineers to study Nechayev and Bakunin instead of the latest React flavor. Those will have a better ROI in the coming years.
- RyEgswuCsn 3y agoIf you need AI to help you program an algorithm, then you shouldn't be using it because you can't tell if AI's solution is correct. If you can tell if a solution is correct or not --- well, then you don't need to have AI write it for you. I think AI programming can only work when the industry begin to treat "almost working" systems backed by human customer service as acceptable.
- sailingparrot 3y ago> If you can tell if a solution is correct or not --- well, then you don't need to have AI write it for you. Did you just solve P=NP? Many things are trivial to verify, but hard/time consuming to code up. You probably shouldn't rely on this to write critical software, no matter the amount of manual QA you throw at it afterwards, but there is an abundance of non-critical use cases where you can quickly check if a solution is good enough for what you care about.
- RyEgswuCsn 3y agoWhat I meant to say is that most people can only verify an algorithm is correct if they already know the correct solution. If they already know the answer then it’s probably more efficient if they write it themselves rather than having AI produce a potentially difficult to verify answer and try to verify it.
- zeroonetwothree 3y agoMost problems in "P" are not "trivial".
- rewgs 3y agoThis. And at a certain point, a prompt might become so specific that you might as well just write the code yourself. After all, a prompt is instructions for a computer, as is code.
- ellis0n 3y agoI wonder how Davin will deal with issues that have remained unfixed for decades
- PodgieTar 3y agoI must say, I'm not HUGELY impressed with a website that lets me, unauthenticated, upload files of an arbitrary size. Just posted a 500mb dmg file to their server. If anyone is practicing for their B1 Dutch exam, feel free to use this link to get the practice paper. https://usacognition--serve-s3-files.modal.run/attachments/460be415-1283-4963-9a52-931ad509afa4/2020%20Lezen%20I%20openbaar%20examen%20tekstboekje%20(digitaal).pdf https://usacognition--serve-s3-files.modal.run/attachments/4...
- 1231232131231 3y agoLooks like they deleted it and restricted file uploads :/
- samstave 3y agoThe first rule of goofy-bug-found on demo is you dont talk about goofy-bug-found on demo
- deleted 3y ago[deleted]
- cxmcc 3y agotime to start writing some cryptic code that AI won't be able to understand
- ramoz 3y agoBearish. These types of tools/agents-chaining will be irrelevant due to lackluster capability until AGI is achieved. At which point, the basis for creating these types of tools/agents will be defunct.
- ein0p 3y agoI know it’s a rigged demo because they pretend AI was able to figure out their broken CUDA situation. :-)
- YeGoblynQueenne 3y ago>> With our advances in long-term reasoning and planning, Devin can plan and execute complex engineering tasks requiring thousands of decisions. They'd better have really advanced reasoning and planning capabilities way beyond everything that anyone else knows how to do with LLMs. There's a growing body of literature that leaves no doubt that LLMs can't reason and can't plan. For a quick summary of some such results see: https://arxiv.org/pdf/2403.04121.pdf https://arxiv.org/pdf/2403.04121.pdf
- E_Bfx 3y ago'plan' is an ambigous word. 'plan' for Devin means break a complexe task in subtask and then generate code. Whereas 'plan' in the paper you link is never define, so I am not sure what his author want to demonstrate, that LLM doesn't have free will ? that LLM are not universal problem solve ? It is a quite confuse paper.
- YeGoblynQueenne 3y agoPlans and planning have precise formal definitions, probably not included in the article I linked because it is a short commentary article. See e.g. the following for the standard definitions: http://www.spacebook-project.eu/pubs/CogSci-13.pdf http://www.spacebook-project.eu/pubs/CogSci-13.pdf
- symlinkk 3y agoIn the video he was having a chat conversation with Devin the whole time, it’s not like Devin did this completely on its own.
- LZ_Khan 3y agoHey! Stop taking our jobs! Side note: I'm kind of offended that something called 'Devin' is going to take my job. If you're going to replace me at least let me keep my dignity by naming it something cool like 'Sora'
- pedalpete 3y agoI'd really like it if Cognition Labs would put the resulting code from the demo into an open-source repository so we could examine it directly. When I was using chatGPT to help guide me through some coding tasks, I'd find it could create somewhat useful code, but where it fell down was that it would put things into variables which would be better put into a class. It is this structuring of a complete system which is important for any real software engineering, rather than just writing code.
- senko 3y ago> I'd really like it if Cognition Labs would put the resulting code from the demo into an open-source repository so we could examine it directly. Yup. It's hard to evaluate things based on the demo. We're building something similar (with an open source core), and publish our examples for everyone to check out, warts and all: https://www.pythagora.ai/examples https://www.pythagora.ai/examples
- swax 3y agoI've been working on something similar, here's one of their same tests where the AI learns how to make a hidden text image. https://www.youtube.com/watch?v=dHlv7Jl3SFI https://www.youtube.com/watch?v=dHlv7Jl3SFI The real problem is coherence (logic and consistency over time) which is what these wrappers try to address. I believe AI could probably be trained to be a lot more coherent out of the box.. working with minimal wrapping.. that is the AI I worry about.
- adabaed 3y agoYou are overreacting. The moment AI can completely replace SWE, our problem won't be having "jobs".
- matthewsinclair 3y agoWe’re still at the “rhyming not reasoning” phase of LLMs. The question of whether we move past rhyming and onto reasoning is a good one, and I’m not sure what I think about it. But I am pretty sure that coding is a lot more like reasoning than it is like rhyming, at least for de novo problems above a certain level of complexity (intellectual challenge) and complication (moving parts). I remain open minded about what’s next and at the rate things are changing, I wouldn’t rule anything out a priori for now.
- devinthenai 3y agoThere's also the alternative: Devin, the NAI. https://docs.google.com/document/d/1byJgu1G_M58QVWmpZeEDthyAB8Bq752RgL_gF7hEJQc/edit https://docs.google.com/document/d/1byJgu1G_M58QVWmpZeEDthyA...
- Havoc 3y agoSurprised how calm and underwhelmed comments are. Sure it is no senior architect but the trajectory is insane. Wasn’t that long ago that LLMs barely managed coherent poems. Now it’s troubleshooting code problems on its own? Sure it’s just a gpt4 wrapper but that implies the same can be done with gpt5 and six etc. Project it forward and that does actually become non trivial
- plutoh28 3y agoThis assumes that AI will continue to improve at the same pace. GPT3 released November 2022 and yes, there have been staggering improvements since then. My intuition is that we’re beginning to reach the bounds of what we can do with LLMs. All we’ve been doing is applying the technology to other mediums (images, video, code) and yeah we’ve been seeing some really awesome/terrifying results. It all looks like it’s 80% of the way there, from which the first thought is that in a couple of years it’ll be 100% of the way there and we’ll all be jobless. This reminds me of the full self driving hype. The demos were so exciting and all that was left was getting these last couple edge cases worked out and then we’d all be in self-driving taxi pods….
- devinthenai 3y agoAlso see Devin the NAI for an old school alternative: https://docs.google.com/document/d/1byJgu1G_M58QVWmpZeEDthyAB8Bq752RgL_gF7hEJQc/edit?usp=drivesdk https://docs.google.com/document/d/1byJgu1G_M58QVWmpZeEDthyA...
- __lbracket__ 3y agoI love the collective pant shitting in this thread.
- mlsu 3y agoAfter devin "figures out" 10 issues, what does the code look like? Those are the easy ones, and if you haven't fixed them cleanly, the next 10 will be more difficult to solve, for human and for robot. Now do this for several years. Can devin create its own bug reports and issues? It better be able to! I'm curious what a large, mature codebase, with complex internals and legacy code looks like after you sick devin on it. Not pretty I suspect. In fact, I think it will become so difficult to fix that nobody -- neither human nor devin -- will be able to clean up the mess. By sheer volume, a broken ball of unfixable spaghetti. I would be immensely pissed off if someone did this to an open source project of mine, or even to a closed-source codebase I'm working on. Not only would it not be useful, it would be moving backwards. Creating an icky vomit mess that we will probably have to spend years cleaning up after bug reports and complaints from customers begin mounting, and competitors can iterate faster. Does that sound like something you want to deal with in your software business?
- mlsu 3y agoIt comes across harshly. The demo is genuinely technically impressive. But may I suggest: apply devin to your own codebase first -- then decide if it's something you want to gift to the rest of us.
- senera 3y ago[flagged]
- Badone404 3y ago[flagged]
- Badone404 3y ago[flagged]
- Angelheller8 3y ago[flagged]
- claire_wu 3y ago[dead]
- playmkr 3y ago- first AI developer - raised $21 million - uses google forms for the onboarding ok
- zackqixun 3y ago[flagged]
- zackqixun 3y ago[flagged]
- aldernet 3y ago[flagged]
- 928711801 3y ago[flagged]
- 928711801 3y ago[flagged]
- leoChat 3y ago[dead]
- leoChat 3y ago[dead]
- isodev 3y agoI'm totally adding "rescue and recovery of projects botched by AI" to my list of services. One thing is certain, it's not going to be cheap.
- DrAgOn200233 3y agoI believe that hosting DEVIN will cost much more GPU time hosting a regular LLM. By inspecting the videos in Cognition Lab's official website, I noticed that DEVIN can take more than one hour to do one step, which is more than an hour of GPU usage. When using GPT-4, we usually get output within 30 seconds, which is less than a minute of GPU usage. In addition, when using GPT-4, I use it only when I have new thoughts, so the GPU occupancy rate is low. I probably use less than 5 hours of GPU time each month. DEVIN is sort of like an intern working for you, so you would probably at least make it work 40hrs/week. These difference in GPU usage would probably make DEVIN 10 times more expensive for the business model to be profitable, that is, if they are using the subscriber business model like GPT-4. I don't think there are any other viable business model for DEVIN - for sure it cannot replace or even reduce the number of human programmer due to LLM's unreliable nature and the necessity of code verification.
- erickmunene 3y agoWow, this is incredible news! Congratulations to the team behind Devin, the first AI engineer! This is a monumental leap forward in technology and innovation. I'm absolutely thrilled to see how Devin will revolutionize the field of engineering. As someone passionate about the potential of AI in tech, I can't wait to see what amazing feats Devin will accomplish. And who knows, maybe one day, companies like Munesoft Technologies will reach similar heights with their own AI-driven advancements. Here's to a future filled with endless possibilities! #DevinAI
- sinopvalisi 3y agothis comment sounds weird and so AI generated.
- deleted 3y ago[deleted]
- MeinanGou 3y ago[flagged]
- MeinanGou 3y ago[flagged]
- fwqAi 3y ago[flagged]
- rohandakua 3y agohey , I am a newbie in field of AI , i want to know about the nearby future of AI . after devin I am little ! or rather I should say deeply terrified about the future (5 to 8 years) of software eng. can someone explain ?
- asasasa123 3y agoWrite a demo of the Milvus vector database
- emawa 3y agoDevin can make a app with end points an join fronsidento backside
- stephen1314 3y ago[flagged]
- stephen1314 3y ago[flagged]
- znhzlq 3y ago[flagged]
- znhzlq 3y ago[flagged]
- deleted 3y ago[deleted]
- shreshth398495 3y agohow will devin pass the CAPTCHA test when it encounters an error while coding? most websites block such automated tools? isn't it?
- zijie-tian 3y ago[dead]
- punkbit 3y agoGood luck finding someone to maintain your fancy AI generated a$$ looking app.
- BIGBOOTYAI 3y ago[flagged]
- BIGBOOTYAI 3y ago[flagged]
- andythedev 3y agoare we not concerned that even though, yes, devin is only solving 13% of issues - it is also an ML model. it is going to learn, potentially very quickly.
- joeevans1000 3y agoThe lack of attention this development is getting on here is astounding. Other developers I am talking to about it brush it off, then change topic. So many comments about how insufficient the tool is. Our heads are really in the sand, I'm afraid.
- deleted 3y ago[deleted]
- pankajdoharey 3y agoSo why are they hiring? https://jobs.ashbyhq.com/cognition https://jobs.ashbyhq.com/cognition cant they just use "Devin" ?
- StickyRibbs 3y agoI've worked on very complex systems - The disney streaming platform (before it was disney), live video streaming systems, banking transaction systems, your run of the mill crud software with kafka clusters piping mind numbing amounts of data, netflix and a few other large engineering heavy companies. No engineering company worth their weight is going to build a world class technology business purely with generative AI in its current state. The risk in doing so currently is total and utter failure. I have a very hard time believing we're any where near that capability. Maybe your mom and pop startup could hire a prompt engineer to build a website and simple tool but we have yet to see those exercises surfaced to the mainstream; it's purely speculative. I say, rest easy programmers. Your careers will be enriched more than axed with generative AI as a support tool for many years to come. Also, if anyone who works in this field has a strong opposing belief, then consider OpenAI engineers are programming themselves out of a job which obviously is not the case.
- deleted 3y ago[deleted]
- zoomin 3y agoAn eval on 25% of the eval dataset is fishy? Why not 100%? Are they training with the eval set? Also that dataset is almost all python.
- plinkplink 3y agoThere have been many tools that were sold as developer replacements over the years. - Microsoft FrontPage - Adobe Dreamweaver - A litany of glorified WYSIWYG editors - WordPress - Wix - Power Apps/SharePoint ...and so on. Business owners have been getting sexually aroused at the prospect of taking a developer's salary and putting it in their own pockets for decades. Each iteration of this wet dream has only locked businesses into "low code" systems that require even more, highly-specialized developers to operate. Right now, and probably for a while, Devon, et al. is on par with the drag-and-drop automagical app building snake oil stuff. LLMs are useful to help developers be more productive, which does translate to lay-offs, but until someone creates an AI that can translate the absolute fevered gibberish that comes out of business people's heads into a profitable piece of software, this is just MS FrontPage v100.0. Just like an entire industry sprang up around fixing WordPress websites that business owners thought they could do themselves, pretty soon we'll start seeing job postings for AI-Generated Spaghetti Unravellers. I'm (half seriously) imagining a future were software engineers are mostly consultants that show up and talk with business folks, then talk with the local robot, and get the project to actually work. Bill $1k per hour.