9 ms·
Tried it a few weeks ago for a task (had a few dozen files in an open source repo I wanted to write tests for in a similar way to each other). I gave it one ex
by yeldarb 2y ago
Tried it a few weeks ago for a task (had a few dozen files in an open source repo I wanted to write tests for in a similar way to each other).
I gave it one example and then asked it to do the work for the other files.
It was able to do about half the files correctly. But it ended up taking an hour, costing >$50 in OpenAI credits, and took me longer to debug, fix, and verify the work than it would have to do the work manually.
My take: good glimpse of the future after a few more Moore’s Law doublings and model improvement cycles make it 10x better, 10x faster, and 10x cheaper. But probably not yet worth trying to use for real work vs playing with it for curiosity, learning, and understanding.
Edit: writing the tests in this PR given the code + one test as an example was the task: https://github.com/roboflow/inference/pull/533 https://github.com/roboflow/inference/pull/533
This commit was the manual example: https://github.com/roboflow/inference/pull/533/commits/93165a81b50b3daa55fbc3359ed6a500825266fc https://github.com/roboflow/inference/pull/533/commits/93165...
This commit adds the partially OpenDevin written ones: https://github.com/roboflow/inference/pull/533/commits/65f51ccda0024f1a7b6d817fab48560861b465e2 https://github.com/roboflow/inference/pull/533/commits/65f51...
- rbren 2y agoOpenDevin maintainer here. This is a reasonable take. I have found it immensely useful for a handful of one-off tasks, but it's not yet a mission-critical part of my workflow (the way e.g. Copilot is). Core model improvements (better, faster, cheaper) will definitely be a tailwind for us. But there are also many things we can do in the abstraction layer _above_ the LLM to drive these things forward. And there's also a lot we can do from a UX perspective (e.g. IDE integrations, better human-in-the-loop experiences, etc) So even if models never get better (doubtful!) I'd continue to watch this space--it's getting better every day.
- anotherpaulg 2y agoAs a comparison, I use aider every day to develop aider. Aider wrote 61% of the new code in its last release. It’s been averaging about 50% since the new Sonnet came out. Data and graphs about aider’s contribution to its own code base: https://aider.chat/HISTORY.html https://aider.chat/HISTORY.html
- Lerc 2y agoHow heavy are the API costs for that? For a project like yours I guess you should be given free credits. I hope that happens, but so far nobody has even given Karpathy a good standalone mic.
- anotherpaulg 2y agoNot much. I spent $25 on Anthropic in July.
- harisec 2y agoIf you use DeepSeek Coder V2 0724 (that is #2 after Claude 3.5 Sonnet on the Aider leaderboard), the costs are very, very small. https://aider.chat/2024/07/25/new-models.html https://aider.chat/2024/07/25/new-models.html
- harisec 2y agoaider is great, i also use it almost daily. thanks for writing it Paul!
- dartos 2y agoIt’d be really great to see a video or cast of you using aider to work on aider. I can’t get anything useful out of these AI tools for my tasks and I’d really like to see what someone who can does. I’d like to know if it’s me or my tasks that aren’t working for the llm.
- 2y ago
- jijji 2y agoinstead of using openAI api, can it use the locally hosted ollama http API?
- davidy123 2y agoYes. It's not really "open" if it depends on a non-libre service. To be legit, they must at least enable this experimentally.
- strangescript 2y agoGuessing you used 4o and not 4o-mini. For stuff like this you are better off letting it use mini which is practically free, and then have it double and triple check everything.
- MattDaEskimo 2y agoIt doesn't work like that. You're more likely to end up with a fractal pattern of token waste, potentially veering off into hallucinations than some actual progress by "double" or "triple checking everything".
- threeseed 2y agoThis assumes that the model knows it is wrong. It doesn't. It only knows statistically what is the most likely sequence of words to match your query. For rarer datasets e.g. I had Claude/OpenAI help out with an IntelliJ plugin it would continually invent methods for classes that never existed. And could never articulate why.
- popinman322 2y agoThis is where supporting machinery & RAG are very useful. You can auto- lint and test code before you set eyes on it, then re-run the prompt with either more context or an altered prompt. With local models there are options like steering vectors, fine-tuning, and constrained decoding as well. There's also evidence that multiple models of different lineages, when their outputs are rated and you take the best one at each input step, can surpass the performance of better models. So if one model knows something the others don't you can automatically fail over to the one that can actually handle the problem, and typically once the knowledge is in the chat the other models will pick it up. Not saying we have the solution to your specific problem in any readily available software, but that there are approaches specific to your problem that go beyond current methods.
- __loam 2y agoThis is a really complicated (and more expensive) setup that doesn't fundamentally fix any of the problems with these systems.
- threeseed 2y ago> 10x better, 10x faster, and 10x cheaper Which is the elephant in the room. There is no roadmap for any of these to happen and a strong possibility that we will start to see diminishing returns with the current LLM implementation and available datasets. At which point all of the hype and money will come out of the industry. Which in turn will cause a lull in research until the next big breakthrough and the cycle repeats.
- Sysreq2 2y agoWhile we have started seeing diminishing returns on rote data ingestion, especially with synthetic data leading to collapse, there is plenty of other work being done to suggest that the field will continue to thrive. Moore’s law isn’t going anywhere for at least a decade - so as we get more computing power, faster memory interconnects, and purpose built processors, there is no reason to suspect AI is going to stagnate. Right now the bottleneck is arguably more algorithmic than compute bound anyways. No one will ever need more than 640kb of RAM, right?
- thwarted 2y agoI feel like the GP and this response are a common exchange right before the next AI Winter hits.
- __loam 2y agohttps://cap.csail.mit.edu/death-moores-law-what-it-means-and-what-might-fill-gap-going-forward https://cap.csail.mit.edu/death-moores-law-what-it-means-and...
- threeseed 2y agoa) It's been widely acknowledged that we are approaching a limit on useful datasets. b) Synthetic data sets have been shown to not be a substitute. c) I have no idea why you are linking Moore's Law with AI. Especially when it has never applied to GPUs and we are in a situation where we have a single vendor not subject to normal competition.
- 2y ago
- __loam 2y agoStrong chance Moores law stops this decade due to the physical limits on the size of atoms lol.
- _w1tm 2y agoI’ve been hearing that for at least a decade.
- dartos 2y agoI’m hopeful that there are some possible model topologies that don’t just stack matmuls. Maybe there’s some wins to be had on the software side still.
- krageon 2y agoI've heard variations on this argument for the past two decades, and it's amusing every time.