7 ms·
While I've not used this product, I've created somewhat similar setup using open source LLMs that runs locally. After having used it for about three months, I c
by singluere 2y ago
While I've not used this product, I've created somewhat similar setup using open source LLMs that runs locally. After having used it for about three months, I can say that debugging LLM prompts was far more annoying than debugging code. Ultimately, I ended up abandoning my setup and going in favor of writing code the good old fashioned way. YMMV
- BoorishBears 2y agoCreating this using Open Source LLMs would be like saying you tried A5 Wagyu by going to Burger King, respectfully. I think benchmarks are severely overselling what open source models are capable of compared to closed source models.
- Zambyte 2y agoI really don't think they're being over sold that much. I'm running llama 3 8b on my machine, and it feels a lot like running claude 3 haiku with a much lower context window. Quality wise it is surprisingly nice.
- BoorishBears 2y agoLlama 3 just came out so they couldn't have used it, and Claude Haiku is the smallest cheapest closed source model out there from what I've seen. Github is likely using a GPT-4 class model which is two (massive) steps up in capabilities in Anthropic's offerings alone
- Zambyte 2y agoYeah I just mentioned Llama to point out that the open weight models have been really catching up. Microsoft is almost certainly using GPT-4 given their relationship with ClosedAI, but I would definitely not put GPT-4 (nor Turbo) "two massive steps up" from Claude 3 Opus. I have access to both through Kagi, and I have found myself favoring the responses of Claude to the point where I almost never use GPT(TM) anymore.
- BoorishBears 2y agoYou're misreading in multiple ways, maybe in a rush to dunk on "Closed AI". Github Copilot is not the same as Copilot Chat which uses GPT-4, there still some uncertainty on if Copilot completions use GPT-4 as outsiders know it (and iirc they've specifically said it doesn't at some point) I also said Haiku is two massive steps behind Anthropic's offerings... which are Sonnet and Opus. Anthropic isn't any more open than OpenAI, and I personally don't attribute any sort of virtue to any major corporation, so I'll take what works best
- Zambyte 2y agoI... don't think I misread you? Maybe you didn't mean what you wrote, but what you said was: > Github is likely using a GPT-4 class model which is two (massive) steps up in capabilities in Anthropic's offerings alone Comparing GPT-4 to Anthropics offerings, which, as you say, includes Sonnet and Opus. > Anthropic isn't any more open than OpenAI, [...] so I'll take what works best I understand that, and same here. I don't prefer Claude for any reason other than the quality of its output. I just think OpenAIs name is goofy with how they actually behave, so I prefer the more accurate derivative of their name :) Regarding what model Copilot Completions is using - point taken, I have no comment on that. My original comment in this thread was only meant to point out that open weight models are getting a lot better. Not saying they're using them.
- BoorishBears 2y agoI used "in Anthropic's capabilities" intentionally: it's two steps up in what they can do from Claude Haiku
- paradite 2y agoLocally running LLMs in Apr 2024 are no where close to GPT-4 in terms of coding capabilities.
- 015a 2y agoAnd GPT-4 is nowhere close to the human brain in terms of coding capabilities, and model advancements appear to be hitting an asymptote. So...
- throwaway4aday 2y agoI don't see a flattening. I see a lot of other groups catching up to OpenAI and some even slightly surpassing them like Claude 3 Opus. I'm very interested in how Llama 3 400B turns out but my conservative prediction (backed by Meta's early evaluations) is that it will be at least as good as GPT 4. It's been a little over a year since GPT 4 was released to the public and in that time Meta and Anthropic seem to have caught up and Google would have too if they spent less time tying themselves up in knots. So OpenAI has a 1 year lead though they seem to have spent some of that time on making inference less expensive which is not a terrible choice. If they release 4.5 or 5 and it flops or isn't much better then maybe you are right but it's very premature to call the race now, maybe 2 years from now with little progress from anyone.
- 015a 2y agoI shouldn't have used the word asymptote; I should have said logarithmic. I don't doubt a best-case situation where we get a GPT-5, GPT-6, GPT-7, etc; each is more capable than the last; just that there will be more months between each, it'll be more expensive to train each, and the gain of function between each will be smaller than the previous. Let me phrase this another way: Llama 3 400B releases and it has GPT-5 level performance. Obviously; we have not seen GPT-5; so we don't have a sense of what that level of performance looks like. It might be that OpenAI simply has a one year lead, but it might also be that all these frontier model developers are stuck in the same capability swamp; and we simply don't have the compute, virgin tokens, economic incentives, algorithms, etc to push through it (yet). So, Meta pulls ahead, but we're talking about feet, not miles.
- lostintangent 2y agoI can definitely echo the challenges of debugging non-trivial LLM apps, and making sure you have the right evals to validate progress. I spent many hours optimizing Copilot Workspace, and there is definitely both an art and a science to it :) That said, I’m optimistic that tool builders can take on a lot of that responsibility, and create abstractions that allow developer to focus solely on their code, and the problem at hand.
- idan 2y agoI'm sure we'll share some of the strategies we used here in upcoming talks. It's, uh, "nontrivial". And it's not just "what text do you stick in the prompt".
- singluere 2y agoFor sure! As a user, I would love to be able to have some sort of debugger like behavior for debugging the LLM's output generation. Maybe some ability for the LLM to keep on running some tests until they pass? That sort of stuff would make me want to try this :)
- slavoglinsky 2y agosee langtail app (I am not maker)
- candiddevmike 2y agoI had ChatGPT output an algorithm implementation in Go (Shamir Secret Sharing) that I didn't want to figure out. It kinda worked, but everytime I pointed out a problem with the code it seemed more bugs were added (and I ended up hating the "Good catch!" text responses...) Eventually, figuring out why it didn't work made me have to read the algorithm spec and basically write the code from scratch, throwing away all of the ChatGPT work. Definitely took more time than doing it the "hard way".
- torginus 2y agoAn alterinative to this workflow that I find myself returning to is the good ol' nicking code from stackoverflow or Github. ChatGPT works really well because the stuff you are looking for is already written somewhere and it solves the needle-in-the-haystack problem of finding it, very well. But I often find it tends to output code that doesn't work but eerily looks like it should, whereas Github stuff tends to need a bit more wrangling but tends to work.
- briHass 2y agoThe big benefit to me with SO is that with a question with multiple answers, the top up voted question likely works, since those votes are probably people that tried it. I also like the 'well, actually' responses and follow up, because people point out performance issues or edge cases I may or may not care about. I only find current LLMs to be useful for code that I could easily write, but I am too lazy to do so. The kind of boilerplate that can be verified quickly by eye.
- fwip 2y agoOne thing that's helped me a little bit, is to open up the spec as context, and then asking the LLM to generate tests based on that spec.
- fragmede 2y agoThe skill in using an LLM currently is in getting you to where you want to be, rather than wasting time convincing the LLM to spit out exactly what you want. That means flipping between having Aider write the code and editing the code yourself, when it's clear the LLM doesn't get it, or you get it better than it does.
- whatever1 2y agoThat is my struggle as well. I need to keep pointing out issues of the llm output, until after multiple iterations it may reach the correct answer. At that point I don't feel I gained anything productivity wise. Maybe the whole point of coding with llms in 2024 is for us to train their models.
- freedomben 2y agoIndeed, and the more niche the use case, the worse it gets.
- lenerdenator 2y agoHonestly, I've found using GH CoPilot chat to be the real value add. It's amazing for rubber ducking. That being said, my employer pays for it. I am still on the fence about which LLM to subscribe to with my own money.
- mountainriver 2y agoMaybe they know something about GPT5