3 ms·
Absolutely not victim blaming, but how do you use it? The biggest issue with local LLM is prompt structure prompting (like how you feed the instructions, not ho
by syntaxing 3y ago
Absolutely not victim blaming, but how do you use it? The biggest issue with local LLM is prompt structure prompting (like how you feed the instructions, not how you say something). If you even deviate a small bit from how it was trained, you’ll get terrible results. I’ve been using codellama 13b + ollama + continue and I’ll be honest, it’s almost on par with GPT-3.5 for my stuff. It’s been amazing as a pair programmer. It’s better to make draft and bounce ideas with it than to ask it to start from scratch. Long story short, try ollama + continue. If you’re using llama cpp by itself, chances are, you’ll get bad results.
- tarruda 3y agoAre you using apple silicon? How much RAM do you have, and how many tokens/second with codellama 13b?
- syntaxing 3y agoYes, 16GB of ram is needed for 13B, 32GB for 34B (both for 4bit). The first time it loads a new models takes some warm up time, I wanna say 30s? After that, the context reading and token generation are usually upward of 8 tk/s. Also, the newer and bigger the die, the faster the token generation. Like a Mac Studio would probably generate 30% or so faster than a MBP
- tarruda 3y ago8tk/s on 34b? I've managed to run Codellama instruct 13b with my laptop's RTX 3070 (8gb VRAM) at 6tk/s by offloading 27 layers into the GPU with llama.cpp I've been considering getting a macbook for running 34b+ LLM inference, but with the speed in which small LLMs are progressing, I think it is better to get a laptop with an RTX 4090 and 16gb vram. Maybe It can run 34b models by offloading layers into the GPU.
- syntaxing 3y agoI only have a 16GB computer so I can’t confirm the 34B performance. I have a 3090 with 24GB of VRAM and 34B just fits and runs above 15 tk/s. If you want a laptop and only plan for inferencing, I think a MBP would be better than a 4090 laptop.
- wahnfrieden 3y agoNo warm up if you switch to metal with no ANE on sonoma
- wokwokwok 3y ago> It’s been amazing as a pair programmer. ... > It’s better to make draft and bounce ideas with it than to ask it to start from scratch. Mm. Look, I'm going to be brutally blunt here. In the long term, chat is an AI-anti-pattern. You can't automate a prompt sequence when the Nth prompt is context dependent on the previous prompt. "Write me XX" ... "No, fix this" ... "no, more like this" ... "I get this error" ... Cool. You get a result and it works. ...but how many interactions did you do to get that? 5? How long did it take? Did you even try 'regenerate answer' and look at some variations? Are you sure the first answer it gave you was the best one? I'm pretty sure it wasn't. Anyway, ok, so now you have 50 functions you need to generate. Now you have 500. What's your plan? Same thing? There are too many human touch points. You know what AI superpower is? Automation. Repeatedly generating output, day in and day out. That's what computers are all about. Don't get me wrong; the interactive style of AI copilot is lovely too, but it's just an incremental improvement on autocomplete, and I'm not interested; I already have autocomplete. > how do you use it? 1) Every code function I want to generate, I create a scaffold that defines the exact function template, like: // Using these imports only import {x, y, z} from "./blah"; /* What does foo do... */ export function foo(a: number, b: number) { ... } Every prompt goes into a `prompts` folder. 2) I create a test harness that defines a set of unit tests that define the behaviour of foo. So, you can literally run: `npx jest ./output/foo.ts` Every prompt has a matching `tests/foo.test.ts` test file. (Yes, I know this sounds like a pain in the ass, it's less annoying when you scaffold tests out an LLM as well. It's not as bad as you might imagine once you get used to the workflow). 3) I process the prompts folder, and for every prompt generate a solution candidate: - I extract the typescript from the markdown output, save it. - I run `npx tsc --strict foo.ts --outDir dist` on it. - If it fails, run a meta 'fix this typescript with these errors' prompt over it. - I run the test suite on the result if it passes. - If the test suite passes, I save the result as a candidate solution. - If it fails, I vary the temperature and generate a new solution. - Eventually if I don't get any candidate solutions, I log an error to revisit and refine the prompt. Look, it's not magic, it's very simple: LLMs generate code. sometimes the code is good, sometimes its not... but you can generate 10 or 20 different variations and it costs literally nothing except time. You just repeat it over and over and over; and maybe run some automated fixes on the outputs. It works fine. I've made a raytracer with it, I've made a little card game with it. I'm building a website with it. Great stuff. ...if I use the openai api. Now, the openai api sucks for lots of reasons, but the big one is that when you use the real AI superpower; ie. automation, it actually starts costing you a not insignificant amount of $$$. So, I've been experimenting with using some offline models; specifically, as I said, code llama, and mistral. The best results I've had are from the q5, q6 codellama (1) 34B model, running using llama.cpp. It's just slow. So, I was experimenting with these smaller models, but... they're not that great for what I'm doing. What you're doing, is not what I'm doing, and not quite I'm trying to do. I get the "you're using it wrong" argument, yup. Fair enough. You're totally right. A lot of people get a lot of value from just having chatGPT open side-by-side with vscode. That's cool... but I'm specifically talking about my difficulties with a different use-case. [1] - https://huggingface.co/TheBloke/Phind-CodeLlama-34B-v2-GGUF https://huggingface.co/TheBloke/Phind-CodeLlama-34B-v2-GGUF
- lmeyerov 3y agoThat feels like damning with faint praise: we encourage Louie.ai users to only do GPT4+ level models for code gen related tasks. Even GPT4 has a lot of ways to go. Saying other models are only around 3.5 for this task isn't great. I'm hopeful for starcoder etc, but still not there yet afaict... Agreed on prompts. We are doing a lot to guide it, and even autorepair loops. Likewise, keeping the interaction model to generating small code likewise helps the chance of any individual step being right and repairable..
- sangnoir 3y ago> Absolutely not victim blaming, but how do you use it? It sounds like OP is trying to replace junior/mid-level SWEs with CodeLLMs where a detailed description of the desired solution goes in, and working code comes out - all hands-off. If there is ever will be a time that LLMs can consistently achieve what OP wants, there will be a reckoning in software engineering. It's not like junior SWEs aren't already having a hard time with the current hiring environment.