10 ms·
Yi-Coder: A Small but Mighty LLM for Code
- smcleod 2y agoWeird they're comparing it to really old deepseek v1 models, even v2 has been out a long time now.
- butterfly42069 2y agoMy exact thoughts, especially because DeepseekV2 is meant to be a massive improvement. It seems to be an emerging trend people should look out for that model release sheets often contain comparisons with out of date models and don't inform so much as just try to make the model look "best." It's an annoying trend. Untrustworthy metrics betray untrustworthy morals.
- bubblyworld 2y agoMy barely-informed guess is that they don't have the resources to run it (it's a 200b+ model).
- regularfry 2y agoThey could compare to DeepSeek-Coder-V2-Lite-Instruct. That's a 16B model, and it comes out at 24.3 on LiveCodeBench. Given the size delta they're respectably close - they're only just behind at 23.4. The full V2 is way ahead.
- smcleod 2y agoThat’s for the larger model, most people running it locally use the -lite model (both of which has lots of benchmarks published)
- theshrike79 2y ago> Continue pretrained on 2.4 Trillion high-quality tokens over 52 major programming languages. I'm still waiting for a model that's highly specialised for a single language only - and either a lot smaller than these jack of all trades ones or VERY good at that specific language's nuances + libraries.
- wiz21c 2y agoIf the LLM training makes the LLM generalize things between languages, then it is better to leave it like it is...
- richardw 2y agoI’d be interested to know if that trade off ends up better. There’s probably a lot of useful training that transfers well between languages, so I wouldn’t be that surprised if the extra tokens helped across all languages. I would guess a top quality single language model would need to be very well supported, eg Python or JavaScript. Not, say, Clojure.
- rfoo 2y agoAn unfortunate fact is, similar to human with infinite time, LLMs usually have better performance on your specific langauge when they are not limited to learn or over-sample one single language. Not unlike the common saying "learning to code in Haskell makes you a better C++ programmer". Of course, this is far from trivial, you don't just add more data and expect it to automatically be better for everything. So is time management for us mere mortals.
- deely3 2y ago> usually have better performance on your specific langauge when they are not limited to learn or over-sample one single language. Source? Im very curious how learning one language helps model to generate code in language with different paradigms. Java, Markdown, JSON, HTML, Fortran?
- cztomsik 2y agoI think around the BLOOM models (2022) it was found out that if you train english-only, the model performs worse than if you have even little mixture of other languages. Also, there were other papers (one epoch is all you need) where it was shown that diverse data is better than multiple epochs, and finally, there was paper (textbooks is all you need) for famous Phi model, with conclusion that high-quality data > lots of data. This by itself is not a proof for your specific question but you can extrapolate.
- mythz 2y agoClaude 3.5 Sonnet still holds the LLM crown for code which I'll use when wanting to check the output of the best LLM, however my Continue Dev, Aider and Claude Dev plugins are currently configured to use DeepSeek Coder V2 236B (and local ollama DeepSeek Coder V2 for tab completions) as it offers the best value at $0.14M/$0.28M which sits just below Claude 3.5 Sonnet on Aider's leaderboard [1] whilst being 43x cheaper. [1] https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/
- dsp_person 2y agoDeepSeek sounds really good, but the terms/privacy policy look a bit sketch (e.g. grant full license to use/reproduce inputs and outputs). Is there anywhere feasible to spin up the 240B model for a similarly cheap price in private? The following quotes from a reddit comment here https://www.reddit.com/r/LocalLLaMA/comments/1dkgjqg/comment/l9i4ujo/ https://www.reddit.com/r/LocalLLaMA/comments/1dkgjqg/comment... > under International Data Transfers (in the Privacy Policy): """ The personal information we collect from you may be stored on a server located outside of the country where you live. We store the information we collect in secure servers located in the People's Republic of China . """ > under How We Share Your Information > Our Corporate Group (in the Privacy Policy): """ The Services are supported by certain entities within our corporate group. These entities process Information You Provide, and Automatically Collected Information for us, as necessary to provide certain functions, such as storage, content delivery, security, research and development, analytics, customer and technical support, and content moderation. """ > under How We Use Your Information (in the Privacy Policy): """ Carry out data analysis, research and investigations, and test the Services to ensure its stability and security; """ > under 4.Intellectual Property (in the Terms): """ 4.3 By using our Services, you hereby grant us an unconditional, irrevocable, non-exclusive, royalty-free, sublicensable, transferable, perpetual and worldwide licence, to the extent permitted by local law, to reproduce, use, modify your Inputs and Outputs in connection with the provision of the Services. """
- yumraj 2y agoThere’s no company info on DeepSeek’s website. Looking at the above, and considering that, it seems very sketchy indeed. Maybe OK for trying out stuff, a big no no for real work.
- ziofill 2y agoAre coding LLMs trained with the help of interpreters?
- willvarfar 2y agoGoogle's Gemini does. I can't find a post that I remember Google published just after all the ChatGPT SQL generation hype happened, but it felt like they were trying to counter that hype by explaining that most complex LLM-generated code snippets won't actually run or work, and that they were putting a code-evaluation step after the LLM for Bard. (A bit like why did they never put an old fashioned rules-based grammar checker check stage in google translate results?) Fast forward to today and it seems it's a normal step for Gemini etc https://ai.google.dev/gemini-api/docs/code-execution?lang=python https://ai.google.dev/gemini-api/docs/code-execution?lang=py...
- redeyedtreefrog 2y agoThat's interesting! Where it says that is will "learn iteratively from the results until it arrives at a final output" I assume it's therefore trying multiple LLM generations until it finds one that works, which I didn't know about before. However, AFAIK it's only ever at inference time, an interpreter isn't included during LLM training? I wonder if it would be possible to fine tune a model for coding with an interpreter. Though if noone has done it yet there is presumably a good reason why not.
- littlestymaar 2y ago> Though if noone has done it yet there is presumably a good reason why not. The field is vast, moving quickly and there are more directions to explore than researchers working at top AI labs. There's lots of open doors that haven't been explored yet but that doesn't mean it's not worth it, it's just not done yet.
- Havoc 2y agoBeats deepseek 33. That’s impressive
- tuukkah 2y agoThey used DeepSeek-Coder-33B-Instruct in comparisons, while DeepSeek-Coder-v2-Instruct (236B) and -Lite-Instruct (16B) are available since a while: https://github.com/deepseek-ai/DeepSeek-Coder-v2 https://github.com/deepseek-ai/DeepSeek-Coder-v2 EDIT: Granted, Yi-Coder 9B is still smaller than any of these.
- cassianoleal 2y agoIs there an LLM that's useful for Terraform? Something that understands HCL and has been trained on the providers, I imagine.
- bloopernova 2y agoCopilot writes terraform just fine, including providers.
- cassianoleal 2y agoThanks. I should have specified, LLMs that can be run locally is what interests me.
- lasermike026 2y agoTry this, https://ollama.com/jeffrymilan/aiac https://ollama.com/jeffrymilan/aiac
- mtrovo 2y agoI'm new to this whole area and feeling a bit lost. How are people setting up these small LLMs like Yi-Coder locally for tab completion? Does it work natively on VSCode? Also for the cloud models apart from GitHub Copilot, what tools or steps are you all using to get them working on your projects? Any tips or resources would be super helpful!
- cassianoleal 2y agoYou can run this LLM on Ollama [0] and then use Continue [1] on VS Code. The setup is pretty simple: * Install Ollama (instructions for your OS on their website - for macOS, `brew install ollama`) * Download the model: `ollama pull yi-coder` * Install and configure Continue on VS Code (https://docs.continue.dev/walkthroughs/llama3.1 https://docs.continue.dev/walkthroughs/llama3.1 <- this is for Llama 3.1 but it should work by replacing the relevant bits) [0] https://ollama.com/ https://ollama.com/ [1] https://www.continue.dev/ https://www.continue.dev/
- suprjami 2y agoIf you have a project which supports OpenAI API keys, you can point it at a LocalAI instance: https://localai.io/ https://localai.io/ This is easy to get "working" but difficult to configure for specific tasks due to docs being lacking or contradictory.
- NKosmatos 2y agoIt would be good if LLMs were somehow packaged in an easy way/format for us "novice" (ok I mean lazy) users to try them out. I'm not interested so much with the response time (anyone has a couple of spare A100s?), but it would be good to be able to try out different LLMs locally.
- nusl 2y agoThis is already possible. There are various tools online you can find and use.
- hosteur 2y agoYou should try GPT4all. It seems to be exactly what you’re asking for.
- suprjami 2y agoOne Docker command if you don't mind waiting minutes for CPU-bound replies: https://localai.io/ https://localai.io/ You can also use several GPU options, but they are not as easy to get working.
- PhilippGille 2y agoWith Mozilla's llamafile you can run LLMs locally without installing anything: https://github.com/Mozilla-Ocho/llamafile https://github.com/Mozilla-Ocho/llamafile
- senko 2y agoLM Studio is pretty good: https://lmstudio.ai/ https://lmstudio.ai/
- dizhn 2y agoI understand your situation. It sounds super simple to me now but I remember having to spend at least a week trying to get the concepts and figuring out what prerequisite knowledge I would need between a continium of just using chatgpt and learning relevant vector math etc. It is much closer to the chatgpt side fortunately. I don't like ollama per se (because i can't reuse its models with other frontends due to it compressing them in its own format) but it's still a very good place to start. Any interface that lets you download models as gguf from huggingface will do just fine. Don't be turned off by the roleplaying/waifu sounding frontend names. They are all fine. This is what I mostly prefer: https://github.com/oobabooga/text-generation-webui https://github.com/oobabooga/text-generation-webui
- Palmik 2y agoThe difference between (A) software engineers reacting to AI models and systems for programming and (B) artists (whether it's painters, musicians or otherwise) reacting to AI models for generating images, music, etc. is very interesting. I wonder what's the reason.
- suprjami 2y agoBecause code either works or it doesn't. Nobody is replacing our entire income stream with an LLM. You also need a knowledge of code to instruct an LLM to generate decent code, and even then it's not always perfect. Meanwhile plenty of people are using free/cheap image generation and going "good enough". Now they don't need to pay a graphic artist or a stock photo licence Any layperson can describe what they want a picture to look like so the barrier to entry and successful exit is a lot lower for LLM image generation than for LLM code generation.
- jodrellblank 2y ago> "Meanwhile plenty of people are using free/cheap image generation and going "good enough". Now they don't need to pay a graphic artist or a stock photo licence" and getting sandwich photos of ham blending into human fingers: https://www.reddit.com/r/Wellthatsucks/comments/1f8bvb8/my_local_sandwich_shop_using_ai_images_for_ads/ https://www.reddit.com/r/Wellthatsucks/comments/1f8bvb8/my_l...
- saurik 2y agoAnd yet, even knowing what I was looking for, I didn't see it long enough that I guessed I misunderstood and swiped to the second image, where it was pointed out specifically. Even if I had noticed myself--presumably because I was staring at it for way too long in the restaurant--I can't imagine I would have guessed what was going on, BUT EVEN THEN it just wouldn't have mattered... clearly, this is more than merely a "good enough" image.
- datavirtue 2y agoAt best it's a prototype and concept generator. It would have to yield assets with layers that can be exported by an illustration or bitmap tool of choice. AI generated images are almost completely useless as-is.
- kleiba 2y agoWhat is the recommended hardware to run a model like that locally on a desktop PC?
- tadasv 2y agoyou can easily run 8b yi coder on 4090 rtx. Probably could do on a smaller gpu (16GB). I have 24gb, and run it through ollama.
- JediPig 2y agoI tested this out on my workload ( SRE/Devops/C#/Golang/C++ ). it started responding about non-sense on a simple write me boto python script that changes x ,y,z value. Then I tried other questions in my past to compare... However, I believe the engineer who did the LLM, just used the questions in benchmarks. One instance after a hour of use ( I stopped then ) it answered one question with 4 different programming languages, and answers that was no way related to the question.
- tmikaeld 2y agoI have the same experience, hallucinates and rambles on and on about "solutions" that are not related. Unfortunately, this has always been my experience with all open source code models that can be self-hosted.
- Gracana 2y agoIt sounds like you are trying to chat with the base model when you should be using a chat model.
- tmikaeld 2y agoNo, I’m using 9b-chat-q8_0 on a 4090
- tmikaeld 2y agoTurns out that Ollama on windows will run multiple models in parallell consuming all available VRAM and RAM. Changing it to 1 fixed the issue, now it's working great! However, the context length for the output is very small - only 1024 tokens.
- Gracana 2y agoThat's some really strange behavior, I don't know why that would cause poor results rather than just poor performance. Can you configure the context size with `/set parameter num_ctx N`? On my laptop with an RTX A3000 12GB I can run `yi-coder:9b-chat` (Q4_0) with 32768 context and it produces good results quickly. That uses 11GB of VRAM so it's maxed out for this setup.
- Tepix 2y agoSounds very promising! I hope that Yi-Coder 9B FP16 and Q8 will be available soon for Ollama, right now i only see the 4bit quantized 9B model. I'm assuming that these models will be quite a bit better than the 4bit model.
- anotherpaulg 2y agoYi-Coder scored below GPT-3.5 on aider's code editing benchmark. GitHub user cheahjs recently submitted the results for the 9b model and a q4_0 version. Yi-Coder results, with Sonnet and GPT-3.5 for scale: 77% Sonnet 58% GPT-3.5 54% Yi-Coder-9b-Chat 45% Yi-Coder-9b-Chat-q4_0 Full leaderboard: https://aider.chat/docs/leaderboards/ https://aider.chat/docs/leaderboards/
- gloosx 2y agoCan someone explain these Aider benchmarks to me? They pass same 113 tests through llm every time. Why they then extrapolate ability of llm to pass these 113 basic python challenges to the general ability to produce/edit code? For me it sounds like this or that model is 70% accurate in solving same hundred python training tasks, but why does it mean that it's good at other languages and arbitrary, private tasks as well? Does anyone ever tried to change them test cases or wiggle conditions a bit to see if it will still hit 70%?
- tarruda 2y agoIt seems this is the problem with most benchmarks, which is why benchmark performance doesn't mean much these days.
- lasermike026 2y agoFirst look seem good. I'll keep hacking with it.
- smokel 2y agoDoes anyone know why the sizes of these models are typically expressed in number of weights (i.e 1.5B and 9B in this case), without mentioning the weight size in bytes? For practical reasons, I often like to know how much GPU RAM is required to run these models locally. The actual number of weights seems to only express some kind of relative power, which I doubt is relevant to most users. Edit: reformulated to sound like a genuine question instead of a complaint.
- tarruda 2y agoSince most LLMs are released as FP16, just the number of parameters is enough to know the total required GPU RAM.
- magnat 2y agoBecause you can quantize a model e.g. from original 16 bits down to 5 bits per weight to fit your available memory constraints.
- GaggiX 2y agoThe weight size depends on the accuracy you are running the model at, you usually do not run a model at fp16 as it would be wasteful.
- zeroq 2y agoEverytime someone tells how AI 10x his programming capabilities I'm like "tell me you're bad at coding without telling me".
- coolspot 2y agoIt allows me to move much faster, because I can write a comment describing something more high-level and get plausible code from it to review & correct.
- jodrellblank 2y agoEverytime someone posts a comment that is just "I'm better than other people", I'm like "what a waste of time reading that was".
- patrick-fitz 2y agoI'd be interested to see how it performs on https://www.swebench.com/ https://www.swebench.com/ Using SWE-agent + Yi-Coder-9B-Chat.
- nathan_tarbert 2y agoThis sounds really cool! I found this Reddit discussion... https://www.reddit.com/r/ArtificialInteligence/comments/1f9mvyu/meet_yicoder_a_small_but_mighty_llm_for_code/ https://www.reddit.com/r/ArtificialInteligence/comments/1f9m...