6 ms·
Hi HN, happy to see this here! I highly recommend to take a look at the technical details of the server implementation that enables large context usage with th
by ggerganov 2y ago
Hi HN, happy to see this here!
I highly recommend to take a look at the technical details of the server implementation that enables large context usage with this plugin - I think it is interesting and has some cool ideas [0].
Also, the same plugin is available for VS Code [1].
Let me know if you have any questions about the plugin - happy to explain. Btw, the performance has improved compared to what is seen in the README videos thanks to client-side caching.
[0] - https://github.com/ggerganov/llama.cpp/pull/9787 https://github.com/ggerganov/llama.cpp/pull/9787
[1] - https://github.com/ggml-org/llama.vscode https://github.com/ggml-org/llama.vscode
- jerpint 2y agoThank you for all of your incredible contributions!
- amrrs 2y agoFor those who don't know, He is the gg of `gguf`. Thank you for all your contributions! Literally the core of Ollama, LMStudio, Jan and multiple other apps!
- sergiotapia 2y agowell hot damn! killing it!
- halyconWays 2y ago[flagged]
- madeforhnyo 2y agoSomeone did? Could you pls share a link?
- kamranjon 2y agoThey collaborate together! Her name is Justine Tunney - she took her “execute everywhere” work with Cosmopolitan to make Llamafile using the llama.cpp work that Giorgi has done.
- halyconWays 2y agoShe actually stole that code from a user named slaren and was personally banned by Gerg from the llama.cpp repo for about a year because of it. Also it was just lazy loading the weights, it wasn't actually a 50% reduction. https://news.ycombinator.com/item?id=35411909 https://news.ycombinator.com/item?id=35411909
- kamranjon 2y agoThat seems like a false narrative, which is strange because you could have just read the explanation from Jart a little further down in the thread: https://news.ycombinator.com/item?id=35413289 https://news.ycombinator.com/item?id=35413289
- kennethologist 2y agoA. Legend. Thanks for having DeepSeek available so quickly in LM Studio.
- nancyp 2y agoTIL: VIM has it's own language. Thanks Georgi for LLAMA.cpp!
- nacs 2y agoVim is incredibly extensible. You can use C or VIMscript but programs like Neovim support Lua as well which makes it really easy to make plugins.
- liuliu 2y agoKV cache shifting is interesting! Just curious: how much of your code nowadays completed by LLM?
- ggerganov 2y agoYes, I think it is surprising that it works. I think a fairly large amount, though can't give a good number. I have been using Github Copilot from the very early days and with the release of Qwen Coder last year have fully switched to using local completions. I don't use the chat workflow to code though, only FIM.
- gloflo 2y agoWhat is FIM?
- jjnoakes 2y agoFill-in-the-middle. If your cursor is in the middle of a file instead of at the end, then the LLM will consider text after the cursor in addition to the text before the cursor. Some LLMs can only look before the cursor; for coding,.ones that can FIM work better (for me at least).
- rav 2y agoFIM is "fill in middle", i.e. completion in a text editor using context on both sides of the cursor.
- menaerus 2y agoInteresting approach. Am I correct to understand that you're basically minimizing the latencies and required compute/mem-bw by avoiding the KV cache? And encoding the (local) context in the input tokens instead? I ask this because you set the prompt/context size to 0 (--ctx-size 0) and the batch size to 1024 (-b 1024). Former would mean that llama.cpp will only be using the context that is already encoded in the model itself but no local (code) context besides the one provided in the input tokens but perhaps I misunderstood something. Thanks for your contributions and obviously the large amount of time you take to document your work!
- deleted 2y ago[deleted]
- halyconWays 2y agoPlease make one for Jetbrains' IDEs!
- bangaladore 2y agoQuick testing on vscode to see if I'd consider replacing Copilot with this. Biggest showstopper right now for me is the output length is substantially small. The default length is set to 256, but even if I up it to 4096, I'm not getting any larger chunks of code. Is this because of a max latency setting, or the internal prompt, or am I doing something wrong? Or is it only really make to try to autocomplete lines and not blocks like Copilot will. Thanks :)
- ggerganov 2y agoThere are 4 stopping criteria atm: - Generation time exceeded (configurable in the plugin config) - Number of tokens exceeded (not the case since you increased it) - Indentation - stops generating if the next line has shorter indent than the first line - Small probability of the sampled token Most likely you are hitting the last criteria. It's something that should be improved in some way, but I am not very sure how. Currently, it is using a very basic token sampling strategy with a custom threshold logic to stop generating when the token probability is too low. Likely this logic is too conservative.
- bangaladore 2y agoHmm, interesting. I didn't catch T_max_predict_ms and upped that to 5000ms for fun. Doesn't seem to make a difference, so I'm guessing you are right.
- attentive 2y agoIs it correct to assume this plugin won't work with ollama? If so, what's ollama missing?
- mistercheph 2y agothis plugin is designed specifically for the llama.cpp server api, if you want copilot like features with ollama, you can use an ollama instance as a drop-in replacement for github copilot with this plugin: https://github.com/bernardo-bruning/ollama-copilot https://github.com/bernardo-bruning/ollama-copilot There is also https://github.com/olimorris/codecompanion.nvim https://github.com/olimorris/codecompanion.nvim which doesn't have text completion, but supports a lot of other AI editor workflows that I believe are inspired by Zed and supports ollama out of the box
- eklavya 2y agoThanks for sharing the vscode link. After trying I have disabled the continue.dev extension and ollama. For me this is wayyyyy faster.