3 ms·
Thanks for the comment. Parity with llama.cpp isn't my goal. What I actually want is MoE on machines that can't fit the model in VRAM, and specifically expert-
by antonellof 2mo ago
Thanks for the comment. Parity with llama.cpp isn't my goal.
What I actually want is MoE on machines that can't fit the model in VRAM, and specifically expert-level residency instead of layer offload: track which experts get hit during decode, keep those resident, evict the rest. Doing that well needs the router, the KV cache and the memory manager to be designed together, which is about the only good reason to write a runtime from scratch.
Yes, let’s see! You are welcome to contribute if you like!
- Alien1Being 2mo agoSince you didn't contribute yourself, that is quite funny. Posting a link to AI slop about a mess of AI vibecode is not contributing.
- antonellof 2mo agoContribute = Human ideas, testing, review etc. The monkey part of writing code by hand = obsolete. Do you still write code without and IDE? you remember all the programming language words? everything? you use stackoverflow? it's called evolution, btw: AI full disclosure This software is developed with strong assistance from Cursor, Grok 4.5, GPT 5.6, and Claude Fable 5, with humans leading the ideas, testing, and debugging. We say this openly because it shaped how the project was built. If you are not happy with AI-developed code, this software is not for you. The acknowledgement below is equally important: this would not exist without llama.cpp and GGML, largely written by hand.
- Alien1Being 2mo agoIf your code is as low quality as your reasoning, I see why you have to vibecode everything. Incidentally, I do not use an IDE, I use vim and emacs. I do not use Stackoverflow. That is for the third raters.
- antonellof 2mo agoAI code quality is not bad, using vim and emacs do not elevate the code quality. do you prefer a typewriter or a computer? it's called future...