2 ms·
Nice job implementing expert caching!
by WithinReason 2mo ago
Nice job implementing expert caching!
- gitpusher42 2mo agoThank you! Under good conditions it achieves approx a 67% cache hit rate with 16 expert slots
- WithinReason 2mo agoThat's great, now I wonder how cache hit rate scales for larger models. Do you have any plans trying Qwen 3.6 or larger?
- gitpusher42 2mo agoCheck for colibri, dwarf star and flash-moe. they do similar things with bigger models https://github.com/JustVugg/colibri https://github.com/JustVugg/colibri https://github.com/antirez/ds4 https://github.com/antirez/ds4 https://github.com/danveloper/flash-moe https://github.com/danveloper/flash-moe