3 ms·
I think the point is that it’s running at all…
by ganelonhb 26d ago
I think the point is that it’s running at all…
- cyanydeez 26d agoQwen3.8-Flash-Next ships with a 51B lookup table that can be read directly from ssd or memory, which greatly improves it's speed and intelligence. It can load at 4bit quant in ~60GB. These demos are maybe useless, but if open models keep progressing, there's going to be some break through that continues whittling down just how much needs to be kept in VRAM, and progressive degredation to regular system ram and to ssds. Afterall, they're not writing anything to these, so saturing all bandwidth could bring models to the masses. all without any help from Zark Muckerberg.
- Argonautlabs 26d agoIt actully does the job. Example: every morning it takes 30-40 minutes to generate reports automatically and these reports are being sent as a pdf to read to Telegram.
- kgeist 26d agoDo those reports require Kimi K3 though? Qwen3.6+ could probably do the same in a few seconds with similar quality.
- Argonautlabs 25d agoOften Deep Seek V4 flash or Qwen should be enough. I wanted to see whether Kimi runs at all on one machine with the full record published, and for long multi-table finance reasoning I wanted the strongest model I could keep on the machine. I did some tests against Deep Seek v4 flash results on my reports and Kimi definitely has some advantages.