3 ms·
Interesting project. The headline number (29 GB of RAM) is for 4k context. From what I've read elsewhere, Kimi K3 is quite verbose in its thinking. At the quot
by logicallee 2mo ago
Interesting project. The headline number (29 GB of RAM) is for 4k context.
From what I've read elsewhere, Kimi K3 is quite verbose in its thinking. At the quoted rate, it would generate only a total of 1.8k tokens in 1 hour. Is that enough for it to get any thinking done and produce output on more complicated prompts?
- 0cf8612b2e1e 2mo agoI saw someone’s excellent idea that if you have a slow system like this, you should communicate by email. It is no longer meant for realtime iteration, but more pointed questions for which there is more effort and time expected on both parties.
- rwz 2mo ago0.5t/s is still too slow even for email. For a moderately large inquiry (1MTok output, let's ignore the 4k context window limitation for now) it'll take the model around 23 days or uninterrupted execution to answer a single email. Real world inquiries are gonna be much slower of course, but this setup is still too slow to do anything meaningfully useful I think.
- springtimesun 2mo agoI’m sure it’s possible, but I really struggle to think of an example that would result in a 1m token output.
- 8note 2mo agoit would be 1M tokens worth of work, with some small report for the end
- 0cf8612b2e1e 2mo agoSure, you cannot do 1M, but there are plenty of useful questions you could ask that are far more modest. Simple Q+A, look at this function, how would you design X? All of those could have few paragraphs of outputs that would finish within a day.