3 ms·
Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s. Al
by npodbielski 16d ago
Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s.
Also model with their draft answered incorrectly.
With MTP it answered correctly.
Question was: "Does MikroTik CRS312-4C+8XG-RM have combo ports?". The answer is Yes.
- nvme0n1p1 16d agoThat's not how you're supposed to use LLMs. You shouldn't expect a tiny little local model to know random facts about every obscure consumer product on earth. That's the job of tool calling. At best a model of this size is just giving you a random guess. You're basically saying "I tried rolling these dice one time, the green dice rolled a 6 and the blue dice rolled a 1, so green dice are better"
- serf 16d agoagreed. a niche knowledge callout is about the worst benchmark one can give a smaller model. smaller models are attempting to distill the useful methodologies, not the license plate number of an obscure extras car on Magnum PI. that said I wonder if there is a small 'trivia' model out there. Seems like the kinda thing Google would tackle.
- npodbielski 16d agoWhich was not he point because I was testing their solution for MPT and it was just funny addition. But of course in internet you always will find some 'well akchually' person straight from the meme.
- npodbielski 16d agoWhen I changed the number of draft tokens to 3 in both, it helped and they Draft is actually performing a bit better: - draft: 67.17 - MTP: 64.18 Why they used those examples? Seems strange.