3 ms·
Can we assume that model performance at 90% of the 256k limit != 90% of 1M token limit? Is this the exact same model just with less VRAM allocated for context
by jscott817 2mo ago
Can we assume that model performance at 90% of the 256k limit != 90% of 1M token limit?
Is this the exact same model just with less VRAM allocated for context window?
- 3836293648 2mo agoNo, you can't assume it. You can trust them as they made that claim outright. Or you can choose to not trust them, I guess