3 ms·Only in the size of model it can run, not speed of token generation.by gehsty 11mo agoOnly in the size of model it can run, not speed of token generation.