2 ms·
>This preview runs a 2B model I guess with 1B or 500M model inference would be even faster?
by DeathArrow 4mo ago
>This preview runs a 2B model
I guess with 1B or 500M model inference would be even faster?
- gaeld 4mo agoIn theory yes, although not in a linearly proportional way, because in practice our memory streaming is not yet perfect. There are still some fixed costs that we did not fully optimize (for now).