4 ms·
thanks! we explain how it scales to larger models in the last section the OP blog post
by gaeld 4mo ago
thanks! we explain how it scales to larger models in the last section the OP blog post
- bcjdjsndon 4mo agoShame you stopped short of actually benchmarking that scale though, eh?
- gaeld 4mo agowill do - we are a small team and it takes time to implement and optimize a new model, whatever the size.
- lostmsu 4mo agoYou don't even need to train the model just to see if you can infer it at the claimed speed
- gaeld 4mo agoTrue, and for third-party models we'll just re-use their public open weights. There is a time-consuming part, though, that is performed manually by our (human) team: implement the logic of the model in C++ and assembly code in a super-optimized way, co-designed for each specific hardware card. This can take months. We hope to accelerate the process with AI agents, but we're not there yet.
- bcjdjsndon 4mo agoOh