4 ms·
Surely there can't be a microsecond to second ratio between the two steps, because then they would never be able to crawl the hundreds of billions of documents
by manholio 4y ago
Surely there can't be a microsecond to second ratio between the two steps, because then they would never be able to crawl the hundreds of billions of documents the model is trained on. Once the model is built, sure, the front end is irrelevant.
- mirker 4y agoInference is latency sensitive so the front end is still relevant.
- pifm_guy 4y agoThe ratio between the tokenizer and the model is constant for both training and inference. And I believe the ratio is that big - you just use a lot of compute for training, and that's why it's so expensive.