4 ms·
The size of the database, the training set, the details of the architecture, as well as results on benchmark tasks should all be considered in the comparison. I
by jayalammar 5y ago
The size of the database, the training set, the details of the architecture, as well as results on benchmark tasks should all be considered in the comparison. I'm also a fan of Behavioral Testing [1].
Parameter count is not very accurate measure of model performance. Mixture of Expert models like the Switch Transformer [2] can be 1 trillion parameters in size, but are not 5X the performance, for example.
They clock the retrieval at 10 ms, unclear if that includes the BERT inference, however. My assumption is that it does not.
[1] https://arxiv.org/abs/2005.04118 https://arxiv.org/abs/2005.04118
[2] https://arxiv.org/abs/2101.03961 https://arxiv.org/abs/2101.03961