3 ms·
(Other author of this blog post here) We actually do CPU inference. The SBERT models have a pretty small memory footprint -- you can fit a couple models on a t
by aaronvg 3y ago
(Other author of this blog post here)
We actually do CPU inference. The SBERT models have a pretty small memory footprint -- you can fit a couple models on a t2.medium instance.
On a C6.Large you can get 75ms inference. T2.medium is more around 100-200ms