5 ms·
Thanks Alex, in this article we focused more on deployment comparisons, for example the cost and latency of what it would take to deploy a BERT based model vs L
by hellovai 3y ago
Thanks Alex, in this article we focused more on deployment comparisons, for example the cost and latency of what it would take to deploy a BERT based model vs LLMs.
In a future article, we're planning on posting accuracy comparisons as well, but here we want to evaluate a few other architectures for comparison. For example, at 1TPS with 1k tokens, chat-gpt-turbo would cost almost $5k vs a simpler BERT model you could run for under $50.
This is probably very obvious to some people, but a lot of people's first experience with any sort of AI is often an LLM, so this is just the first of many posts we hope to share.