4 ms·
They’ve drowned in the LLM noise, but they’re definitely still relevant. - Generative model outputs are not always desirable, and often even undesirable - BER
by deepsquirrelnet 2y ago
They’ve drowned in the LLM noise, but they’re definitely still relevant.
- Generative model outputs are not always desirable, and often even undesirable
- BERT models are smaller and can run with lower latency and serve larger batches with lower vram requirements
- BERT models have bidirectional attention, which can improve performance in many applications
LLMs are “cheap” in the sense that they work well generically, without requiring fine tuning. Where they overlap with BERT models is mostly that they may work better in low training data environments due to better generalization capabilities.
But mostly companies like them because they don’t “require” ML engineers or data scientists on staff. For the lack of care given to evaluation that I see around LLM apps, I suspect that’s going to prove to be a faulty premise.
- antononcube 2y ago> - BERT models are smaller and can run with lower latency and serve larger batches with lower vram requirements The most recent version of Wolfram Language (aka Mathematica) uses by default BERT models for embedding. (Say, for this function: https://reference.wolfram.com/language/ref/CreateSemanticSearchIndex.html https://reference.wolfram.com/language/ref/CreateSemanticSea... .)