6 ms·
A Replacement for BERT
- jbellis 2y agoLooks great, thanks for training this! - Can I fine tune it with SentenceTransformers? - I see ColBERT in the benchmarks, is there an answerai-colbert-small-v2 coming soon?
- gunalx 2y agoSeems like it. They even have example training scripts available. https://github.com/AnswerDotAI/ModernBERT/blob/main/examples/train_st.py https://github.com/AnswerDotAI/ModernBERT/blob/main/examples... Check out their documentation page linked on the bottom of the article. https://huggingface.co/docs/transformers/main/en/model_doc/modernbert https://huggingface.co/docs/transformers/main/en/model_doc/m...
- jph00 2y agoThe creator of answerai-colbert-small-v2 (bclavie) is also the person that launched the ModernBERT project, so yes, you can expect to see a lot of activity in this space! :D (Also yes, it works great with ST and we provide a full example script.)
- zelias 2y agomissed opportunity to call it ERNIE
- amrrs 2y agoI remember back in the day there was an Ernie model
- axpy906 2y agoDon’t forget ELMO. The bi-lstm.
- lrog 2y agoyep, too late: https://huggingface.co/docs/transformers/en/model_doc/ernie https://huggingface.co/docs/transformers/en/model_doc/ernie
- behnamoh 2y agoI never liked the names BERT and its derivatives. Of all the names on the world, they chose words that are ugly, specific to one culture, and frankly childish.
- Cthulhu_ 2y agoSesame Street has been broadcast in 140 countries; Bert (and Ernie) have been localized to 18 languages, including Arabic, Hindi, Japanese, Hebrew and Chinese, with China having an AI called ERNIE because of course. Or to make an overly worded / researched reply to a petulant comment short, they are very much not specific to one culture.
- Dalewyn 2y ago>they are very much not specific to one culture. It's American culture. Or as Civilization would put it: We made them buy our blue jeans and listen to our pop music. Note: I disagree with all the other points like "ugly" and "childish".
- chriswarbo 2y agoTangentially: ERNIE is probably the most famous "computer" in the UK, which has been picking winners for the UK's premium bonds scheme since the 1950s. It was heavily marketed, to get the public used to the new-fangled idea of electronics, and is sometimes considered one of the first computers; though (a) it was more of a special-purpose random number generator rather than a computer, and (b) it descended from the earlier Colossus code-breaking machines of World War II (though the latter's existence was kept secret for decades). The latest ERNIE is version 5, which uses quantum effects to generate its random numbers (earlier versions used electrical and thermal noise). https://en.wikipedia.org/wiki/Premium_Bonds#ERNIE https://en.wikipedia.org/wiki/Premium_Bonds#ERNIE
- timClicks 2y agoMore generally, using the prefix "Modern" haunts every product name that uses it. Technologies move fast and modern becomes antiquated very quickly.
- int_19h 2y agoIt'll just get shortened to Mobert in the long run anyway.
- bclavie 2y agoWe had a bit of a discussion around it, but I figured that 6 years warranted the prefix, and it's easier to remember in the sea of new acronyms popping up everyday. Besides, PostModernBERT will be there for us for the next generational jump.
- pantsforbirds 2y agoAwesome news and something I really want to checkout for work. Has anyone seen any RAG evals for ModernBERT yet?
- cubie 2y agoNot yet - these are base models, or "foundational models". They're great for molding into different use cases via finetuning, better than common models like BERT, RoBERTa, etc. in fact, but like those models, these ModernBERT checkpoints can only do one thing: mask filling. For other tasks, such as retrieval, we still need people to finetune them for it. The ModernBERT documentation has some scripts for finetuning with Sentence Transformers and PyLate for retrieval: https://huggingface.co/docs/transformers/main/en/model_doc/modernbert#resources https://huggingface.co/docs/transformers/main/en/model_doc/m... But people still need to make and release these models. I have high hopes for them.
- Arcuru 2y agoI'm not sure I am understanding where exactly this slots in, but isn't this an embedding model? Shouldn't they be comparing it to a service like Voyage AI? - https://docs.voyageai.com/docs/embeddings https://docs.voyageai.com/docs/embeddings
- spott 2y agoEmbedding models are frequently based on Bert style models, but Bert models can be finetuned to do a lot more than just embeddings. So an embedding focused finetune of modern Bert should be compared to something like voyageai, but not modern Bert itself.
- KTibow 2y agoWhat are the people who keep downloading Bert doing then? Are they the minority who directly use it for embeddings?
- spott 2y agoI’m honestly not sure why Bert-based-uncased is so popular… the model isn’t that useful on its own. From their huggingface page: > You can use the raw model for either masked language modeling or next sentence prediction, but it's mostly intended to be fine-tuned on a downstream task. See the model hub to look for fine-tuned versions of a task that interests you. > Note that this model is primarily aimed at being fine-tuned on tasks that use the whole sentence (potentially masked) to make decisions, such as sequence classification, token classification or question answering. For tasks such as text generation you should look at model like GPT2.
- metanonsense 2y agoI am out of the game for a year or so (and was never completely in the game), but back then BERT was the basis for lots of interesting applications. The original Vision Transformer (ViT) was based (or at least inspired by) BERT, it was used for graph transformers, visual language understanding, etc.
- jph00 2y agoHi gang, Jeremy from Answer.AI here. Nice to see this on HN! :) We're very excited about this model release -- it feels like it could be the basis of all kinds of interesting new startups and projects. In fact, the stuff mentioned in the blog post is only the tip of the iceberg. There's a lot of opportunities to fine tune the model in all kinds ways, which I expect will go far beyond what we've managed to achieve in our limited exploration so far. Anyhoo, if anyone has any questions, feel free to ask!
- ZQ-Dev8 2y agoJeremy, this is awesome! Personally excited for a new wave of sentence transformers built off ModernBERT. A poster below provided the link to a sample ST training script in the ModernBERT repo, so that's great. Do you expect the ModernBERT STs to carry the same advantages over ModernBERT that BERT STs had over the original BERT? Or would you expect caveats based on ModernBERT's updated architecture and capabilities?
- jph00 2y agoYes absolutely the same advantages -- in fact the maintainer of ST is on the paper team, and it's been a key goal from day one to make this work well.
- data_ders 2y agowhat’s ST stand for here? I googled and only got results for BERT STS (semantic text similarity)
- bclavie 2y agoSentence Transformers (https://sbert.net/ https://sbert.net/), the most used library for embedding models (similarity, retrieval.)
- derbaum 2y agoHey Jeremy, very exciting release! I'm currently building my first product with RoBERTa as one central component, and I'm very excited to see how ModernBERT compares. Quick question: When do you think the first multilingual versions will show up? Any plans of you training your own?
- carschno 2y agoThe model cars says only English, is that correct? Are there any plans to publish a multilingual model or monolingual ones for other languages?
- amunozo 2y agoYes, the paper says that is only English.
- readthenotes1 2y agoI guess the next release is going to be postmodern bert.
- janalsncm 2y ago> encoder-only models add up to over a billion downloads per month, nearly three times more than decoder-only models This is partially because people using decoders aren’t using huggingface at all (they would use an API call) but also because encoders are the unsung heroes of most serious ML applications. If you want to do any ranking, recommendation, RAG, etc it will probably require an encoder. And typically that meant something in the BERT/RoBERTa/ALBERT family. So this is huge.
- EGreg 2y agoCan you go into detail for those of us who aren't as well versed in the tech? What do the encoders do vs the decoders, in this ecosystem? What are some good links to learn about these concepts on a high level? I find all most of the writing about different layers and architectures a bit arcane and inscrutable, especially when it comes to Attention and Self-Attention with multiple heads.
- cubie 2y agoOn a very high level, for NLP: 1. an encoder takes an input (e.g. text), and turns it into a numerical representation (e.g. an embedding). 2. a decoder takes an input (e.g. text), and then extends the text. (There's also encoder-decoders, but I won't go into those) These two simple definitions immediately give information on how they can be used. Decoders are at the heart of text generation models, whereas encoders return embeddings with which you can do further computations. For example, if your encoder model is finetuned for it, the embeddings can be fed through another linear layer to give you classes (e.g. token classification like NER, or sequence classification for full texts). Or the embeddings can be compared with cosine similarity to determine the similarity of questions and answers. This is at the core of information retrieval/search (see https://sbert.net/ https://sbert.net/). Such similarity between embeddings can also be used for clustering, etc. In my humble opinion (but it's perhaps a dated opinion), (encoder-)decoders are for when your output is text (chatbots, summarization, translation), and encoders are for when your output is literally anything else. Embeddings are your toolbox, you can shape them into anything, and encoders are the wonderful providers of these embeddings.
- dmezzetti 2y agoGreat news here. Will takes some time for it to trickle downstream but expect to see better vector embeddings models, entity extraction and more.
- cubie 2y agoSpot on
- crimsoneer 2y agoAnswer.ai team are DELIVERING today. Well done Jeremy and team!
- wenc 2y agoCan I ask where BERT models are used in production these days? I was given to understand that they are a better alternative to LLM type models for specific tasks like topic classification because they are trained to discriminate rather than to generate (plus they are bidirectional so they can “understand” context better through lookahead). But LLMs are pretty strong so I wonder if the difference is negligible?
- vietvu 2y agoLLMs like GPT are heavy and costly (and BERT are LLMs too, params can up to like 1.5B). For niche problems like classification on a small domain, BERT like models are much better, cheaper. You don't need all knowledge gen AI LLM has. I have seen many companies using DeBERTa or RoBERTa for text classification, not using GPT/LLaMA.
- ganeshkrishnan 2y agoLLMs dont have the same usecase as encoder only models. Lets assume you have around million keywords and you want to find the most similar to a keyword that the user input. In pre-processing you would have calculated the vector encoding of all the million keywords before hand and now with the keyword the user input, you calculate the vector and then find the most similar vectors LLM is used by end user, encoders are used by devs in app to search/retrieve text.
- deepsquirrelnet 2y agoI read your paper this morning, and am just thrilled with the work. Love the added local attention layers. I’ve experimented with them for years (lucidrains repo), and was always surprised they didn’t go further. Inference speeds are awesome on this model. Scrapping NSP, awesome. Increased masking, awesome. RoPE and longer context, again, bravo. There’s so many great incremental improvements learned over the years and you guys made so many good decisions here. I’d love to distill a “ModernTinyBERT”, but it seems a bit more complex with the interleaved layers.
- anon373839 2y ago> I’d love to distill a “ModernTinyBERT That’s a question I’m interested in as well! DistilBERT and friends have been terribly useful at the edge. I wonder if/when we may see something similar for ModernBERT.
- mark_l_watson 2y agoI saw this early this morning. About for or five years ago I used BERT models for summarization, etc. BERT seemed like a miracle to me back then. I am going to wait until Ollama has this in their library, even though consuming HF is straight forward. The speedup is impressive, but then so are the massive speed improvements for LLMs recently. Apple has supported BERT models in their SDKs for Apple developers for years, it will be interesting to see how quickly they update to this newer tech.
- vietvu 2y agoSo that what's Jeremy Howard was teasing about. Nice one.
- GaggiX 2y agoIt would be really cool to have a model like this but multilingual, it would really help with things like moderation.
- neodypsis 2y agoHow does it compare to Jina V3 [0], which also has 8192 context length? 0. https://arxiv.org/abs/2409.10173 https://arxiv.org/abs/2409.10173
- bclavie 2y agoThey perform different roles, so they're not directly comparable. Jina V3 is an embedding model, so it's a base model, further fine-tuned specifically for embedding-ish tasks (retrieval, similarity...). This is what we call "downstream" models/applications. ModernBERT is a base model & architecture. It's not supposed to be out of the box, but fine-tuned for other use-cases, serving as their backbone. In theory (and, given early signal, most likely in practice too), it'll make for really good downstream embeddings once people build on top of it!
- Labo333 2y agoSad that it is English only, not multilingual.
- shahjaidev 2y agoThe community would benefit a lot from a multilingual ModernBERT. Pretraining on a multilingual corpus is crucial for a ranking/retrieval model to be deployed in many industry settings.Simply extending the vocab and fine tuning the en checkpoint won’t quite work. Any plans to release a multilingual checkpoint ?
- 303bookworm 2y agoReally excited to see this! 2 Questions: 1. Did you try using RTD (Electra like pretraining)? Or did you skip that for reasons of compatability? 2. Why not incorporate jamba like Mamba2 alternating layers?