4 ms·
So i have played around with this architecture specifically for entity linking. Two BERT encoders: one for the query text and another for the candidates. Initia
by rdedev 3y ago
So i have played around with this architecture specifically for entity linking. Two BERT encoders: one for the query text and another for the candidates. Initially the two encoders were separate but trained together. When I tried using the same BERT model for both, the accuracy jumped by 4 percentage points. Was pretty suprised by this and i guess it got me thinking that maybe a simple cosine similarity loss function is not enough information for the model to shared latent space. Maybe we also need some weights to be the same between encoders. Granted in my use case above they are the same modality but if we are building a model with image and text encoders it might be helpful to try and tie the weights in the last layers of those two encoders
- refulgentis 3y agoCheck into SBERT, sounds perfect for what you're trying for: same encoder, asymmetric search