6 ms·
Even if you choose a more sophisticated similarity measure as suggested by another commenter, you’ll still need to set a threshold on that metric to perform bin
by refibrillator 3y ago
Even if you choose a more sophisticated similarity measure as suggested by another commenter, you’ll still need to set a threshold on that metric to perform binary classification.
In my experience there are two paths forward, the one I recommend is to train an MLP classifier on the embeddings to produce a binary classification (ie similar or not). The advantage is that you no longer need to set a numeric threshold on a distance metric, however you will need labeled training data to define what is “similar” in the context of your use case.
The other path is to calculate the statistics of pair-wise distances for every record in some unlabeled dataset you have, and use the resulting distribution to inform choice of threshold. This will at least give you an estimate of what % of records will be classified as similar for a given threshold value.
- whakim 3y agoI’ve had a lot of success training a classifier on top of these embeddings. An MLP works well, as does an SVM. If you don’t have labeled data, I’ve also found that various active learning techniques can produce extremely strong results while only requiring a small amount of human labeling.
- abhgh 3y agoI second the point about learning a classifier over them - it is practically quite useful. But I'd caution against believing Active Learning (AL) would be unequivocally be useful; most positive performances seem to arise in very specific circumstances [1]. Esp. see Table 1: a Linear SVM with MPNet (one of the embeddings sbert supports) has indeed strong performance, but, in general, random sampling (over an AL strategy) performs quite well! [1] https://arxiv.org/pdf/2403.15744v1.pdf https://arxiv.org/pdf/2403.15744v1.pdf
- whakim 3y agoExtremely interesting paper that runs counter to my experience. In my previous role I ran tens of thousands of classification pipelines similar to those described in the paper, using bert-based models and then running SVMs on top as classifiers. In almost all cases AL techniques such as Least Certainty and Query by Committee dramatically outperformed random sampling. Reading the paper a bit more closely, the two differences that stand out are the use of multiclass classifiers (we transformed multiclass classification problems into a number of binary classification problems in order to improve performance), and the relative simplicity of the datasets used in the paper. That being said, anyone who is doing AL should of course first consider model performance with a randomly-selected subsample of the data!
- abhgh 3y agoAppreciate the reply! We have 2 binary datasets in there [1], but of course, those results are not separately presented - you see the averages. The choice of datasets here is based on what other papers report, to have comparable results. In our experience with real data, AL doesn't fare any better on average - our pipelines use different representations (not just BERT based, includes some legacy ones like n-grams+tf-idf) and classifiers. We do see some strong peformance with RoBERTa (Fig 6) and the margin strategy, which is an uncertainty based strategy like least confidence (Fig 7) - these probably correspond to your observations (to an extent). It is a little hard to comment on differences without looking closely (I'd be happy to go through any document/tech. report you might be able to share) but in our experience, we have seen inconsistencies arise from: * Not performing model selection at each iteration. * Not calibrating the selected model at each iteration. In these cases, a non-random query strategy sometimes just makes up for the lack of proper model selection, making it seem like the query strategy is stronger than it really is. Another point of difference is the seed data size - that can have a significant effect on "priming" the query strategy. On a related note, specifically wrt uncertainty-based strategies, sampling bias is a known problem [2]. [1] On the Fragility of Active Learning, https://arxiv.org/abs/2403.15744 https://arxiv.org/abs/2403.15744 [2] Section 1.2 in Two faces of active learning, https://cseweb.ucsd.edu/~dasgupta/papers/twoface.pdf https://cseweb.ucsd.edu/~dasgupta/papers/twoface.pdf
- whakim 3y agoUnfortunately I don't have anything to share, as much of the data now rests with my former employer. That being said, a few comments: * Understood that dataset selection exists within the context of what others are doing and trying to provide comparable results. That being said, binary sentiment classifiers are (and should be) a pretty low bar for any reasonable technique in 2024, so it necessarily makes the reader wonder whether the conclusions are actually applicable to problems they might be facing. The point is taken that you're also looking at some older methods like TF-IDF which are prevalent in the (historical) AL literature. * Performing model selection and calibration (cross-validation) at each iteration makes intuitive sense, but doesn't really make practical sense. If the goal of AL in the real world is getting frequent updates from a human oracle, new models need to be trained in seconds (at absolute maximum). Even if performance improves by spending time doing model selection/CV, humans simply won't tolerate waiting around for new responses to label. * In your first link, the sample sizes seem extremely...small? If I was consistently faced with datasets only containing 1000 or 2000 data points, I'd probably just do supervised learning because the effort to set up a good AL pipeline (not to mention dealing with stopping criteria etc.) probably wouldn't be worth it. My experience has mostly been with datasets containing tens of thousands of data points (if not more). * 100% agree that sampling bias is a huge issue with a number of AL strategies and one we worked on extensively. A component of that is trying to design the interaction such that the human labeler has some insight into parts of the data distribution that the model is making assumptions about; your second link is very nice and we tried a number of clustering-type techniques alongside our active learner to try and offset sample bias. More generally, I think this points to the importance of considering any human-in-the-loop ML problem in a way that takes into account more than just algorithmic efficiency.
- lmeyerov 3y agoYep, we do precisely this, st + SVM => pre filtering / routing => RAG pipelines for some of louie.ai In our cases, < 100 labels on modern models (per task) goes far. We can do better, but more useful problems to solve. Impressive how far these have come!
- kurt_goedel 3y agoCould you maybe tell me more about this approach? How do I have to build my training dataset? Any paper or other source about this approach?
- anon373839 3y agoThese are excellent suggestions, and much appreciated. Thanks! If you train a binary classifier on the embeddings, have you found that the resulting probabilities are also good for ranking? Or do you stick with a distance measure for that?
- jmalicki 3y agoThe simplest way of building an MLP on top of the embeddings is simply to concatenate the embeddings and put some dense layers on top. However, if you use a "two towers" approach and have several additional MLP layers on top of each embedding separately, and then a dense MLP on the concatenations of each tower, the individual tower MLP layers are an embedding transformation, and will improve retrieval. Like this: Final MLP / \ QP MLP DP MLP | | QE DE Now, you can apply the document preprocessor ("DP MLP") to your document embeddings ("DE") before storing them in the vector database, and apply the query preprocessor ("QP MLP") to your query embeddings ("QE") before querying your vector database. This should improve the precision and recall of your vector retrieval step beyond using e.g. raw LLM embeddings. Even better is for the final layer to just be cosine similarity, maximum inner product, or L2 distance, rather than having an MLP so you can just use a raw threshold (it's at least worth trying).
- anon373839 2y agoInteresting. This sounds like it would be similar to fine-tuning the embeddings, but with added benefit of learning different representations for the query and document. If you keep a distance/similarity measure as the final layer, then I'm assuming this isn't going to work with binary labels?
- jmalicki 2y agoIf you have e.g. cosine distance as the final layer, if the label is 1 you reward the cosine distance being close to 1, if the label is 0 you reward the cosine distance being close to 0. The finetuning here is specific to optimizing for retrieval, which may be different than just matching documents, which can be an advantage. You may want to force the query and document finetunings to be the same, which makes a lot of sense, but the advantage in them differing can be that query strings are often rather short, and in a different sort of language structure than documents, so the differing query and document tunings can in some sense "normalize" queries and documents to be in the same space, when it works well.
- 2y ago