3 ms·
Vector embeddings are key in tech and science. Our research shows that Mixture of Experts (MoE) transformers beat traditional sentence similarity methods. A hug
by warship 3y ago
Vector embeddings are key in tech and science. Our research shows that Mixture of Experts (MoE) transformers beat traditional sentence similarity methods. A huge result. Why?
LLMs often misread nuances in scientific texts. Our scalable solution upgrades pre-trained language models to MoE versions, with expert groups matching multiple model performances across various benchmarks.
Focusing on the cardiovascular disease and chronic obstructive pulmonary disease subfields, we generated training datasets using co-citations (publications citing each other) as a similarity metric.
These scalable datasets required no manual labeling, yet took advantage of the expert collective intelligence of the scientific community (in terms of citation networks), allowing our models to outperform all other tested models.
We believe this new approach marks significant and timely advancements in scientific text classification approaches and holds promise for enhancing vector database tasks, which is an active area of research (and teaching) in our lab.
Congratulations to co-authors Rohan Kapur, Logan Hallee, and Arjun Patel for innovating with this tour-de-force framework in the post-GPT era we all live in today. We hope you may find our work useful in the context of your own AI/machine learning research.