7 ms·
Show HN: Wordllama – Things you can do with the token embeddings of an LLM
After working with LLMs for long enough, I found myself wanting a lightweight utility for doing various small tasks to prepare inputs, locate information and create evaluators. This library is two things: a very simple model and utilities that inference it (eg. fuzzy deduplication). The target platform is CPU, and it’s intended to be light, fast and pip installable — a library that lowers the barrier to working with strings semantically. You don’t need to install pytorch to use it, or any deep learning runtimes.
How can this be accomplished? The model is simply token embeddings that are average pooled. To create this model, I extracted token embedding (nn.Embedding) vectors from LLMs, concatenated them along the embedding dimension, added a learnable weight parameter, and projected them to a smaller dimension. Using the sentence transformers framework and datasets, I trained the pooled embedding with multiple negatives ranking loss and matryoshka representation learning so they can be truncated. After training, the weights and projections are no longer needed, because there is no contextual calculations. I inference the entire token vocabulary and save the new token embeddings to be loaded to numpy.
While the results are not impressive compared to transformer models, they perform well on MTEB benchmarks compared to word embedding models (which they are most similar to), while being much smaller in size (smallest model, 32k vocab, 64-dim is only 4MB).
On the utility side, I’ve been adding some tools that I think it’ll be useful for. In addition to general embedding, there’s algorithms for ranking, filtering, clustering, deduplicating and similarity. Some of them have a cython implementation, and I’m continuing to work on benchmarking them and improving them as I have time. In addition to “standard” models that use cosine similarity for some algorithms, there are binarized models that use hamming distance. This is a slightly faster, similarity algorithm, with significantly less memory per embedding (float32 -> 1 bit).
Hope you enjoy it, and find it useful. PS I haven’t figured out Windows builds yet, but Linux and Mac are supported.
- dspoka 2y agoLooks cool! Any advantages to the mini-lm model - it seems better on most mteb tasks but wondering if maybe inference or something is better.
- lennxa 2y agolooks like it's the size of the model itself, more lightweight and faster. mini-lm is 80mb while the smallest one here is 16mb.
- authorfly 2y agoMini-lm isn't optimized to be as small as possible though, and is kind of dated. It was trained on a tiny amount of similarity pairs compared to what we have available today. As of the last time I did it in 2022, Mini-lm can be distilled to 40mb with only limited loss in accuracy, so can paraphrase-MiniLM-L3-v1 (down to 21mb), by reducing the dimensions by half or more and projecting a custom matrix optimization(optionally, including domain specific or more recent training pairs). I imagine today you could get it down to 32mb (= project to ~156 dim) without accuracy loss.
- byefruit 2y agoWhat are some recent sources for high quality similarity pairs?
- deepsquirrelnet 2y agoMini-lm is a better embedding model. This model does not perform attention calculations, or use a deep learning framework after training. You won’t get the contextual benefits of transformer models in this one. It’s not meant to be a state of the art model though. I’ve put in pretty limiting constraints in order to keep dependencies, size and hardware requirements low, and speed high. Even for a word embedding model it’s quite lightweight, as those have much larger vocabularies are are typically a few gigabytes.
- ryeguy_24 2y agoWhich do use attention? Any recommendations?
- deepsquirrelnet 2y ago
- authorfly 2y agoNice. I like the tiny size a lot, that's already an advantage over SBERTs smallest models. But it seems quite dated technically - which I understand is a tradeoff for performance - but can you provide a way to toggle between different types of similarity (e.g. semantic, NLI, noun-abstract)? E.g. I sometimes want "Freezing" and "Burning" to be very similar (1) as in regards to say grouping/clustering articles in a newspaper into categories like "Extreme environmental events", like on MTEB/Sentence-Similarity, as classic Word2Vec/GloVe would do. But if this was a chemistry article, I want them to be opposite, like ChatGPT embeddings would be. And sometime I want to use NLI embeddings to work our the causal link between two things. Because the latter two embedding types are more recent (2019+), they are where the technical opportunity is, not the older MTEB/semantic similarity ones which have been performant enough for many use cases since 2014 and 2019 received a big boost with mini-lm-v2 etc. For the above 3 embedding types I can use SBERT but the dimensions are large, models quite large, and having to load multiple models for different similarity types is straining on resources, it often takes about 6GB because generative embedding models (or E5 etc) are large, as are NLI models.
- deepsquirrelnet 2y agoGreat ideas - I’ll run some experiments and see how feasible it is. I’d want to see how performance is if I train on a single type of similarity. Without any contextual computation, I am not sure there are other options for doing it. It may require switching between models, but that’s not much of an issue.
- refulgentis 2y agoIts a 17 MB model that benchmarks obviously worse than MiniLM v2 (which is SBERT). I run V3 on ONNX on every platform you can think of with a 23 MB model. I don't intend for that to be read as dismissive, it's just important to understand work like this in context - here, it's that there's a cool trick where if you get to an advanced understanding of LLMs, you notice they have embeddings too, and if that is your lens, it's much more straightforward to take a step forward and mess with those, than take a step back and survey the state of embeddings.
- curl-up 2y ago
- anonymousfilter 2y agoHas anyone thought of using embeddings to solve Little Alchemy? #sample-use
- jcmeyrignac 2y agoAny plan for languages other than english? This would be a perfect tool for french language.
- deepsquirrelnet 2y agoIt’s certainly feasible. I’d need to put together a corpus for training and I’m not terribly familiar with what’s available for French language. I have done some training with the Mistral family of models, and that’s probably what I’d think to try first on a French corpus. Feel free to open an issue and I’ll work on it as I find time.
- Ey7NFZ3P0nzAe 2y agoVery interested in a multilingual version too! FYI huggingface hosts datasets too. And wikipedia has a nice portal for datasets : https://en.m.wikipedia.org/wiki/List_of_datasets_for_machine-learning_research https://en.m.wikipedia.org/wiki/List_of_datasets_for_machine...
- ttpphd 2y agoThis is great for game making! Thank you!
- warangal 2y agoEmbeddings capture a lot of semantic information based on the training data and objective function, and can be used independently for a lot of useful tasks. I used to use embeddings from the text-encoder of CLIP model, to augment the prompt to better match corresponding images. For example given a word "building" in prompt, i would find the nearest neighbor in the embedding matrix like "concrete", "underground" etc. and substitute/append those after the corresponding word. This lead to a higher recall for most of the queries in my limited experiments!
- nostrebored 2y agoYup, and you can train these in-domain contextual relationships into the embedding models. https://www.marqo.ai/blog/generalized-contrastive-learning-for-multi-modal-retrieval-and-ranking https://www.marqo.ai/blog/generalized-contrastive-learning-f...
- deepsquirrelnet 2y agoThat’s a really cool idea. I’ll think about it some more, because it sounds like a feasible implementation for this. I think if you take the magnitude of any token embedding in wordllama, it might also help identify important tokens to augment. But it might work a lot better if trained on data selected for this task.
- visarga 2y agoThis shows just how much semantic content is embedded in the tokens themselves.
- Der_Einzige 2y agoI wrote a set of "language games" which used a similar set of functions years ago: https://github.com/Hellisotherpeople/Language-games https://github.com/Hellisotherpeople/Language-games
- xk3 2y agoInteresting... looks like this uses pymagnitude https://github.com/plasticityai/magnitude https://github.com/plasticityai/magnitude
- johnthescott 2y agohmm ... postgresql extension?
- xk3 2y agoWith a large corpus (10,000+ sentences--each sentence is a "document" in my usecase) I can get similar results by kmeans clustering TF-IDF spmatrix vectors but it looks like this has a lot of utilities for making the kmeans part faster (binarization, etc). Looking forward to doing some benchmarking over the next couple weeks