3 ms·
Thanks for pointing out those models. I see from a quick Huggingface search that the bge model is available in GGML format. You can trivially add new GGML forma
by eigenvalue 3y ago
Thanks for pointing out those models. I see from a quick Huggingface search that the bge model is available in GGML format. You can trivially add new GGML format models to the code by simply adding the direct download link to this line:
https://github.com/Dicklesworthstone/llama_embeddings_fastapi_service/blob/c620f0b0ad347c57c6bf23060b3212722d73ed60/llama_2_embeddings_fastapi_server.py#L405 https://github.com/Dicklesworthstone/llama_embeddings_fastap...
So to add the base bge model, you could just add this URL to the list:
https://huggingface.co/maikaarda/bge-base-en-ggml/resolve/main/ggml-model-f32.bin https://huggingface.co/maikaarda/bge-base-en-ggml/resolve/ma...
I will add that as an additional default.
- jerrygenser 3y agoI think it's still overkill though for semantic embedding, SBERT is on order of ~250mm parameters while smallest llama at 7b parameters.
- eigenvalue 3y agoIf all you want to do is make some basic semantic search, that’s probably true. But I strongly suspect we are only just now starting to scratch the surface of what’s possible with embeddings that come from much more powerful LLMs like Llama2 that can clearly manifest much greater demonstrated “understanding” of sentences they are shown (whatever that means, but intuitively, it seems obvious to me). That’s partly why I made this tool—- to aid in my investigations of LLM embeddings in a convenient and performant way.
- gsuuon 3y agoI'm really curious to see where that investigation leads - have you done any comparisons between Llama 2 and the embedding focused models? I wonder if it'll be better able to provide more 'intuitively correct' similarities?