6 ms·
Maybe I’m missing something but I’ve created vector embeddings for all of English Wikipedia about a dozen times and it costs maybe $10 of compute on Colab, not
by gfourfour 2y ago
Maybe I’m missing something but I’ve created vector embeddings for all of English Wikipedia about a dozen times and it costs maybe $10 of compute on Colab, not $5000
- emmelaich 2y agoGot any details?
- gfourfour 2y agoNothing too crazy, just downloading a dump, splitting it into manageable batch sizes, and using a lightweight embedding model to vectorize each article. Using the best GPU available on colab it takes maybe 8 hours if I remember correctly? Vectors can be saved as NPY files and loaded into something like FAISS for fast querying.
- j0hnyl 2y agoHow big is the resulting vector data?
- gfourfour 2y agoLike 8 gb roughly
- abetusk 2y agoThis probably deserves its own article and might be of interest to the HN community.
- gfourfour 2y agoI will probably make a post when I launch my app! For now I’m trying to figure out how I can host the whole system for cheap because I don’t anticipate generating much revenue
- solarengineer 2y agoI'm interested in hearing about what you will be hosting. Would Digital Ocean or Hetzner meet your needs?
- gfourfour 2y agoI was going to use ec2 and s3, should I look at digital ocean or hetzner instead?
- riku_iki 2y agoWhat is the end task(e.g. RAG, or just vector search for question answering), are you satisfied with results in terms of quality?
- gfourfour 2y agoThe end result is a recommendation algorithm, so basically just vector similarity search (with a bunch of other logic too ofc). The quality is great, and if anything a little bit of underfitting is desirable to avoid the “we see you bought a toilet seat, here’s 50 other toilet seats you might like” effect.
- thomasfromcdnjs 2y agoDid you chunk the articles? If so, in what way?
- gfourfour 2y agoYes. I split the text into sentence and append sentences to a chunk until the max context window is reached. The context window size is dynamic for each article so that each chunk is roughly the same size. Then I just do a mean pool of the chunks for each article.
- thomasfromcdnjs 2y agoThanks for the answer. Wikipedia has a lot of tables so I was wondering if content-aware sentence chunking would be good enough for Wikipedia. https://www.pinecone.io/learn/chunking-strategies/ https://www.pinecone.io/learn/chunking-strategies/
- gfourfour 2y agomwparserfromhell can parse the text content without including tables
- hivacruz 2y agoDid you do use the same method, i.e. split by chunks each article and vectorize each chunk?
- gfourfour 2y agoYes
- dudus 2y agoThat's the only way to do it. You can't index the whole thing. The challenge is chunking. There are several different algorithms to chunk content for vectorization with different pros and cons.
- minimaxir 2y agoYou can do much bigger chunks with models that support RoPE embeddings, such as nomic-embed-text-1.5 which has a 8192 context length: https://huggingface.co/nomic-ai/nomic-embed-text-v1.5 https://huggingface.co/nomic-ai/nomic-embed-text-v1.5 In theory this would be an efficiency boost but the performance math can be tricky.
- qudat 2y agoAs far as I understand it, context length degrades llm performance, so just because an llm "supports" a large context length it basically just clips a top and bottom chunk and skips over the middle bits.
- rahimnathwani 2y agoWhy would you want chunks that big for vector search? Wouldn't there be too much information in each chunk, making it harder to match a query to a concept within the chunk?
- nostrebored 2y agoThe problem is that often semantic meaning depends on state multiple paragraphs or sections away. This is a coarse way to tackle that
- DataDaemon 2y agoHow?
- bunderbunder 2y agoThis is covering 300+ languages, not just English, and it's specifically using Cohere's Embed v3 embeddings, which are provided as a service and currently priced at US$0.10 per million tokens [1]. I assume if you're running on Colab you're using an open model, and possibly a relatively lighter weight one as well? [1]: https://cohere.com/pricing https://cohere.com/pricing
- traverseda 2y agoThis is pretty early in the game to be relying on proprietary embeddings, don't you think? If if they are 20% better, blink and there will be a new normal. It's insane to me that someone, this early in the gold rush, would be mining in someone else's mine, so to speak
- bunderbunder 2y agoI have no idea. But that wasn't the question I was answering. It was, "how does the article's author estimate that would cost $5000?" And I think that's how. Or at least, that gets to a number that's in the same ballpark as what the author was suggesting. That said, first guess, if you do want to evaluate Cohere embeddings for a commercial application, using this dataset could be a decent basis for a lower-cost spike.
- jbellis 2y agoYes, that is how I came up with that number.
- janalsncm 2y agoIt’s not just that. Embeddings aren’t magic. If you’re going to be creating embeddings for similarity search, the first thing you need to ask yourself is what makes two vectors similar such that two embeddings should even be close together? There are a lot of related sources of similarity, but they’re slightly different. And I have no idea what Cohere is doing. Additionally, it’s not clear to me how queries can and should be embedded. Queries are typically much shorter than their associated documents, so they typically need to be trained jointly. Selling “embeddings as a service” is a bit like selling hashing as a service. There are a lot of different hash functions. Cryptographic hashes, locality sensitive hashes, hashes for checksum, etc.
- janalsncm 2y agoAlso, if you’re spending $5000 to compute embeddings, why are you indexing them on a laptop?
- syllogistic 2y agoHe's not though because cohere stuck the already embedded dataset on huggingface https://huggingface.co/datasets/Cohere/wikipedia-22-12-en-embeddings https://huggingface.co/datasets/Cohere/wikipedia-22-12-en-em...
- bytearray 2y agoDo you have a link to the notebook?
- gfourfour 2y agoNo haha just a rats nest of a bunch of notebooks