3 ms·
RAG can cost a lot of money if not done thoughtfully. Most embedding and chat completion model providers charge by the token (think number of words in the reque
by chuckhend 3y ago
RAG can cost a lot of money if not done thoughtfully. Most embedding and chat completion model providers charge by the token (think number of words in the request). You'll pay to have the data in the database transformed into embeddings, that is mostly a one-time fixed cost. Then every time there is a search query in RAG, that question needs to be transformed. The chat completion model (like ChatGPT 4), will charge for the number of tokens in the request + number of tokens in the response.
Self-hosting can be a big advantage for cost control, but it can be complicated too. Tembo.io's managed service provides privately hosted embedding models, but does not have hosted chat completion models yet.
- nostrebored 3y agoDo you work for Tembo? I only ask because making embedding model consumption based charges seem like the norm is off to me — there are a ton of open source, self managed embeddings that you can use. You can spin up a tool like Marqo and get a platform that handles making the embedding calls and chunking strategy as well.