3 ms·
This reminds me of an idea I had for an OpenAI proxy that transparently handles batching of requests. The use case is that OpenAI has rate limits not only on to
by HellsMaddy 3y ago
This reminds me of an idea I had for an OpenAI proxy that transparently handles batching of requests. The use case is that OpenAI has rate limits not only on tokens but also requests per minute. By batching multiple requests together you can avoid hitting the requests limit.
This isn’t really feasible to implement if your app runs on lambda or edge functions, you’d need a persistent server.
Here’s a diagram I drew of a simple approach that came to mind: https://gist.github.com/b0o/a73af0c1b63fccf3669fa4b00ac4be52 https://gist.github.com/b0o/a73af0c1b63fccf3669fa4b00ac4be52
It would be awesome to see this functionality built into BricksLLM.
- computerex 3y agoOpenAI API doesn't support batching afaik.
- HellsMaddy 3y agoThey do: https://platform.openai.com/docs/guides/rate-limits/batching-requests https://platform.openai.com/docs/guides/rate-limits/batching...
- treprinum 3y agoEmbeddings can be batched.
- swyx 3y agohow exactly are you intending to batch different prompts together in the openai api? its not like they accept an array of parallel inputs
- heyn05tradamu5 3y agoThey’ve recently added this functionality to AWS Bedrock thankfully. Doesn’t support OpenAI models, but does support Anthropic. https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-bedrock-batch-inference/ https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-be...
- te_chris 3y agoIf you can get Claude approved.