7 ms·
Dummy calculations aren't even needed, if you allow the LLMs to pre-compute on the given context before inference: https://arxiv.org/abs/2504.13171 https://arx
by x-complexity 1y ago
Dummy calculations aren't even needed, if you allow the LLMs to pre-compute on the given context before inference:
https://arxiv.org/abs/2504.13171 https://arxiv.org/abs/2504.13171
It should be noted that this type of inference is less useful on time-sensitive tasks, but most tasks truthfully don't require such time sensitivity (there exists slack time between when the task is given & when questions are asked).
- wongarsu 1y agoThere are already some providers offering cheap LLM services that will give you a response within 24 hours instead of within seconds. That allows them to schedule tasks during low-request hours when they have spare capacity and use better batching. For some automated tasks this is perfectly acceptable. A bit of effort to accommodate, but easy to justify when it halves your inference costs
- sandis 1y agoAny examples of such providers?
- dghlsakjg 1y agoCertain tasks at OpenAI when I checked a few months ago. Embedding for one.
- wongarsu 1y agoOpenAI [1] as well as Azure OpenAI, Anthropic [2], as well as Parasail [3] for all the "open source" models. There are others that I was thinking of, but those are the first I could find without my notes. Typically the batch API is 50% cheaper than live inference 1: https://platform.openai.com/docs/guides/batch https://platform.openai.com/docs/guides/batch 2: https://docs.anthropic.com/en/docs/build-with-claude/batch-processing https://docs.anthropic.com/en/docs/build-with-claude/batch-p... 3: https://docs.parasail.io/parasail-docs/batch/batch-quickstart https://docs.parasail.io/parasail-docs/batch/batch-quickstar...