3 ms·
Sending two identical parallel requests is the classic approach. But, logically speaking, it should also double the cost. I would send a second request if the
by nine_k 2mo ago
Sending two identical parallel requests is the classic approach. But, logically speaking, it should also double the cost.
I would send a second request if the first request fails to return the first token within, say, 1 second. Then there's a chance the first request is stalling, which is an infrequent event.
I wonder if higher-availability tiers of LLM providers do a similar thing internally.
- ImPostingOnHN 2mo agoToken caching might help here, but if it returns the same result, faster, for the same price as priority, seems good
- awwaiid 2mo agoI wonder how parallel token caches are, like when exploring a tree of sample continuations.