3 ms·
This approach only works for small context requests. For large context and relatively smaller output (say understanding a huge code base), the cost will mainly
by yiyingzhang 2mo ago
This approach only works for small context requests. For large context and relatively smaller output (say understanding a huge code base), the cost will mainly be on prefill, and sending the large context to multiple models will only increase the cost, possibly by some factor.