3 ms·
We did two separate studies, one of accuracy and cost the other of accuracy and latency. Most of our studies used cost as the definition is pretty clear and les
by physicalrobot 1y ago
We did two separate studies, one of accuracy and cost the other of accuracy and latency. Most of our studies used cost as the definition is pretty clear and less sensitive to environmental conditions (e.g. LLM provider changes) or definition (e.g. typical vs. worst case). But latency is a concern for many usecases, so we did want to investigate that problem as well.
Note we used Random LLM which had a 0.84 correlation with human labeling rather than 0.90 from only using gpt-4o-mini.