3 ms·
Let’s not forget that every round trip with the LLM costs latency (and extra input tokens). We now have parallel tool calls which sometimes works in some models
by blixt 1y ago
Let’s not forget that every round trip with the LLM costs latency (and extra input tokens). We now have parallel tool calls which sometimes works in some models[1]. But it’s great because now a model can say “write these 3 files then read these 2 files” before the time-to-first token latency is incurred once more (not to mention input token cost).
I think LLMs will indirectly move towards being fuzzy VMs that output tokens much like VM instructions so they can prepare multiple conditional branches of tool calling, load/unload useful subprograms, etc. It might not be expressed exactly like that, but I think given how LLMs today are very poor at reusing things in their context window, we will naturally add features that take us in this direction. Also see frameworks like CodeAct[2] etc.
[1] This can be converted to a single tool call with many arguments instead, which you’ll see providers do in their internal tools, but it’s just messier.
[2] https://machinelearning.apple.com/research/codeact https://machinelearning.apple.com/research/codeact