3 ms·
They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into acc
by Phemist 1mo ago
They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatever is actually required to produce coherent speech, sadly it is all rather entangled so this is not so easy).
- perching_aix 1mo agoBut there are already voice models that do a reasonable job at a fraction of the throughput available? The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much. I half expect Boston Dynamics to show something like this off in Q4 or whatever.
- Phemist 1mo agoI am not arguing that there are perhaps other models that can run at the same quality, can coordinate between the different modalities, but are way less power hungry. My point is exactly about the comparison between the token output of the LLM running on the jalapeno chip, and sneaking in the power "usage" of the brain in the "token output" of human speech.