5 ms·
I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress
by fraboniface 1mo ago
I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.
- danishanish 1mo agoI mean, surely when quality is accounted for the difference is significantly higher
- GaggiX 1mo agoOr maybe significantly lower.
- plasticchris 1mo agoProbably not when you consider the training cost and upkeep expenses, not to mention the depreciation…
- Phemist 1mo agoThe 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.
- phoghed 1mo agokind of a moot point if you can't get your brain to not do everything else. I think it's a fun comparison, even if it's not a 100% equivalence.
- CooCooCaCha 1mo agoAnd the brain is literally only producing electrochemical signals. I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.
- Phemist 1mo ago> I don’t see how tokens can’t produce speech or track metabolic needs. It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.
- cmrdporcupine 1mo agoRight, I can do the talked about ~3 tok/sec output and drive a car, hold my bladder, and eat chips at the same time. Take that, Jalapeno!
- falcor84 1mo agoFor what it's worth, LLMs don't really suffer from incontinence, so at least that part is pretty much a solved problem.
- pantalaimon 1mo agoThey sometimes leak their system prompt
- undersuit 1mo agoSo why are their water cooling systems filled with leak detectors? /s
- desterothx 1mo agoAlso, I can talk at a lot more than 3 tok/sec, its just going to be more gibberish (see thinking traces)
- perching_aix 1mo ago> At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model. This is about the same rate you get out of Sol Ultraspeed. Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not a single user thing anymore, which is exactly what they have in the cooker with Astra.
- Phemist 1mo agoThey are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatever is actually required to produce coherent speech, sadly it is all rather entangled so this is not so easy).
- perching_aix 1mo agoBut there are already voice models that do a reasonable job at a fraction of the throughput available? The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much. I half expect Boston Dynamics to show something like this off in Q4 or whatever.
- Phemist 1mo agoI am not arguing that there are perhaps other models that can run at the same quality, can coordinate between the different modalities, but are way less power hungry. My point is exactly about the comparison between the token output of the LLM running on the jalapeno chip, and sneaking in the power "usage" of the brain in the "token output" of human speech.
- kemiller 1mo agoI wonder how that stacks up if you consider all the time you have to keep the body alive when it’s not actively producing “tokens”.
- jdiff 1mo agoCareful, let's not put the whole matrix into stasis outside of business hours. Productivity is not the only reason to let these meatbags burn oxygen.
- DoctorOetker 1mo agoI couldn't source the parameters from the screenshot or the nearby graphs, but from the nearby graphs you can see that at concurrency C=1, tokens/Joule (vertical axis) has totally plummeted, and obviously concurrent inference is much more efficient by batching. Divide the memory by the bandwidth and thats how long it takes to dump the full RAM contents through the chip. Do you want to do this once per token for a single conversation, or do you want to progress multiple conversations if you're going through all the weights anyway? The peak in the graphs is easily 22x more efficient than the low bottom right part on the graphs. So in batched mode its already more efficient than human speech.
- jstummbillig 1mo agoAt just inference! Which both a human and a model can not do without training, but while training rounds to zero for the model, for humans it scales linearly. I am relatively certain we have already squarely been beaten in net efficiency at scale.
- saagarjha 1mo agoYou’re missing the factor for intelligence/token.
- nojs 1mo ago> Humans are still 22x more efficient, which is not that far considering the rate of progress in this area. Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison
- falcor84 1mo agoWhat exactly are you questioning?
- nojs 1mo agoThe claim that tok/s independent of quality is a useful comparison (I can get thousands of tok/s on a suitable small model), and secondarily that humans can’t output “tokens” faster than than in some sense, which I am less confident about
- xyzsparetimexyz 1mo agoI do believe that this is the trade off. We are more efficient but slower in terms of thinking (at the same level of intelligence). Some animals go much further in terms of that trade off, see https://en.wikipedia.org/wiki/Portia_(spider) https://en.wikipedia.org/wiki/Portia_(spider) for example.
- freakynit 1mo agoJust checked wikipedia page... they have like 100K neurons only.. wtf!!! How can nature cramp all senses, including spatial, motion, life maintenance and general thinking into just 100K neurons?
- levocardia 1mo agoA lot of the low level stuff is outsourced to biochemistry: the physical properties of proteins, and the various self-regulating biochemical systems of an animal, can "encode" a lot of intelligence, easing up on the computational demands of the brain proper.
- walrus01 1mo agoFairly amazing when you think about it, like human intellect can run on a bowl of rice and a chicken yakitori skewer.
- sinuhe69 1mo ago22 times more efficient is not like 22 times more powerful. It’s extremely harder to close the gap in power efficiency than in raw power. Simply because the power efficiency we see is the result of billion years evolution optimization. But the true number is IMO far bigger: orders of magnitude greater if we think in terms of equivalent performance.