3 ms·
A design that bakes the architecture into silicon would be 10x faster, and imagine a version that does all the multiplication ops using single log-amp addition
by LarsDu88 1mo ago
A design that bakes the architecture into silicon would be 10x faster, and imagine a version that does all the multiplication ops using single log-amp addition versus dozens of transistors to cut down the amount of silicon used by 50x. The ceiling for AI optimized hardware is extremely high.
Stack on top of that the fact that diffusion based models like the ones made by Inception Labs are far faster and more efficient than autoregressive LLMs and have an even higher ceiling of optimization (single step path prediction via model distillation versus 50 step denoise is currently an active area for image diffusion)
The human brain is soon neither going to be more powerful nor energy efficient than the stuff we use to run AI.