4 ms·
"I've discovered that attention heads in transformer models can be approximated by simple MLPs using only 5% of the original parameters while maintaining nearly
by MikeBee 1y ago
"I've discovered that attention heads in transformer models can be approximated by simple MLPs using only 5% of the original parameters while maintaining nearly identical performance. This could significantly reduce the power consumption of LLMs by 95%. My research includes a working demonstration with full code. I'd love to discuss how this approach could benefit Anthropic's efficiency goals. Read the full paper here: https://medium.com/@mbonsign/attention-heads-can-be-approximated-by-simple-neural-networks-972a37ee2ec0 https://medium.com/@mbonsign/attention-heads-can-be-approxim..."