3 ms·
Towards understanding multiple attention sinks in LLMs
- deleted 7mo ago[deleted]
- thw20 7mo agoThis project reveals an interesting phenomena, where LLM converts semantic non-informative tokens to attention sinks through middle layer MLP. The converted sinks are termed secondary attention sinks as they are weaker then BOS attention sinks. This might be related to layer specialisation in LLM!
- deleted 7mo ago[deleted]
- thw20 7mo agoThe up to date paper documenting and analysing the observation is now available on arxiv!