3 ms·
Most AI models blend meaning — but this one selects. GD-Attention is a provably nonlinear attention mechanism with: ・ No softmax ・ No averaging ・ A unique s
by GhostDrift 1y ago
Most AI models blend meaning — but this one selects.
GD-Attention is a provably nonlinear attention mechanism with:
・ No softmax
・ No averaging
・ A unique semantic jump point $s^*$
Verified independently by Gemini, GitHub Copilot, and GPT-4.
→ Softmax isn't just suboptimal — it's structurally incapable of what this model does.