3 ms·
Not sure what “official” means but would direct you to the GCP MaxText [0] framework which is not what this GDM paper is referring to but rather this repo conta
by yarri 1y ago
Not sure what “official” means but would direct you to the GCP MaxText [0] framework which is not what this GDM paper is referring to but rather this repo contains various attention implementations in MaxText/layers/attentions.py
[0] https://github.com/AI-Hypercomputer/maxtext https://github.com/AI-Hypercomputer/maxtext