4 ms·
How are these performing so well compared to Llama 2, are there any documents on the architecture and differences, is it MoE? Also note some of the links on th
by neximo64 3y ago
How are these performing so well compared to Llama 2, are there any documents on the architecture and differences, is it MoE?
Also note some of the links on the blog post don't work, e.g debugging tool.
- kathleenfromgdm 3y agoWe've documented the architecture (including key differences) in our technical report here (https://goo.gle/GemmaReport https://goo.gle/GemmaReport), and you can see the architecture implementation in our Git Repo (https://github.com/google-deepmind/gemma https://github.com/google-deepmind/gemma).