2 ms·
>I couldn't find a lot of information as to size of the network used In the original paper,[0] in the section called 'Implementation of MLC', there's a descrip
by 8bitsrule 3y ago
>I couldn't find a lot of information as to size of the network used
In the original paper,[0] in the section called 'Implementation of MLC', there's a description of sorts (Greek to me):
"Both the encoder and decoder have 3 layers, 8 attention heads per layer, input and hidden embeddings of size 128, and a feedforward hidden size of 512. Following GPT63, GELU64 activation functions are used instead of ReLU. In total, the architecture has about 1.4 million parameters."
[0] https://www.nature.com/articles/s41586-023-06668-3 https://www.nature.com/articles/s41586-023-06668-3