4 ms·
My understanding and suspicion is mostly less than you think. Llama 3 architecture has the following changes on GPT-2: 1. delete the absolute positional encodi
by karpathy 2y ago
My understanding and suspicion is mostly less than you think. Llama 3 architecture has the following changes on GPT-2:
1. delete the absolute positional encoding and replace with RoPE
2. delete all biases in all layers (in LayerNorms, they
turn into RMSNorm)
3. GeLU -> SwiGLU non-linearity in the MLP
4. longer context length
5. architecture hyperparameter changes, e.g. slightly different aspect ratios
And there was a paper that I can't find the reference to anymore that claimed that if you train long enough, the gap becomes even lower. Possibly because the absolutely positional encoding has enough time to train more fully, where as the RoPE layer benefits from the "inductive bias" it adds in the earlier stages of training.
But I don't have full confidence on the above claim, maybe someone has tried or has better/concrete reference.
- jorlow 2y agoNote llama's feed forward is a bit different too: self.w2(F.silu(self.w1(x)) * self.w3(x)) I.e. the nonlinearity is a gate. https://github.com/meta-llama/llama3/blob/14aab0428d3ec3a9596f1dea06d9c564f9c0e35f/llama/model.py#L219 https://github.com/meta-llama/llama3/blob/14aab0428d3ec3a959...
- soraki_soladead 2y agoFwiw, that's SwiGLU in #3 above. Swi = Swish = silu. GLU is gated linear unit; the gate construction you describe.