4 ms·
> Hardware: StableLM Zephyr 3B was trained on the Stability AI cluster across 8 nodes with 8 A100 80GBs GPUs for each nodes. I might be missing it but do they
by filterfiber 3y ago
> Hardware: StableLM Zephyr 3B was trained on the Stability AI cluster across 8 nodes with 8 A100 80GBs GPUs for each nodes.
I might be missing it but do they say the number of training tokens that was used to train this?
This would help with efforts like TinyLlama in trying to figure out how well the scaling works with training tokens vs parameter size and challenging the chinchilla model.
- emadm 3y agoWe included full training details for the base model on 4 trillion tokens including wandb etc https://stability.wandb.io/stability-llm/stable-lm/reports/StableLM-3B-4E1T--VmlldzoyMjU4?accessToken=u3zujipenkx5g7rtcj9qojjgxpconyjktjkli2po09nffrffdhhchq045vp0wyfo https://stability.wandb.io/stability-llm/stable-lm/reports/S...