3 ms·
It's a model 40% of the size of Mistral's designed to be transparent (full training details, datasets & evals here: https://stability.wandb.io/stability-llm/sta
by emadm 3y ago
It's a model 40% of the size of Mistral's designed to be transparent (full training details, datasets & evals here: https://stability.wandb.io/stability-llm/stable-lm/reports/StableLM-3B-4E1T--VmlldzoyMjU4?accessToken=u3zujipenkx5g7rtcj9qojjgxpconyjktjkli2po09nffrffdhhchq045vp0wyfo https://stability.wandb.io/stability-llm/stable-lm/reports/S...) and work on edge devices.
There are improved versions coming but this is the best 3b model and Mistral is the best 7b model.
- brrrrrm 3y agohow much faster is 3b in practice? Seems like an uncommon size, so it would make sense to have the title of "best 3b" lol
- lhl 3y agoThe rule of thumb is that inference speed halves with every doubling of parameter size (and obviously a doubling of memory size). You can check out real world performances on devices here: https://llm.mlc.ai/ https://llm.mlc.ai/
- imjonse 3y agoI wonder if it's behind the subscription, but I see no reference source code for StableLM.
- omneity 3y agoThank you for sharing a well-written and transparent training report! Do you have plans to train/release a fine-tuned 3b chat version or other variants?