4 ms·
In the span of a few months, with a small team of researchers and engineers, we trained a 70B parameter model from scratch on our own infrastructure that outper
by thejash 2y ago
In the span of a few months, with a small team of researchers and engineers, we trained a 70B parameter model from scratch on our own infrastructure that outperformed zero-shot GPT-4o on reasoning-related tasks. Using our cluster for high performance training meant that every component — InfiniBand, Ethernet, GPUs, and the nodes themselves — had to work perfectly. If even a single one of the over 12,000 connections was a little flaky, it could slow down the entire training run.
We're sharing open-source scripts and an end-to-end guide for infrastructure set-up that details the process of making everything work perfectly, and ensuring that it stays that way.
This is one of a three-part toolkit on training a 70b model from scratch. The other two sections focus on evaluations and CARBS, our hyperparameter optimizer; you can find them here: https://imbue.com/research/70b-intro/ https://imbue.com/research/70b-intro/
Thoughts and questions welcome! :)
- ipsum2 2y agoWhat happened to the Minecraft-like 3d world your team built? Did you guys pivot?
- deleted 2y ago[deleted]
- Flumio 2y agoNice. Tx for the write up
- chx 2y ago> If even a single one of the over 12,000 connections was a little flaky, it could slow down the entire training run It's an unusual enough sentence to be remarkable and I was like "I read this exact same sentence before". Indeed, this and most of the writeup appeared on Twitter, LinkedIn, Reddit it seems word-by-word. Is this just spam ? https://x.com/imbue_ai/status/1805629547473518695 https://x.com/imbue_ai/status/1805629547473518695 https://reddit.com/r/learnmachinelearning/comments/1dobgbs/training_a_70b_model_from_scratch_open_source/ https://reddit.com/r/learnmachinelearning/comments/1dobgbs/t... https://www.linkedin.com/posts/mattboulos_training-a-70b-model-from-scratch-open-source-activity-7211401491293589507-UECW/ https://www.linkedin.com/posts/mattboulos_training-a-70b-mod...
- bottled_poe 2y ago[flagged]
- knowaveragejoe 2y agoHaving listened to the person who wrote this speak at length about the subject, it is not BS or grifting.
- leothetechguy 2y agoThe same company reports multiple times on a finding they've made through multiple social media channels? Shocking. /s
- neilv 2y agoI'd rather some company copy&paste the same text multiple places -- if the alternative was that those places would instead get obfuscation of the same information to appear novel each time (so I'd have to read all of them to realize they're all just the same info).
- lolinder 2y agoThis is the kind of criticism that could only come from someone without much formal writing experience. This is a very normal workflow: You write a full-length text detailing the project you worked on. You then trim it down to a summary which you share with a group of people X. You then trim it down into a different summary which you share with a group of people Y. When you do this multiple times you unsurprisingly end up with some sentences that make it into multiple summaries because they're that important to the thesis! (Also, the summaries on Twitter and Reddit aren't anything close to "most of the writeup"—the full text is 6000+ words!)
- chx 2y agoSo you agree this is a formally written PR piece copy-pasted and as such it's spam. Got it. What would've been not spam? Spontaneous writing with a link to the puff piece.
- 2y ago
- highfrequency 2y ago> outperformed zero-shot GPT-4o Cool stuff! Does this do RLHF or just pretraining? If the latter, how did you manage to beat GPT 4?
- vessenes 2y agoLoved this and the detail - thank you. It’s the best inside detail on the engineering work behind these models I’ve ever read. Two things I’m curious about- first, what, if any difference would you imagine in training a 400b parameter model? It seems that you have plenty of vram across the cluster, but I want to know what you think. Second, do you think this sort of architecture is the end game for model training? It seems sooo fragile. Are there better shared training mechanisms/architectures? Are there better cluster geometries? Thanks again - great read.