4 ms·
Compute-efficient pretraining and scaling to trillion-parameter models
- simonw 24d ago> We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200. If this holds up that's a really big deal.
- wayfwdmachine 24d agoHuge if true. As it were.
- ismael_rr 24d agoSuper awesome. Wish they would release the paper about what they did to achieve this. I remember nous released the token superposition paper which improved pretraining FLOPs some, but not 50x: https://nousresearch.com/token-superposition https://nousresearch.com/token-superposition. Wondering if they also found some cool tokenization strategiesa
- pixl97 24d ago"Write a paper" < "Sell to a big AI lab for $$$" Going to be interesting to see what happens to discoveries like this in the future.
- monneyboi 24d agoImagine the sheer amount of power you could save by releasing the paper.
- speedgoose 24d agoBut thanks to the Jevon Paradox, the global power consumption would probably increase. https://en.wikipedia.org/wiki/Jevons_paradox https://en.wikipedia.org/wiki/Jevons_paradox
- aaroninsf 24d agoSolar is now not only the cheapest energy it's cheaper to intitially deploy than non-renewables. Doesn't mean we should waste energy; it does mean that we have crossed a threshold beyond which energy concerns change shape.
- pixl97 24d ago"Honey, why is there a solar panel-maximizer converting our car?"
- pvillano 24d agoSunlight no longer reaches the earth's surface but at least we know P vs NP
- pixl97 23d agoI mean P=NP might be a fair trade.
- corysama 23d agoYou just need to work your way down to a deeper layer in the https://en.wikipedia.org/wiki/Matrioshka_brain https://en.wikipedia.org/wiki/Matrioshka_brain Not the deepest layer. Humans would fry instantly there. But, the warm glow of the third layer down upon the fourth is quite pleasant.
- vatsachak 24d agoCool story. If it's true the company will be bought by open AI/Anthropic and Chinese labs will discover the trick and open source it by next quarter.
- FailMore 24d ago[flagged]
- asadm 24d agonot a good enough submarine
- pvillano 24d agoA lot of people are betting their money on infinite growth forever of AI performance, compute usage, user base, subscription price. I think cost will decrease forever.
- bilater 24d agoCost will keep dropping, but the frontier will keep getting pushed. The whole "give me today's model 10x cheaper and I'm good" line is a fallacy. It isn't true now and it never will be for the top 1% of tasks, which will create the most economic gains.
- adam_arthur 24d agoThere are an enormous number of tasks that can get by on good enough. If you need image recognition, and a 30B model saturates the use case with 100% accuracy, you absolutely wouldn't continue to use the next frontier model as they come out. And I'd argue most economically meaningful tasks will be saturated by cheaper models than those requiring frontier. Think about what today's models can do with pretty close to 100% accuracy, and then consider that they will be orders of magnitudes cheaper over the years. 5.6 Sol can already obviate tons of labor, and why would you pay 2x or more for no meaningful gain? The relative gap between frontier and non frontier also continues to shrink, so it's not like you take a meaningful performance loss by rewinding to models from 3-6 months ago. And soon that gap will expand to 12-24 months. I get the impression the majority of people on here only think about coding, which net net will be a tiny volume of overall AI use in the end.
- bilater 24d agoThat can all be true but the frontier models will still have a huge market. You're thinking of all tasks as a fixed pie. The top 1% of intelligence opens up a whole new pie, stuff nobody does today because it's too expensive: daily cancer scans instead of one every few years, asteroid mining missions that need ten thousand PhD-hours of planning, custom drugs designed for your specific tumor, a personal lawyer and doctor for every person on earth, auditing every line of code in every bank and hospital continuously and so on.
- ansk 24d agoI don't know enough about the specific models they're comparing against to say this definitively, but it looks to me like they're comparing their pre-trained models with others' post-trained models. The metric upon which their 10x claim is based (bits-per-byte) is exactly the metric which is optimized during pre-training. Post-trained models are fine-tuned to optimize other metrics, which is known to be detrimental to performance on bits-per-byte evaluations. So bits-per-byte evaluations will always make a pre-trained model look favorable in comparison to a comparable model which has also undergone post-training. Can someone confirm whether the models they are comparing against (DeepSeek V4, Kimi K2, and Nemotron 3 Ultra) have been post-trained?
- brrrrrm 24d agothey say they're looking at base models, so I think it's fairly compared as written.
- KaushikR2 23d agoSee footnote 2: > “Base models” are pretrained models that have not yet undergone reinforcement learning, SFT, or other post-training. They are highly sensitive to prompting, making sampling-based evals unreliable. Instead, we measured bits-per-byte loss on heldout data, which does not suffer from prompt sensitivity and smooths measurement of otherwise emergent abilities. As a side note, we were surprised that Nemotron 3 outperforms DeepSeek V4 Pro across the board but found this to be consistent across domains and inference engines. This might indicate that Nemotron’s weak performance on benchmarks after RL is due to weaker post-training, but the pretrain was ahead of Chinese open-weight competitors.
- brrrrrm 24d agothis is basically the only thing pre-training teams work on in labs. compute efficiency is the metric, the assumption that scaling = intelligence is considered a given.
- mohsen1 24d ago[dead]
- vkaku 24d agoThis is great. All algorithmic efficiencies are amazing! One thing I'd remind all scientists and the wonderful people here is this wonderful meme/line from Jurassic Park: "Your scientists were so preoccupied with whether they could they didn't stop to think if they should." What is the actual amount of data that needs to be pre-trained and what is not? Nobody has come up with great answers to this question, and I'm already seeing amazing 0.5b-2b parameter models working very well with n-Gram corpuses of data. So, how many parameters do you really need for a given workload?
- pvillano 24d agoImagine yourself the CEO of a big AI company. It takes about a month to develop and train a model, so you release a new model every month. A startup says they can 10x your efficiency. What does that get you? You can't release a new model every three days. You can't 10x R&D either. You definitely can't tell investors that you are growing at the same rate, but selling off assets and cancelling purchasing contracts. So you just never improve efficiency enough to use less energy than the previous model version. I don't believe this is actually happening.
- awestroke 24d ago> What does that get you? Cheaper model training runs? Ability to scale training to larger model sizes without extending training time?
- pvillano 24d agoYes, but a 10x larger model is only marginally better, and 10x cheaper training runs is only useful if you can find a use for 10x as many.
- aaronblohowiak 24d agomore experiments.
- BoorishBears 24d agoThis seems weirdly pessemistic: frontier labs have much stronger pretraining than most open weights models And reading the release it feels very obvious this is also a ton of aligning their data mix with coding and science: we don't know that this model doesn't have terrible world knowledge or is ruined for anything related to subjective preference They also repeatedly mention knowledge almost as if they saw that skepticism coming, but then limit knowledge to topics where more understanding of how code/scientific writing looks would produce the same graph as having actual world knowledge maintained. That's not nefarious (they literally build coding models), but it also means the resulting model isn't necessarily competitive with a frontier model in a broader way. This feels like the inverse approach to what Thinking Machines did with Inkling (trying to train as "un-spikey" a base model as possible)
- woadwarrior01 24d agoThey have a history of making grandiose claims like this[1] from 2024, with no visible products or research. [1]: https://magic.dev/blog/100m-token-context-windows https://magic.dev/blog/100m-token-context-windows (also linked to in their blogpost)
- riazrizvi 24d agoVery helpful thanks.