4 ms·
We're talking 16 GPUs for ~6 hrs for inference, and 48 hrs for pre-training. This is not an exorbitant amount of compute. A GPU costs $1-2/hr on the cloud mark
by wholehog 2y ago
We're talking 16 GPUs for ~6 hrs for inference, and 48 hrs for pre-training. This is not an exorbitant amount of compute.
A GPU costs $1-2/hr on the cloud market. So, ~$100-200 for inference, and ~$800-1600 for pre-training, which amortizes across chips. Cloud prices are an upper bound -- most CS labs will have way more than this available on premises.
In an industry context, these costs are completely dwarfed by the rest of the chip design process. (For context, the licensing costs alone for most commercial EDA software are in the millions of dollars.)
- bushbaba 2y agoh100 GPU instances are multiple orders of magnitude more expensive.
- radq 2y agoNot true, H100s cost $2-3/GPU/hr on the open market.
- menaerus 2y agoYes, they even do at $1/GPU/hr. However, 8xH100 cluster at full utilization is ~8kWh of electricity and costs almost ~0.5M$. 16xH100 cluster is probably 2x of that. How many years before you break-even at ~24$/GPU/day income?
- Jabbles 2y ago7 https://www.google.com/search?q=0.5e6%2F8%2F24%2F365 https://www.google.com/search?q=0.5e6%2F8%2F24%2F365
- menaerus 2y agoDid you really not understand rethoric nature of my question and assumed that I can't do 1st grade primary school math?
- solidasparagus 2y agoWho cares? That's someone else's problem. I just pay 2-3$/hr and the H100s are usable
- menaerus 2y agoLook, I understand that some people are short-sighted and can hardly think out of the box and that is totally fine by me. I don't judge you for being that so I kindly ask you not to judge my question. Learn to give some benefit of the doubt.
- sangnoir 2y agoYou should care about counterparty risks. If your business model depends on unsustainable 3rd party prices powered by VC largesse and unrealizable dreams of dominance, the very least you can do is plan for the impending reckoning, after which GPU proces will be determined by costs.
- YetAnotherNick 2y agoH100 GPUs are more or less similar in price/performance. It is 2-3x more expensive per hour for 2-3x higher performance.
- vighneshiyer 2y agoYou are correct. For commercial use, the GPUs used for training and fine-tuning aren't a problem financially. However, if we wanted to rigorously benchmark AlphaChip against simulated annealing or other floorplanning algorithms, we have to afford the same compute and runtime budget to each algorithm. With 16 GPUs running for 6 hours, you could explore a huge placement space using any algorithm, and it isn't clear if RL will outperform the other ones. Furthermore, the runtime of AlphaChip as shown in the Nature paper and ISPD was still significantly greater than Cadence's concurrent macro placer (even after pre-training, RL requires several hours of fine-tuning on the target problem instance). Arguably, the runtime could go down with more GPUs, but at this point, it is unclear how much value is coming from the policy network / problem embedding vs the ability to explore many potential placements.
- Jabbles 2y agoYou're saying that if the other methods were given the equivalent amount of compute they might be able to perform as well as AlphaChip? Or at least that the comparison would be fairer? Are the other methods scalable in that way?
- pclmulqdq 2y agoYes, they are. The other approaches usually look like simulated annealing, which has several hyperparameters that control how much computing is used and improve results with more compute usage.
- wholehog 2y agoSee my comment above - the Nature authors already did this, and tried a huge hyperparameter sweep for SA, and RL still won. See appendix of the Nature article: rdcu.be/cmedX
- pclmulqdq 2y agoI understand and have read the article. Running 80 experiments with a crude form of simulated annealing is at most 0.0000000001% of the effort that has been spent on making that kind of hill climb work well by traditional EDA vendors. That is also an in-sample comparison, where I would believe the Google thing pre-trained on Google chips would do well, while it might have a harder time with a chip designed by a third party (further from its pre-training). The modern versions of that hill climb also use some RL (placing and routing chips is sort of like a game), but not in the way Jeff Dean wants it to be done.