3 ms·
Ultimately, what they needed to do was train to convergence, and they didn't do that. They should have made it extremely clear in their original manuscript the
by warbaker 4y ago
Ultimately, what they needed to do was train to convergence, and they didn't do that.
They should have made it extremely clear in their original manuscript their CT results are against a non-pretrained version that was not always trained to convergence, and so the CT results presented are a just a lower bound on performance.
They got confirmation from a Google engineer on only one block, which they trained for 1 million steps. For the other blocks, they didn't train anywhere close to that many steps, and their tensorboards don't show convergence on total cost.
ML is difficult! It's not enough to confirm that your implementation and installation are correct - you also need to train to convergence. You are free to complain that this is onerous or expensive. It's a common problem in academia vs industry in other areas of ML -- ML takes a lot of compute, and academic budgets are extremely tight. That said, CT is still orders of magnitude smaller / cheaper than most modern ML models (not surprising! CT is pretty old at this point by ML standards (3 years is a long time in this field)).
Kahng's FAQ admits that pre-training led to a better final proxy cost in the Nature paper, and that it significantly reduced compute time, and that this is the method that the Nature authors recommend and used in their Table 1 evals.
This whole situation is such a shame. Prof Kahng wrote a really great article on CT a few years back, and it gave me hope that the whole field was going to embrace ML and really take chip design to the next level: https://www.nature.com/articles/d41586-021-01515-9 https://www.nature.com/articles/d41586-021-01515-9
Most of the field in fact has, but Kahng himself seems to be backsliding here.
Unfortunately, I suspect that Kahng's been listening to some pretty toxic people (the same day his paper was posted to arxiv, it appeared in an SB author's court filing). It seems like his group ran into some difficulties jumping into machine learning, and rather than take that as an opportunity to grow and learn, they took that as an opportunity to attack.
While I'm glad that others in the field have built on the work successfully (see Appendix of https://www.annagoldie.com/home/statement https://www.annagoldie.com/home/statement ), it would be really tragic for Kahng himself to get left behind here, and so so unnecessary!
- ipso_facto 4y agoThe impact of pre-training is covered in Kahng's FAQ, "6. Did you use pre-trained models? How much does pre-training matter?" https://github.com/TILOS-AI-Institute/MacroPlacement#new-faqs-after-the-release-of-our-ispd-2023-paper-here-and-on-arxiv https://github.com/TILOS-AI-Institute/MacroPlacement#new-faq... It is also covered in the ISPD'22 paper, "Scalability and Generalization of Circuit Training for Chip Floorplanning", by Summer Yue, Ebrahim M. Songhori, Joe Wenjie Jiang, Toby Boyd, Anna Goldie, Azalia Mirhoseini, and Sergio Guadarrama, which was written by authors from the Nature paper. https://dl.acm.org/doi/pdf/10.1145/3505170.3511478 https://dl.acm.org/doi/pdf/10.1145/3505170.3511478 In their ISPD'22 paper, link above, Figure 7 shows a 5x speedup but 0.3% worse placement results with pre-training on diverse designs vs. running from scratch. Pre-training on previous versions of the same design does improve results by 1.7% vs. running from scratch, and has similar runtime benefit to pre-training with diverse designs. Given that the Circuit Training approach requires 41x to 646x wall time and an even higher factor of compute time vs. other standard techniques, the 5x speedup with pre-training is insufficient to close that very large gap. Furthermore, pre-training does not improve results vs. training from scratch, unless you are pre-training on previous versions of the same design but even then a 1.7% improvement in the circuit training results is not adequate to demonstrate improvement vs. standard techniques as can be seen by examining the results from Professor Kahng's group's work. This concern about pre-training has been answered in multiple forums, including the discussion in ISPD'23 Session 8, when Mr. Washauer-Baker raised the same issue. It's unfortunate to see Gabriel Warshauer-Baker, Anna Goldie's spouse, making repeated personal attacks on Professor Kahng. Please focus on the scientific discussion.
- warbaker 4y agoYou don't get to cut parts of the method and then claim to have compared against it! Sorry.
- true_north 3y ago[dead]