8 ms·
This work from Google (original Nature paper: https://www.nature.com/articles/s41586-021-03544-w https://www.nature.com/articles/s41586-021-03544-w) has been cr
by vighneshiyer 2y ago
This work from Google (original Nature paper: https://www.nature.com/articles/s41586-021-03544-w https://www.nature.com/articles/s41586-021-03544-w) has been credibly criticized by several researchers in the EDA CAD discipline. These papers are of interest:
- A rebuttal by a researcher within Google who wrote this at the same time as the "AlphaChip" work was going on ("Stronger Baselines for Evaluating Deep Reinforcement Learning in Chip Placement"): http://47.190.89.225/pub/education/MLcontra.pdf http://47.190.89.225/pub/education/MLcontra.pdf
- The 2023 ISPD paper from a group at UCSD ("Assessment of Reinforcement Learning for Macro Placement"): https://vlsicad.ucsd.edu/Publications/Conferences/396/c396.pdf https://vlsicad.ucsd.edu/Publications/Conferences/396/c396.p...
- A paper from Igor Markov which critically evaluates the "AlphaChip" algorithm ("The False Dawn: Reevaluating Google's Reinforcement Learning for Chip Macro Placement"): https://arxiv.org/pdf/2306.09633 https://arxiv.org/pdf/2306.09633
In short, the Google authors did not fairly evaluate their RL macro placement algorithm against other SOTA algorithms: rather they claim to perform better than a human at macro placement, which is far short of what mixed-placement algorithms are capable of today. The RL technique also requires significantly more compute than other algorithms and ultimately is learning a surrogate function for placement iteration rather than learning any novel representation of the placement problem itself.
In full disclosure, I am quite skeptical of their work and wrote a detailed post on my website: https://vighneshiyer.com/misc/ml-for-placement/ https://vighneshiyer.com/misc/ml-for-placement/
- s-macke 2y agoWhen I first read about AlphaChip yesterday, my first question was how it compares to other optimization algorithms such as genetic algorithms or simulated annealing. Thank you for confirming that my questions are valid.
- gdiamos 2y agoCriticism is an important part of the scientific process. Whichever approach ends up winning is improved by careful evaluation and replication of results
- jeffbee 2y agoIt seems like this is multiple parties pursuing distinct arguments. Is Google saying that this technique is applicable in the way that the rebuttals are saying it is not? When I read the paper and the update I did not feel as though Google claimed that it is general, that you can just rip it off and run it and get a win. They trained it to make TPUs, then they used it to make TPUs. The fact that it doesn't optimize whatever "ibm14" is seems beside the point.
- clickwiseorange 2y agoGood question. It's not just ibm14, but everything people outside Google tried shows that RL is much worse than prior methods. NVDLA, BlackParrot, etc. There is a strong possibility that Google pre-trained RL on certain TPU designs then tested in them, and submitted to Nature.
- Workaccount2 2y agoTo be fair, some of these criticisms are a few years old. Which normally would be fair game, but the progress in AI has been breakneck. Criticism of other AI tech from 2021 or 2022 are pretty dated today.
- jeffbee 2y agoIt certainly looks like the criticism at the end of the rebuttal that DeepMind has abandoned their EDA efforts is a bit stale in this context.
- anna-gabriella 2y agoDated or not, if half of the criticisms are right, the original paper may need to be retracted. No progress on RL for chip design was published by Google since 2022, as far as I can tell. So, it looks like most if not all criticisms remain valid.
- joshuamorton 2y ago> No progress on RL for chip design was published by Google since 2022, as far as I can tell. This makes sense given that both authors of the paper left Google in 2022. And one no longer seems to work in the chip design space, plausibly because of the bullying by entrenched folks. Then again, since rejoining Google the other author has produced around one patent per month in chip design with RL in 2023 and 2024, so perhaps they feel there is a marketable tool here that they don't want to share.
- deleted 2y ago[deleted]
- porphyra 2y agoThe Deepmind chess paper was also criticized for unfair evaluation, as they were using an older version of Stockfish for comparison. Apparently, the gap between AlphaZero and that old version of Stockfish (about 50 elo iirc) was about the same as the gap between consecutive versions of Stockfish.
- lacker 2y agoIndeed, six years later, the AlphaZero algorithm is not the best performing algorithm for chess. LCZero (uses AlphaZero algorithm) won some TCECs after it came out but for the past few years Stockfish (does not use AlphaZero algorithm) has been winning consistently. https://en.wikipedia.org/wiki/Top_Chess_Engine_Championship https://en.wikipedia.org/wiki/Top_Chess_Engine_Championship So perhaps the critics had a point there.
- vlovich123 2y agoThere’s a lot of codevelopment happening in the space where the positions are evaluated by Leela and then used to train the NNUE net within stockfish. And Leela comes from AlphaZero. So basically AlphaZero was directly responsible for opening up new avenues of research for a more specialized chess engine to reach new levels than it could have without it. > Generally considered to be the strongest GPU engine, it continues to provide open data which is essential for training our NNUE networks. They released version 0.31.1 of their engine a few weeks ago, check it out! [1] I’d say the impact AlphaZero has had on chess and go can’t be understated considering it’s a general algorithm that at worst is highly competitive with purpose built engines. And that’s ignoring the actual point of why DeepMind is doing any of this which is for GAI (that’s why they’re not constantly trying to compete with existing engines) [1] https://lichess.org/@/StockfishNews/blog/stockfish-17-is-here/7LRhV68w https://lichess.org/@/StockfishNews/blog/stockfish-17-is-her...
- theodorthe5 2y agoDo you understand StockFish filled the gap only after using NNs as well in the evaluation function? And that was a direct consequence of AlphaZero research.
- nemonemo 2y agoWhat is your opinion of the addendum? I think the addendum and the pre-trained checkpoint are the substance of the announcement, and it is surprising to see little mention of those here.
- dogleg77 2y agoFair point. I looked at the addendum, and it doesn't really address the critiques. They show an "ablation study" on one additional circuit without describing how big it is, etc. The few numbers they give suggest that this circuit is unlike those in the Nature paper: possibly a newer technology (good) but much smaller in the total length of wires and hence components (bad!). They are trying to debunk the Kahng work, but they aren't addressing many other complaints. Without showing results on public benchmarks in a reproducible way, they are just rehashing some excuses. Maybe their results are no good, maybe they stopped working on this. But with claimed runtimes under 6 hours, they don't need 3 years to add thorough benchmarking results. Anyone who designs competitive chips is doing benchmarking, but Google isn't. Draw your own conclusions.
- negativeonehalf 2y agoFD: I have been following this whole thing for a while, and know personally a number of the people involved. The AlphaChip authors address criticism in their addendum, and in a prior statement from the co-lead authors: https://www.nature.com/articles/s41586-024-08032-5 https://www.nature.com/articles/s41586-024-08032-5 , https://www.annagoldie.com/home/statement https://www.annagoldie.com/home/statement - The 2023 ISPD paper didn't pre-train at all. This means no learning from experience, for a learning-based algorithm. I feel like you can stop reading there. - The ISPD paper and the MLcontra paper both used much larger older technology node sizes, which have pretty different physical properties. TPU has a sub 10nm technology node size, whereas ISPD uses 45nm and 12nm. These are really different from a physical design perspective. Even worse, MLcontra uses a truly ancient benchmark with >100nm technology node size. Markov's paper just summarizes the other two. (Incidentally, none of ISPD / MLcontra / Markov were peer reviewed - ISPD 2023 was an invited paper.) There's a lot of other stuff wrong with the ISPD paper and the MLcontra paper - happy to go into it - and a ton of weird financial incentives lurking in the background. Commercial EDA companies do NOT want a free open-source tool like AlphaChip to take over. Reading your post, I appreciate the thoroughness, but it seems like you are too quick to let ISPD 2023 off the hook for failing to pre-train and using less compute. The code for pre-training is just the code for training --- you train on some chips, and you save and reuse the weights between runs. There's really no excuse for failing to do this, and the original Nature paper described at length how valuable pre-training was. Given how different TPU is from the chips they were evaluating on, they should have done their own pre-training, regardless of whether the AlphaChip team released a pre-trained checkpoint on TPU. (Using less compute isn't just about making it take longer - ISPD 2023 used half as many GPUs and 1/20th as many RL experience collectors, which may screw with the dynamics of the RL job. And... why not just match the original authors' compute, anyway? Isn't this supposed to be a reproduction attempt? I really do not understand their decisions here.)
- clickwiseorange 2y agoOh, man... this is the same old stuff from the 2023 Anna Goldie statement (is this Anna Goldie's comment?). This was all addressed by Kahng in 2023 - no valid criticisms. Where do I start? Kahng's ISPD 2023 paper is not in dispute - no established experts objected to it. The Nature paper is in dispute. Dozens of experts objected to it: Kahng, Cheng, Markov, Madden, Lienig, Swartz objected publically. The fact that Kahng's paper was invited doesn't mean it wasn't peer reviewed. I checked with ISPD chairs in 2023 - Kahng's paper was thoroughly reviewed and went through multiple rounds of comments. Do you accept it now? Would you accept peer-reviewed versions of other papers? Kahng is the most prominent active researcher in this field. If anyone knows this stuff, it's Kahng. There were also five other authors in that paper, including another celebrated professor, Cheng. The pre-training thing was disclaimed in the Google release. No code, data or instructions for pretraining were given by Google for years. The instructions said clearly: you can get results comparable to Nature without pre-training. The "much older technology" is also a bogus issue because the HPWL scales linearly and is reported by all commercial tools. Rectangles are rectangles. This is textbook material. But Kahng etc al prepared some very fresh examples, including NVDLA, with two recent technologies. Guess what, RL did poorly on those. Are you accepting this result? The bit about financial incentives and open-source is blatantly bogus, as Kahng leads OpenROAD - the main open-source EDA framework. He is not employed by any EDA companies. It is Google who has huge incentives here, see Demis Hassabis tweet "our chips are so good...". The "Stronger Baselines" matched compute resources exactly. Kahng and his coauthors performed fair comparisons between annealing and RL, giving the same resources to each. Giving greater resources is unlikely to change results. This was thoroughly addressed in Kahng's FAQ - if you only could read that. The resources used by Google were huge. Cadence tools in Kahng's paper ran hundreds times faster and produced better results. That is as conclusive as it gets. It doesn't take a Ph.D. to understand fair comparisons.
- smokel 2y agoI don't really understand all the fuss about this particular paper. Nearly all papers on AI techniques are pretty much impossible to reproduce, due to details that the authors don't understand or are trying to cover up. This is what you get if you make academic researchers compete for citation counts. Pretraining seems to be an important aspect here, and it makes sense that such pretraining requires good examples, which unfortunately for the free lunch people, is not available to the public. That's what you get when you let big companies do fundamental research. Would it be better if the companies did not publish anything about their research at all? It all feels a bit unproductive to attack one another.
- gabegobblegoldi 2y ago[dead]
- bsder 2y agoEDA claims in the digital domain are fairly easy to evaluate. Look at the picture of the layout. When you see a chip that has the datapath identified and laid out properly by a computer algorithm, you've got something. If not, it's vapor. So, if your layout still looks like a random rat's nest? Nope. If even a random person can see that your layout actually follows the obvious symmetric patterns from bit 0 to bit 63, maybe you've got something worth looking at. Analog/RF is a little tougher to evaluate because the smaller number of building blocks means you can use Moore's Law to brute force things much more exhaustively, but if things "looks pretty" then you've got something. If it looks weird, you don't.
- glitchc 2y agoThat doesn't mean the fabricated netlist doesn't work. I'm not supporting Google by any means, but the test should be: Does it fabricate and function as intended? If not, clearly gibberish. If so, we now have computers building computers, which is one step closer to SkyNet. The truth is probably somewhere in between. But even if some of the samples, with the terrible layouts, are actually functional, then we might learn something new. Maybe the gibberish design has reduced crosstalk, which would be fascinating.