8 ms·
The root of the problem is that the methods used by LLMS (even deep learning for that matter) are considered by many academics as "the brute force approach" and
by rusticpenn 3y ago
The root of the problem is that the methods used by LLMS (even deep learning for that matter) are considered by many academics as "the brute force approach" and that the researchers are throwing GPUs at the problem. The speech recognition by deep neural networks bet almost every "intelligent" approach and a performance barrier which had been there for 30 years to be even considered good enough.
- benreesman 3y agoThey’re considered by many people whose experience is entirely industrial and zero academic to be trivially wasteful via the efficacy of quantization strategies directly or indirectly based on Hessian local geometry, to name but one example. It’s unclear to me what either the widely observed empirical observation that modern LLMs and their immediate offshoots (ViTs, all of it) are wildly entropically redundant via too many arguments to succinctly enumerate has to do with my arguments about the comparative efficacy of e.g. CERN and the way we do AI research in the large. As for academics, many academics regard calling AI scaling in an industrial or quasi-industrial setting “research” is generous in the extreme. It’s eye-popping engineering but we’re a long way from anyone so much as whispering about any Fields-level math coming out of “AI”.