3 ms·
Yes, my point is that there was a need/involvement of big companies to help the seed of these projects. Too much money definitely hurts any research as the "po
by rusticpenn 3y ago
Yes, my point is that there was a need/involvement of big companies to help the seed of these projects.
Too much money definitely hurts any research as the "politicians" take over.
- benreesman 3y agoThere was a need for some form of funding from somewhere to allow LLM research certainly, and there was probably a need for someone to subsidize a few expensive runs that didn't pan out before it became common enough knowledge how to train a foundation LLM. But the cat's out of the bag on that now, you might not get Opus, but that capability band between e.g. GPT-3.5 and GPT-4 (and it's various previews) is crowded with stuff that is well-explained in significant detail in purely public sources. So IDK, yeah, maybe we needed big giant companies to kickstart this thing, maybe not, the government used to do effective public/private partnerships or even useful stuff directly with academia and in places still does: CERN did the LHC without like, crazy direct corporate sponsorship written all over it (there may be indirect stuff that I don't know about not being involved in the LHC project). And anyone who thinks AI is more complicated or should be more capital-intensive than high-energy particle physics done at CERN and places like it is getting high on their own supply.
- rusticpenn 3y agoThe root of the problem is that the methods used by LLMS (even deep learning for that matter) are considered by many academics as "the brute force approach" and that the researchers are throwing GPUs at the problem. The speech recognition by deep neural networks bet almost every "intelligent" approach and a performance barrier which had been there for 30 years to be even considered good enough.
- benreesman 3y agoThey’re considered by many people whose experience is entirely industrial and zero academic to be trivially wasteful via the efficacy of quantization strategies directly or indirectly based on Hessian local geometry, to name but one example. It’s unclear to me what either the widely observed empirical observation that modern LLMs and their immediate offshoots (ViTs, all of it) are wildly entropically redundant via too many arguments to succinctly enumerate has to do with my arguments about the comparative efficacy of e.g. CERN and the way we do AI research in the large. As for academics, many academics regard calling AI scaling in an industrial or quasi-industrial setting “research” is generous in the extreme. It’s eye-popping engineering but we’re a long way from anyone so much as whispering about any Fields-level math coming out of “AI”.