4 ms·
It required some leaked weights to start right?
by rusticpenn 3y ago
It required some leaked weights to start right?
- benreesman 3y agoDepends on how you count. Meta/FAIR realizing de-facto that open LLaMa was in their advantage is debatably causally related to everything that happened after, but there's a case people would have figured it out anyways. In this narrow instance it's not true at all: Mistral (pre-MSFT Mistral, we'll see going forward) released "Mixtral" 8x7B with de-facto "available weight" licensing and enough arch literature to make it the go-to for a long time. Dolphin is an Orca-style "non-operator de-alignment / operator-alignment" of that model, based on open work on this out of MSFT. And they did it on a very modest budget by the standards of foundation models in any modality but especially natural language and adjacent things. It's basically established at this point at least to my satisfaction that too much money hurts AI research. It's too easy to just throw money at it, and if you have that kind of backing, you have to show results on a very specific schedule (usually).
- rusticpenn 3y agoYes, my point is that there was a need/involvement of big companies to help the seed of these projects. Too much money definitely hurts any research as the "politicians" take over.
- benreesman 3y agoThere was a need for some form of funding from somewhere to allow LLM research certainly, and there was probably a need for someone to subsidize a few expensive runs that didn't pan out before it became common enough knowledge how to train a foundation LLM. But the cat's out of the bag on that now, you might not get Opus, but that capability band between e.g. GPT-3.5 and GPT-4 (and it's various previews) is crowded with stuff that is well-explained in significant detail in purely public sources. So IDK, yeah, maybe we needed big giant companies to kickstart this thing, maybe not, the government used to do effective public/private partnerships or even useful stuff directly with academia and in places still does: CERN did the LHC without like, crazy direct corporate sponsorship written all over it (there may be indirect stuff that I don't know about not being involved in the LHC project). And anyone who thinks AI is more complicated or should be more capital-intensive than high-energy particle physics done at CERN and places like it is getting high on their own supply.
- rusticpenn 3y agoThe root of the problem is that the methods used by LLMS (even deep learning for that matter) are considered by many academics as "the brute force approach" and that the researchers are throwing GPUs at the problem. The speech recognition by deep neural networks bet almost every "intelligent" approach and a performance barrier which had been there for 30 years to be even considered good enough.
- benreesman 3y agoThey’re considered by many people whose experience is entirely industrial and zero academic to be trivially wasteful via the efficacy of quantization strategies directly or indirectly based on Hessian local geometry, to name but one example. It’s unclear to me what either the widely observed empirical observation that modern LLMs and their immediate offshoots (ViTs, all of it) are wildly entropically redundant via too many arguments to succinctly enumerate has to do with my arguments about the comparative efficacy of e.g. CERN and the way we do AI research in the large. As for academics, many academics regard calling AI scaling in an industrial or quasi-industrial setting “research” is generous in the extreme. It’s eye-popping engineering but we’re a long way from anyone so much as whispering about any Fields-level math coming out of “AI”.