5 ms·
Thanks for sharing this. How do you think the local LLM movement will evolve? Especially as in the post, you mentioned startups and VCs both hoarding GPUs to at
by mchiang 3y ago
Thanks for sharing this. How do you think the local LLM movement will evolve? Especially as in the post, you mentioned startups and VCs both hoarding GPUs to attract talent.
There seems to be a good demand behind tools like llama.cpp or ollama (https://github.com/jmorganca/ollama https://github.com/jmorganca/ollama) to run models locally.
Maybe as the local runners become more efficient, we'll start seeing more trainings for smaller models or fine-tuning done locally? I too am still trying to wrap my head around this.
- CuriouslyC 3y agoThe current generation of models (Llama/Llama2) seem to pass a threshold of "good enough" for the majority of use cases at 60b+ parameters. Quantized 30b models that can run in 24gb of GPU VRAM are good enough for many applications but definitely show their limitations frequently. It is likely that we will eventually see good fine tunes for Llama 30b that produce usable output for code/other challenging domains, but until we get GPUs with 48g+ VRAM we're going to have to make do with general models that aren't great at anything, and fine tunes that only do one very narrow thing well.
- swyx 3y agoapart from the obvious GGML, we've done podcasts with both MLC/TQ Chen (https://www.latent.space/p/llms-everywhere https://www.latent.space/p/llms-everywhere) and Tiny/George Hotz (https://www.latent.space/p/geohot https://www.latent.space/p/geohot) who are building out more tooling for the Local LLM space! there's actually already a ton of interest, and arguably if you go by the huggingface model hub it's actually a very well developed ecosystem.. just that a lot of the usecases tend to be NSFW oriented. still, i'm looking to do more podcasts in this space, please let me know if any good guests come to mind.
- FanaHOVA 3y agoTraining smaller models can be really compute intensive (a 7B models should get trained on 1.4T tokens to follow the "LLaMA laws"). So that would be C = 6 * 1.4T * 7B = 58.8T FLOP-seconds. That's 1/5th the compute of GPT3 for example, but it's still a lot. We asked Quentin to do a similar post but for fine tuning math; that's still a very underexplored space. (Not to self plug too much, but this is exactly what last episode's with Tianqi Chen was about if you're interested :) https://www.latent.space/p/llms-everywhere#details https://www.latent.space/p/llms-everywhere#details
- swyx 3y agothat said there's more data efficiency to be gained in the smol models space - Phi-1 achieved 51% on HumanEval with only a 5.3x token-param ratio: https://www.latent.space/p/cogrev-tinystories#details https://www.latent.space/p/cogrev-tinystories#details
- arugulum 3y agoI want to jump in and correct your usage of "LLaMA Laws" (even you are using it informally, but I just want to clarify). There is no "LLaMA scaling law". There are a set of LLaMA training configurations. Scaling laws describe the relationship between training compute, data, and expected loss (performance). Kaplan et al., estimated one set of laws, and the Chinchilla folks refined that estimate (mainly improving it by adjusting the learning rate schedule). The LLaMA papers do not posit any new law nor contradict any prior one. They chose a specific training configuration that still abide by the scaling laws but with a different goal in mind. (Put another way: a scaling law doesn't tell you what configuration to train on. It tells you what to expect given a configuration, but you're free to decide on whatever configuration you want.)
- FanaHOVA 3y agoYep, +1. That's why I used the quotes. :) Thanks for expanding!
- arugulum 3y agoYep I understood that you were using it informally, just trying to keep things informative for other folks reading too.
- swyx 3y agothere frankly needs to be a paper calling this out tho, because at this point there are a bunch of industry models following “llama laws” and nobody’s really done the research, its all monkey see monkey do