3 ms·
Please don't reduce LLM down to ChatGPT (or generative models more generally). People are using LLM for real-world problems every day. BERT and its descendants/
by armoredkitten 4y ago
Please don't reduce LLM down to ChatGPT (or generative models more generally). People are using LLM for real-world problems every day. BERT and its descendants/variants are used all over the place for many different problems in natural language processing. I and my team have used it on dozens of different projects, mainly in classifying text documents and inputs. And it works very well. Multilingual LLMs are responsible for the huge improvements in machine translation; my team has to deal with text in multiple languages, and these models are vital there too. We have used LLM on real-world problems that are in production now and are saving hundreds of person-hours of tedious work.
ChatGPT? Yeah, it's neat. I'm sure people will find some useful niche for it. And I do think generative models will eventually have a big impact, once researchers find good ways to ground them to data and facts. This is already an active area of research -- combining generative LLMs with info retrieval methods, or targeting it to a specific context. (Meta just gave a talk last week at the NeurIPS conference about teaching a model to play Diplomacy, a game that mostly involves talking and negotiating deals with the other players. ChatGPT is too broad for that -- they just need a model that can talk about the state of the game board.) So in general, I'm optimistic about generative LLMs. But ChatGPT...is just a toy, really. It's not the solution -- it's one of the signposts along the way toward the real solution. It's a measure of progress.
- jerrygenser 4y agoDo you feel like for some tasks though that people might be too eager to reach for GPU and more expensive inference and training costs? In some projects I'm working on, I'm seeing close to 95F1 score using deep learning to do token classification based NER. However, using non-deep learning approaches (still training statistical models) I can get to 91F1 on my use case but have a much faster inference and not need to use a GPU. I'm similarly optimistic about LLM, but fear that a lot of other really useful workhouse algorithms and strategies are going to be pushed aside and people will forget about them. Back to my use case, some strategies I'm looking at are using cheap and fast CPU powered models/smaller models for inference and then based on certain signals decide whether a particular instance should be passed to a GPU based model for better accuracy.
- armoredkitten 4y agoI don't know if I'm too concerned about that, to be honest. Yeah, there's a huge cost in terms of training the LLMs, and then there can be a cost for downstream inference, but I think it depends on the use case. In some cases, performance is the absolute top priority; in other cases, you might be willing to trade off some performance for better inference time, or model size, etc. If you need to put the model on cell phones or offline low-power devices, that's a key constraint that might make you reach for a different tool. The nice thing about more "classical" approaches -- a simple BoW random forest or MLP, for example -- is that they're typically quick to train and experiment with, and they make for great baselines, if nothing else. So I doubt that we're in danger of people forgetting about them entirely. If people do, they're leaving quick, easy solutions on the table. I do like your idea about triaging inference between smaller CPU vs. larger GPU models based on whatever signals. I haven't tried that before, but a project my colleagues worked on did some triaging between regex pattern-matching vs. model inference. Basically, the regex pulled some of the data out first if it matched very specific, known patterns, and then the rest was handled probabilistically. I guess the effectiveness of that sort of triaging approach depends on how strong and clear your signals are that let you choose one path over the other.
- hodgesrm 4y agoI wouldn't undersell ChatGPT. It's like a repl for a particular LLM. Maybe there are others but it's the first time many people have gotten direct access to the technology. Sometimes the medium is the message.
- armoredkitten 4y agoThat's fair -- perhaps we could frame it as a large-scale beta test of sorts. Researchers are building LLMs to solve problems, but new technologies can often end up solving problems they were never designed for. Once people get their hands on them, they test and tinker and find new uses for them. Sometimes it turns out not to be a good solution to the initial problem, but a great solution for something completely different. For instance, while I'm still generally of the opinion that generative models have limited use unless they're grounded to reality...I did see a post on Reddit about someone using ChatGPT to generate story ideas for their D&D game. So yeah...don't need to be tethered to reality to make a fantasy story! That's not something I would have thought of (even though I'm a DM!), and it's still relatively niche, but it's a great story of how getting something into people's hands to play with can generate lots of new ideas.