17 ms·
Against LLM Maximalism
- mark_l_watson 3y agoI agree with much of the article. You do need to take great care to make code with embedded LLM use modular and easily maintainable, and otherwise keep code bases tidy. I am a fan of tools like LangChain that bring some software order to using LLMs. BTW, this article is a blog hosted by the company who writes and maintains the excellent spaCy library.
- passion__desire 3y agoIs anyone working on a OS LLM layer? e.g. consider a program like gimp. It would feed in its documentation and workflow details in LLM and get embeddings which would be installed with the program just like man-pages. Users could just express what they want to do in natural languages and Gimp would just query llm and create a workflow that might achieve the task.
- mark_l_watson 3y agoApple's CoreML is a large collection of regular models, deep learning models, etc. that are easy to use in macOS/iOS/iPadOS apps.
- __loam 3y ago> You do need to take great care to make code with embedded LLM use modular and easily maintainable, and otherwise keep code bases tidy. Sure makes sense. > I am a fan of tools like LangChain that bring some software order to using LLMs. Lmao. I feel like tools like LangChain that are really just very thin wrappers for the LLM APIs are quite complex for what they supposedly do for you. Lots of leaky abstractions and indirection for very little gained over just calling the APIs themselves.
- Grimburger 3y ago[flagged]
- davepeck 3y agoExplosion is an old school machine learning company by the people who built the spaCy natural language library. They’re serious practitioners whose work predates the “hype-train” you’re concerned about. The blog post might be worth a gander.
- Grimburger 3y ago> They’re serious practitioners From the article in question: https://explosion.ai/static/1863c4dfa57ad28dbbd68e432bde34e9/2523c/llm-maximalism_meme.webp https://explosion.ai/static/1863c4dfa57ad28dbbd68e432bde34e9...
- deleted 3y ago[deleted]
- davepeck 3y agoSerious practitioners are permitted to have fun and/or use goofy memes.
- Grimburger 3y agoSure, but readers of "serious" content are also permitted to be turned off by them and with expert prompt engineering content like this half written by the AI itself that it purports to explain, I think it's fair to be dismissive. I've done so much work for AI adjacent stuff now that I'm completely numb. There's very little left that is original at a small scale and the actual "good stuff" has a literal army of third worlders behind it working for $2/hr, on demand, for whatever needs adjusting as may be. There's a massive dark underbelly that no one wants to talk about, so let's just pretend it's all an api :|
- EGreg 3y agoI predicted that AI will be the next Web3 — hugely promising but increasingly ignored by HN. There will be waves of innovation in the coming years. Web3 solutions will mostly enrich people or at worst be zero-sum. While AI solutions will redistribute wealth from the working class to the top 1% and corporations, as well as giving people ways to take advantage of vulnerable people and systems at a scale never seen before.
- phillipcarter 3y agoSo I think this is an excellent post. Indeed, LLM maximalism is pretty dumb. They're awesome at specific things and mediocre at others. In particular, I get the most frustrated when I see people try to use them for tasks that need deterministic outputs and the thing you need to create is already known statically. My hope is that it's just people being super excited by the tech. I wanted to call this out, though, as it makes the case that to improve any component (and really make it production-worthy), you need an evaluation system: > Intrinsic evaluation is like a unit test, while extrinsic evaluation is like an integration test. You do need both. It’s very common to start building an evaluation set, and find that your ideas about how you expect the component to behave are much vaguer than you realized. You need a clear specification of the component to improve it, and to improve the system as a whole. Otherwise, you’ll end up in a local maximum: changes to one component will seem to make sense in themselves, but you’ll see worse results overall, because the previous behavior was compensating for problems elsewhere. Systems like that are very difficult to improve. I think this makes sense from the perspective of a team with deeper ML expertise. What it doesn't mention is that this is an enormous effort, made even larger when you don't have existing ML expertise. I've been finding this one out the hard way. I've found that if you have "hard criteria" to evaluate (i.e., getting the LLM to produce a given structure rather than an open-ended output for a chat app) you can quantify improvements using Observability tools (SLOs!) and iterating in production. Ship changes daily, track versions of what you're doing, and keep on top of behavior over a period of time. It's arguably a lot less "clean" but it's way faster, and because it's working on the real-world usage data, it's really effective. An ML engineer might call that some form of "online test" but I don't think it really applies. At any rate, there are other use cases where you really do need evaluations, though. The more important correct output is, the more it's worth investing in evals. I would argue that if bad outputs have high consequences, then maybe LLMs also aren't the right tech for the job, but that'll probably change in a few years. And hopefully making evaluations will be easier too.
- syllogism 3y ago(Author here) It's true that getting something going end-to-end is more important than being perfectionist about individual steps -- that's a good practical perspective. We hope good evaluation won't be such an enormous effort. Most of what we're trying to do at Explosion can be summarised as trying to make the right thing easy. Our annotation tool Prodigy is designed to scale down to smaller use-cases for instance ( https://prodigy.ai https://prodigy.ai ). I admit it's still effort though, and depending on the task, may indeed still take expertise.
- alexvitkov 3y agoSorry if this is a bit ignorant, I don't work in the space, but if a single LLM invocation is considered too slow, how could splitting it up into a pipeline of LLM invocations which need to happen in sequence help? Same with reliability - you don't trust the results of one prompt, but you trust multiple piped one into another? Even if you test the individual components, which is what this approach enables and this article heavily advocates for, I still can't imagine that 10 unreliable systems, which have to interact with rach other, are more reliable than one. 80% accuracy of one system is 80% accuracy. 95% accuracy on 10 systems is 59% accuracy in total if you need all of them to work and they fail independently.
- peter_l_downs 3y agoI think the idea behind breaking down the task into a composable pipeline is that you then replace the LLM steps in a pipeline with supervised models that are much faster. So you end up with a pipeline of non-LLM models, which are faster and more explainable.
- cmcaleer 3y ago> you don't trust the results of one promt, but you trust multiple piped one into another? This is really not at all unusual. Take aircraft for instance. One system is not reliable, for a multitude of reasons. A faulty sensor could be misleading, a few bits could get flipped by cosmic rays causing ECC to fail, the system itself could be poorly calibrated, there are far too many unacceptable risks. But add TMR[0][1] and suddenly you are able to trust things a lot more. This isn't to say that TMR is bullet proof e.g. incidents like [2], but redundancy does make it possible to increase trust in a system, and assign blame to what part of a system is faulty (e.g. if 3 systems exist, and 1 appears to be disagreeing wildly with 2 and 3, you know to start investigating system 1 first). Would it work here? I don't know! But it doesn't seem like an inherently terrible or flawed idea if we look at past applications. Ensembling different models is a pretty common technique to get better results in ML, and maybe this approach would make it easier to find weak links and assign blame. [0]: https://en.wikipedia.org/wiki/Triple_modular_redundancy https://en.wikipedia.org/wiki/Triple_modular_redundancy [1]: https://en.wikipedia.org/wiki/Air_data_inertial_reference_unit https://en.wikipedia.org/wiki/Air_data_inertial_reference_un... [2]: https://www.atsb.gov.au/media/news-items/2022/pitot-probe-covers https://www.atsb.gov.au/media/news-items/2022/pitot-probe-co... causing total confusion among the TMR
- peterldowns 3y agoSpacy [0] is a state-of-art / easy-to-use NLP library from the pre-LLM era. This post is the Spacy founder's thoughts on how to integrate LLMs with the kind of problems that "traditional" NLP is used for right now. It's an advertisement for Prodigy [1], their paid tool for using LLMs to assist data labeling. That said, I think I largely agree with the premise, and it's worth reading the entire post. The steps described in "LLM pragmatism" are basically what I see my data science friends doing — it's hard to justify the cost (money and latency) in using LLMs directly for all tasks, and even if you want to you'll need a baseline model to compare against, so why not use LLMs for dataset creation or augmentation in order to train a classic supervised model? [0] https://spacy.io/ https://spacy.io/ [1] https://prodi.gy/ https://prodi.gy/
- famouswaffles 3y ago>what I see my data science friends doing — it's hard to justify the cost (money and latency) in using LLMs directly for all tasks, and even if you want to you'll need a baseline model to compare against, so why not use LLMs for dataset creation or augmentation in order to train a classic supervised model? The NLP infrastructure and pipelines we have today aren't there because they are necessarily the best way to handle the tasks you want. They're in place because computers simply could not understand text the way we would like and shortcuts, approximations were necessary. Borrowing from the blog, Since you could not simply ask the computer, "How many paragraphs in this review say something bad about the acting? Which actors do they frequently mention?", separate processes of something like tagging names, linking them to a knowledge base, and paragraph-level actor sentiment etc were needed. The approximations are cool and they do work rather well for some use cases but they fall apart in many others. This is why automated resume filtering, moderation etc is still awful with the old techniques. You simply can't do what is suggested above and get the same utility.
- PheonixPharts 3y ago> why not use LLMs for dataset creation or augmentation in order to train a classic supervised model? Or, as I mentioned in another comment, just use the embeddings directly. This also does a lot to remove the "cost (money and latency)" part of the problem since you can batch queries to be lightening fast and the dollar cost of generating the embeddings is effectively zero (~3000 pages of text per $1) for most traditional NLP tasks that require a vector representation.
- sudb 3y agoI've had a fair amount of success at work recently with treating LLMs - specifically OpenAI's GPT-4 with function calling - as modules in a larger system, helped along powerfully by the ability to output structured data. > Most systems need to be much faster than LLMs are today, and on current trends of efficiency and hardware improvements, will be for the next several years. I think here I disagree with the author here though, and am happy to be a technological optimist - if LLMs are used modularly, what's to stop us in a few years (presumably still hardware requirement costs, on reflection) eventually having small, fast specialised LLMs for the things that we find them truly useful/irreplaceable?
- syllogism 3y agoNothing's to stop us, and in fact we can do that now! This is basically what the post advocates for: replacing the LLM calls for task-specific things with smaller models. They just don't need to be LLMs.
- famouswaffles 3y agoI'll just say there's no guarantee training or fine-tuning a smaller bespoke model will be more accurate (Certainly though, it may be accurate enough). Minerva and Med-Palm are worse than GPT-4 for instance.
- syllogism 3y agoThis is where the terminology being used to discuss LLMs today is a touch awkward and imprecise. There's a key distinction between smaller models trained with transfer-learning, and just fine-tuning a smaller LLM and still using in-context learning. Transfer learning means you're training an output network specifically for the task you're doing. So like, if you're doing classification, you output a vector with one element per class, apply a softmax transformation, and train on a negative log likelihood objective. This is direct and effective. Fine-tuning a smaller LLM so that it's still learning to do text generation, but it's better at the kinds of tasks you want to do, is a much more mixed experience. The text generation is still really difficult, and it's really difficult to learn to follow instructions. So all of this still really favours size.
- famouswaffles 3y agoRight that is a good distinction. Fair enough. Still stand that you could train a worse model depending on the task. Translation, Nuanced Classification are all instances where i've not seen bespoke models outright better than GPT-4. although, like i said it could still be good enough for speed, compute requirements.
- skybrian 3y agoI don’t understand this heuristic and I think it might be a bit garbled. Any idea what the author meant? How do you get 1000? > A good rule of thumb is that you’ll want ten data points per significant digit of your evaluation metric. So if you want to distinguish 91% accuracy from 90% accuracy, you’ll want to have at least 1000 data points annotated. You don’t want to be running experiments where your accuracy figure says a 1% improvement, but actually you went from 94/103 to 96/103.
- akprasad 3y agoMy guess is that this should be something like "If you have n significant digits in your evaluation metric, you should have at least 10^(n+1) data points."
- wrs 3y agoAvoiding the term “significant digits” completely: Distinguishing 91 vs 90 is a difference of 1 on a 0-100 scale. 100x10=1000. If you wanted to distinguish 91.0 vs 90.9, that’s 1 on a 0-1000 scale, so you’d want 10,000 points.
- forward-slashed 3y agoAll of this is quite difficult without the DSL to explore and construct pipelines for LLMs. Current approaches are very slow in terms of iteration.
- PheonixPharts 3y agoI personally still think most people (not necessarily the author) miss out on the biggest improvement LLMs have to offer: powerful embeddings for text representation for text classification. All of the prompting stuff is, of course, incredible, but the use of these models to create text embeddings of virtually any text document (from a sentence to a news paper article) allows for incredibly fast iteration on many traditional ML text classification problems. Multiple times I've taken cases where I have ~1,000 text documents with labels, run them through ada-002, and stuck that in a logistic model and gotten wildly superior performance to anything I've tried in the past. If you have an old NLP classification problem that you couldn't quite solve satisfactorily enough a few years ago, it's worth just mindlessly running it through the OpenAI embeddings API and sticking using those embeddings on your favorite off the shelf classifier. Having done NLP work for many years, it is insane to me to consider how many countless hours I spent doing tricky feature engineering to try to squeeze the most information I could out of the limited text data available, to realize it can now be replaced with about 10 minutes of programming time and less than a dollar. An even better improvement is the trivial ability to scale to real documents. It wasn't long ago that the best document models were just sums/averages of word embeddings.
- jiggawatts 3y agoI keep hearing about text classification but I can’t think of many specific use cases. If you don’t mind me asking: what are you using text classification for?
- stormfather 3y agoAsk ChatGPT and you'll get a ton. A few I've run into personally: TSLA is going to the moon. ^ Is this tweet bearish or bullish regarding the asset it mentions? (Acute) hepatitis C Hepatitis B; Acute ^ Do the above refer to the same disease? The Federal Reserve decides to abolish interest rates on leap years. Is it a leap year? New policy from the Fed says no interest if so. ^ Do these refer to the same news story, or different ones? So you can see that text classification is useful for consolidating and integrating streams of textual information, and extracting actionable meaning.