13 ms·
Classifying customer messages with LLMs vs traditional ML
- rossirpaulo 3y agoThis is great! We had a similar thought and couldn't agree more with "LLMs prefer producing something rather than nothing." We have been consistently requesting responses in JSON format, which, despite its numerous advantages, sometimes imposes an obligation for an output even if it shouldn't. This frequently results in hallucinations. Encouraging NULL returns, for example, is a great way to deal with that.
- galleywest200 3y agoHave you tried using GPT-4s new Function Call feature? The "killer" portion of this is guaranteed JSON based on a schema you pass to the model.
- hellovai 3y agoThat's a good point! We're actually working on integrating this as well, but in practice, what we've found is that LLM's in general don't like to respond with empty strings for example. My hypothesis here is that due to RLFH, there's likely some implicit learning that tangentially related content is better than no content. Given that, you'd likely still get better results with your schema being: "string | null" so the LLM can output a null instead of "" since there is probably not as much training data that gives "" high log prob values. But we're looking forward to evaluating the functions call, and seeing what the metrics show!
- guhidalg 3y agoI integrated the function calling feature into my personal project and wrote a blog post about it here: https://letscooktime.com/Blog/ai,/machine/learning,/chatgpt,/ocr/2023/06/27/chatgpt.html https://letscooktime.com/Blog/ai,/machine/learning,/chatgpt,... Hopefully this saves you some time!
- CallMeMarc 3y agoThanks for the post! Really liked it being short and precise to the point. Also looking to integrate the new function feature and now already got some learnings out of the post without even starting to code.
- Der_Einzige 3y agoConstrained generation should not require calling supplemental functions. It's as simply as banning or reducing the weight of the naughty tokens. There are several libraries which enable this without function calling (microsoft guidance, jsonformer, lmql)
- msp26 3y agoThe output is not 100% guaranteed. Be careful about that and have another layer to check the output. I had a schema with a string enum property to categorise some inputs. One of the category names was "media/other" or something to that effect. Sometimes the output would stop at just media even though it wasn't a valid option in the schema.
- rolisz 3y agoNope, it's not guaranteed. They warn you in the OpenAI docs that it might hallucinate inexistent parameters.
- caesil 3y agoI've found that this is best dealt with along two axes with constrained options. i.e., request both a string and a boolean, and if you get boolean false you can simply ignore the string. So when the LLM ignores you and prints a string like "This article does not contain mention of sharks", you can discard that easily. If you tell it "Return what this says about sharks or nothing if it does not mention them", it will mess up.
- LawTalkingGuy 3y agoHave you tried this sort of prompt? User text: "Blah blah ... Sharks ... Surfing ..." Instruction: Return an JSON object containing an array of all sentences in the user text which mention sharks directly or by implication. Response: {"list_of_shark_related_sentences": [ Stop token: ']}' It'll try to complete the JSON response and it'll try to end it by closing the array and object as shown in the stop token. This severely limits rambling, and if it does add a spurious field it'll (usually) still be valid JSON and you can usually just ignore the unwanted field. wrt OpenAI, text-davinci-003 handles this well, the other models not so much.
- dontupvoteme 3y agoMaking it rank multiple attributes on a scale of 1-10 also works decent in my experience. Then one can simply k-means cluster (or similar) and evaluate the grouping to see how accurate its estimations are
- caesil 3y agoYes, agreed. I'm doing this as well. Works excellently for NLP classifier tasks. Funnily enough, there is a certain propensity for it to output round numbers (50, 100, etc.) so I have to ask it not to do this and provide examples ("like 27, 63, or 4"). Now that I think about it I should probably randomize those.
- dontupvoteme 3y agoInteresting, I've just been doing 1-10 (maybe i should include 0) -- Do you get the same result if you floatify the larger integers, e.g. 0.000 - 10.000?
- com2kid 3y agoI've run into the same issue, but you can turn it into an advantage if you are careful enough. Basically, give the LLM a schema that is loose enough for the LLM to expand where it feels expansion is needed. Saying always "return a number" is super limiting if the LLM has figured out you need a range instead. Saying "always populate this field" is silly because sometimes the field doesn't need to be populated.
- crazygringo 3y agoThis is really interesting. I'm really wondering when LLM's are going to replace humans for ~all first-pass social media and forum moderation. Obviously humans will always be involved in coming up with moderation policy and judging gray areas and refining moderation policy... but at what point will LLM's do everything else more reliably than humans? 6 months from now? 3 years from now?
- zht 3y agothis is some black mirror stuff imagine Google's general approach to customer service/moderation, but applied all over the place by companies small and large I shudder at the thought
- ghaff 3y agoEspecially with fairly systemic labor shortages, it seems inevitable that we'll see more and more self-service and automation with the corollary that getting an actual human involved will become more difficult.
- Xenoamorphous 3y agoI’ve found that it’s pretty much impossible to talk to a person in most customer services in the past few years, it’s always a “robot”. And this has been going since well before LLMs.
- HWR_14 3y agoI find it easy to talk to a person. A US based person, no. A person who can resolve my problem, maybe. But a person, yes.
- hypothesis 3y agoYup, most companies are now simply providing “emotional” support using sufficiently apologetic, but powerless agents. That is, if you can defeat chat bot screen or IVR maze. The hope is for user to give up..
- i-am-agi 3y agoWohoo this is amazing! I have been using the Autolabel (https://news.ycombinator.com/item?id=36409201 https://news.ycombinator.com/item?id=36409201) library so far for labeling a few classification and question answering datasets and have been seeing some great performance. Would be interested in giving gloo a shot as well to see if it helps performance further. Thanks for sharing this :)
- deleted 3y ago[deleted]
- alexmolas 3y agoWhere's the comparison with traditional ML? In the article I only see the good things about using LLM, but there's no mention to traditional ML besides from the title. It would be nice to see how compares this "complex" approach against a "simple" TF-IDF + RF or SVM.
- jonathankoren 3y agoYeah, I also find the lack of the comparison suspicious. As is the talk about “hallucinated class labels” being “helpful”. If I had to take a guess, I suspect the LLM might perform a touch better, but we’re taking fractional percent better. Which is fine, if you have the volume, but a wash otherwise
- famouswaffles 3y agohttps://news.ycombinator.com/item?id=36685921 https://news.ycombinator.com/item?id=36685921
- specproc 3y agoYeah, my thoughts exactly. If you're running 500k in tokens through through someone else's hallucination-prone computer and paying for the privilege, I want to know why that's any better than something like SetFit. All I saw were attempts to reproduce some chatgpt output.
- espe 3y ago+1 for setfit. a baseline that's hard to beat.
- hellovai 3y agoSetFit is fairly good, and we do help train SetFit like models for the results you get, however, the issue with SetFit is that its latency and cost benefits come at the price of flexibility. If you want to add a new class, to update an existing one, it requires training a new model. Sometimes this is ok and sometimes it's not. This is why we generally prefer a hybrid approach where some classes are using traditional models (BERT based) while others are determined by the LLM.
- wilg 3y agoClassic HN website nitpick: Logo should link to home page. In this case it is a link but just goes to the current page. However, points for being able to easily get to the main product page from the blog, usually that's buried.
- hellovai 3y agooh! Good catch! Fixed this, and will update in the release.
- r_singh 3y agoI have been using LLMs for ABSA, text classification and even labelling clusters (something that had to be done manually earlier on) and I couldn't be happier. It was turning out to be expensive earlier but with optimising the prompt a lot, reduced pricing by OpenAI and now also being able to run Guanaco 13/33B locally has made it even more accessible in terms of pricing for millions of pieces of text.
- hellovai 3y agoThat's very interesting! What sort of direction did you head in with prompt optimization? Was it mostly in shrinking it and then using multi-shot examples? We found that shorter prompts (empirically) perform better than longer prompts.
- nestorD 3y agoLLMs are significantly slower than traditional ML, typically costlier and, I have been told, tend to be less accurate than a traditional model trained on a large dataset. But, they are zero/few shot classifiers. Meaning that you can get your classification running and reasonably accurate now, collect data and switch to a fine-tuned very efficient traditional ML model later.
- hellovai 3y agoThat's a great summary and insight. We should likely use that verbiage to help make it more crystal clear :)
- famouswaffles 3y agoCurrent State of the art (GPT-4) is not going to be less accurate than whatever bespoke option you can cook up. https://news.ycombinator.com/item?id=36685921 https://news.ycombinator.com/item?id=36685921
- withinboredom 3y agoI wouldn't be so sure of that.
- golergka 3y agoIt will be much slower and costlier.
- mplewis 3y agoThis is absolutely untrue.
- famouswaffles 3y agoFeel free to show otherwise
- chaxor 3y agoI think they mean that it is likely to heavily depend upon the task. Decoder models are good at generation of language, and of course they're going to do well where that counts. But if you want to do typical NER+Relation extraction and then normalize to an in-house dictionary of 10 million IDs? You can't do that as effectively with GPT-4. You need a local model (right now). There are a lot of things that GPT-4 doesn't touch in terms of data domain, so for many projects, yes local (typically encoder) models are 1) way faster 2) way" cheaper and 3) have better metrics. Of course, the capability could* be there with a large decoder model, but if it's a task that needs your specific (large amount of) data, you have to make something locally.
- rckrd 3y agoI just released a zero-shot classification API built on LLMs https://github.com/thiggle/api https://github.com/thiggle/api. It always returns structured JSON and only the relevant categories/classes out of the ones you provide. LLMs are excellent reasoning engines. But nudging them to the desired output is challenging. They might return categories outside the ones that you determined. They might return multiple categories when you only want one (or the opposite — a single category when you want multiple). Even if you steer the AI toward the correct answer, parsing the output can be difficult. Asking the LLM to output structure data works 80% of the time. But the 20% of the time that your code parses the response fails takes up 99% of your time and is unacceptable for most real-world use cases. [0] https://twitter.com/mattrickard/status/1678603390337822722 https://twitter.com/mattrickard/status/1678603390337822722
- caycep 3y agowhat's "traditional ML"?
- 19h 3y agoWe’re classifying gigabytes of intel (SOCMINT / HUMINT) per second and found semantic folding or better in classification quality vs throughput than BERT / LLMs. How it works — imagine you’re having these sentences: “Acorn is a tree” and “acorn is an app” You essentially keep record of all word to word relations internal to a sentence: - acorn: is, a, an, app, tree Etc. Now you repeat this for a few gigabytes of text. You’ll end up with a huge map of “word connections”. You now take the top X words that other words connect to (I.e. 16384). Then you create a vector of 16384 connections, where each word is encoded as 1,0,1,0,1,0,0,0, … (1 is the most connected to word, 0 the second, etc. 1 indicates “is connected” and 0 indicates “no such connection). You’ll end up with a vector that has a lot of zeroes — you can now sparsify it (I.e. store only the positions of the ones). You essentially have fingerprints now — what you can do now is to generate fingerprints of entire sentences, paragraphs and texts. Remove the fingerprints of the most common words like “is”, “in”, “a”, “the” etc. and you’ll have a “semantic fingerprint”. Now if you take a lot of example texts and generate fingerprints off it, you can end up with a very small amount of “indices” like maybe 10 numbers that are enough to very reliably identify texts of a specific topic. Sorry, couldn’t be too specific as I’m on the go - if you’re interested drop me a mail. We’re using this to categorize literally tens of gigabytes per second with 92% precision into more than 72 categories.
- lgas 3y agoNot to dogpile on all the other "isn't this just" messages, but isn't this just sparse embeddings?
- lmeyerov 3y agoYeah I'm struggling to understand at a fundamental level how this is better both in math + engineering, doubly so by the time you get to sentence embeddings . (Genuinely, it seems to use the same ideas, so curious what the specific trick is vs mature embedding packagings already doing much of this afaict.)
- mmcwilliams 3y agoYou're not wrong. This sounds curiously close to the ways I've seen word2vec used in production.
- YetAnotherNick 3y agoInterested in knowing how you are running BERT model with $35/month? Cheapest GPU instance costs $200-300/month AFAIK.
- aaronvg 3y ago(Other author of this blog post here) We actually do CPU inference. The SBERT models have a pretty small memory footprint -- you can fit a couple models on a t2.medium instance. On a C6.Large you can get 75ms inference. T2.medium is more around 100-200ms
- m3kw9 3y agoProb cheaper with ML but you need training, with transfer learning though you can use a pub trained model and use way less data to train up a classifier like single digit thousands may be ok with 2-5 sentiments
- avereveard 3y agoOne can use the LLM to generate the label to distill a model to the desired precision, I used that approach and it worked quite well and the model runs locally, including creating the sentence embeddings, faster than the LLM, at a fraction of the cost. Now certain problem space may be large enough to require models where the runtime makes it non economical to run it locally, but Ml is still a game of heuristics, see each problem requires some experimentation.
- Animats 3y agoWhat's the application? If you're using this to direct messages to approximately the correct department, it doesn't have to be that complicated. If you're doing this to evaluate customer sentiment, you could probably just select a few hundred messages at random and read them. (There are many "big data" problems which are only big due to not sampling.)
- andrewgazelka 3y agoMy understanding was training on ChatGPT output was against OpenAI ToS. Is this incorrect for this use case (training BERT)?
- katrinaroberta 3y ago[dead]