4 ms·
I find it hard to believe that anything like this will be feasible or effective beyond a certain level of complexity. It seems like a willful denial of the comp
by lsy 2y ago
I find it hard to believe that anything like this will be feasible or effective beyond a certain level of complexity. It seems like a willful denial of the complexity and ambiguity of natural language, and I am not looking forward to some poor developer trying to reason their way out of a two-hundred-step paradox that was accidentally created.
And for a use-case simple enough for this system to work (e.g. regurgitate a policy), it seems like the LLM is unnecessary. After all, if your system can perfectly interpret the question and answer and see if this rule set applies, then you can likely just use the rule set to generate the answer rather than wasting resources with a giant language model.
- vineyardmike 2y agoI don’t think this is a concern, but I do understand what you see, I think this really just is a new way for a computer to be the “bad guy” in customer support systems. First, they have a pretty low token limit for a “policy” so there won’t be anything too complex. Second, they explicitly say they don’t support synonyms. Seems very likely it’ll just reject anything that doesn’t fit closely, so you’ll end up with “I’m sorry. I don’t know what the ‘bought it’ date is, please provide purchase date?” Until the customer does the work of using the exact language. It looks like it takes a policy “returns must be processed within 30 days of purchase” and turns it into a pseudo-code type logic “if {purchase date} < {today-30d} => reject”. Then it seems to parse the LLM query and apply the logic. Considering my first two points, it’ll just be used to turn GPUs into another inhuman system to help companies avoid having to be human about customer support, while sounding more human.
- sdesol 2y agoI'm working on a rather naive approach that is focused on identifying errors in a LLM response by using LLMs. What I can share right now are screenshots with regards to how it works. The basic idea is you can use other high-quality models to validate and compare against to find irregularities or errors. You can see what it looks like below: https://app.gitsense.com/--/images/options.png https://app.gitsense.com/--/images/options.png https://app.gitsense.com/--/images/validate.png https://app.gitsense.com/--/images/validate.png https://app.gitsense.com/--/images/models.png https://app.gitsense.com/--/images/models.png The basic idea behind my chat system is, every model can be wrong, but it is unlikely that all will be wrong at the same time. This chat system is based on what I've learned when building my spelling and grammar checker. If you look at the following links, you can see that even the best models can get it wrong, but it is unlikely that others will get it wrong at the same time. https://app.gitsense.com/?doc=6c9bada92&model=GPT-4o&samples=5 https://app.gitsense.com/?doc=6c9bada92&model=GPT-4o&samples... https://app.gitsense.com/?doc=905f4a9af74c25f&model=Claude+3.5+Sonnet&samples=5 https://app.gitsense.com/?doc=905f4a9af74c25f&model=Claude+3...
- WhitneyLand 2y agoWhen will we be able to give it a try? I’m playing around with similar ideas, sometimes called ensembling techniques.
- sdesol 2y agoProbably in a couple of weeks. It's taken a while to finalize the UX but I know what it should look like now.
- nomel 2y ago> but it is unlikely that all will be wrong at the same time. Here's a prompt that proves this untrue, for now at least: > A woman and her biological son are gravely injured in a car accident and are both taken to the hospital for surgery. The surgeon is about to operate on the boy when they say "I can’t operate on this boy, he’s my biological son!" How can this be? Makes sense considering they're things of most-likely statistics, after all.
- thekyle 2y agoI tried this one with ChatGPT o1 and it seemed to get it right > The surgeon is the boy’s biological father. While the woman injured in the accident is the boy’s biological mother, the surgeon is his father, who realizes he cannot operate on his own son. https://chatgpt.com/share/674fc638-cd0c-8012-a4c4-9f1cad204054 https://chatgpt.com/share/674fc638-cd0c-8012-a4c4-9f1cad2040...
- fnordpiglet 2y agoClaude Sonnet also gets it right, but not reliably. It seems to be over aligned against gender assumptions and keeps assuming this is a gender assumption trick - that a surgeon isn’t necessarily male. This is probably the clearest case I’ve seen of alignment interfering with model performance.
- sdesol 2y agoI think anything requiring strong reasoning will probably have issues. However, I think most Enterprises is only interested in knowing that the summary of a document doesn't contain hallucinations, which I think most models will probably get right. If you go by a super majority rule and use 5 models, I think most business will be satisfied that the summary that it was given doesn't contain hallucinations. However, like you said, we are dealing with a non-deterministic system so the best we can hope for is a statistically likely answer.
- jimmySixDOF 2y ago> It seems like a willful denial of the complexity and ambiguity of natural language There is a paper and set of work recently that uses a measurement of entropy on the set of returned logits to detect a "certainty" estimate for outputs and flag hallucinations. It is a lot more rigorous than the OP but like everything in this space needs further testing.
- fzzzy 2y agoI've been thinking a lot about whether this would work lately. Do you have a link?
- jimmySixDOF 2y agoThis group of conversations is what I was thinking of: https://x.com/seb_far/status/1803446067343556872 https://x.com/seb_far/status/1803446067343556872 https://oatml.cs.ox.ac.uk/blog/2024/06/19/detecting_hallucinations_2024.html https://oatml.cs.ox.ac.uk/blog/2024/06/19/detecting_hallucin... Some people in the Open Source community I think implemented it a few months ago as Shrek Entropy Sampeling and it got a pop in circulation (https://github.com/xjdr-alt/entropix https://github.com/xjdr-alt/entropix) best of luck!
- fzzzy 2y agoThank you very much. I found this last night when looking into some of the keywords you mentioned. Looks like some of the same authors you referenced. Exciting! https://arxiv.org/abs/2406.15927 https://arxiv.org/abs/2406.15927