4 ms·
For what it pertains the finances, we simply will have to agree to disagree here. To be convinced I would require detailed financial data that you cannot and sh
by ADeerAppeared 2y ago
For what it pertains the finances, we simply will have to agree to disagree here. To be convinced I would require detailed financial data that you cannot and should not share with random strangers.
But to say something useful, let me try to elaborate my general criticism here:
> Prior systems either dumped massive amounts of cognitive load in the investigators face or took man years of effort to create a specific workflow, and in an adversarial dynamic space like fraud you need a much more dynamic approach to different types of new attacks.
This begs a question: Why didn't a computer system to summarize this data already exist? Or rather, what stopped the prior systems from doing this work? (And I'll consider conventional machine learning; classifiers and the like, as traditional computer systems here)
And there's generally two options here:
1. Conventional computer systems absolutely could do this work, but they just haven't been built. (Say, because nobody signed off on the R&D but would sign off on AI hype R&D)
2. The LLM system is doing a task the conventional computer system cannot do.
Number one's problem is simple: It's just inefficient and wasteful. Number two is a red flag: There's very little overlap between the things a conventional computer system cannot do, and the things you can trust an LLM to do reliably.
As you describe this system, selecting which data is relevant for fraud investigation is a very traditional classification task. Using normal machine learning for that is basically industry standard.
So what's the LLM actually doing? Subtract the hard logic of normal software, and the classification of machine learning, and the answer is generally: A complex nuanced reasoning task.
But that's precisely what LLMs are not to be trusted for, because they are incapable of that kind of reasoning.
> Listen. When John Henry battled the steam drill he did win, but it killed him.
You're missing the point I was making with that remark. It's not about firing people or not.
It's that these systems are dangerous to evaluate from a high level. It's very easy to miss externalities that'll tip the entire endeavour into a net-negative. You need the investigation of what exactly the AI systems are doing, on a specific detailed level.
E.g.:
> This is useful if say you have business people or whatever writing effectiveness or whatever testing where they can provide a specification of policy and a well prompted LLM can generate pretty exhaustive cucumber tests (which can be pretty redundant and formulaic when asserting positive and negative cases exhaustively) which can then be revised by hand as needed.
"A specification that has been prompted into sufficient detail" is just a program. You're describing the most inefficient declarative programming stack on the planet.
Granted, the programming stack to actually declare business rules this way isn't very good, but using AI here is just an error-prone transpiler.
It's very easy to "looks good to me" these tests and claim the project a success, yet miss subtle errors in the generated tests. I remain skeptical about how well these tests will hold up in the longer term.
- fnordpiglet 2y agoYou misunderstand. We are one of the top shops for ML based fraud detection. But when someone is accused of fraud they get to appeal it. Then a human is in the loop and the model scores and all inputs are investigated and compared against many other sets of data and policy etc. LLMs facilitate this effort by making the investigators job considerably easier in navigating the enormous amount of complex information. The LLMs role is not to make decisions but to assist in navigating and understanding a lot of high cognitive load information. We have been doing a lot for many years using traditional techniques to make this process easier. But LLMs unlocked a level of dynamism and responsive UX that has blown the lid off our ability to adjudicate appeals. This has significant economic gain for us as offboarding legit customers for fraud causes a lot of losses over a long term. The LLM isn’t used for reasoning at all. The human does all the reasoning. The LLMs task is summarization and semantic analysis of relevance which LLMs are fantastic about especially in a well managed and fine tuned environment with guard rails and context scoping. It’s a true copilot scenario and the LLM takes direction from the human and answers questions only. All decisions are investigator driven. This is the right relationship. The LLM coupled with IR tools does information retrieval and summarization and the human makes decisions and reasons.