5 ms·
Show HN: Paramount – Human Evals of AI Customer Support
Hey HN, Hakim here from Fini (YC S22), a startup focused on providing automated customer support bots for enterprises that have a high volume of support requests.
Today, one of the largest use cases of LLMs is for the purpose of automating support. As the space has evolved over the past year, there has subsequently been a need for evaluations of LLM outputs - and a sea of LLM Evals packages have been released. "LLM evals" refer to the evaluation of large language models, assessing how well these AI systems understand and generate human-like text. These packages have recently relied on "automatic evals," where algorithms (usually another LLM) automatically test and score AI responses without human intervention.
In our day to day work, we have found that Automatic Evals are not enough to get the required 95% accuracy for our Enterprise customers. Automatic Evals are efficient, but still often miss nuances that only human expertise can catch. Automatic Evals can never replace the feedback of a trained human who is deeply knowledgeable on an organization's latest product releases, knowledgebase, policies and support issues. The key to solve this is to stop ignoring the business side of the problem, and start involving knowledgeable experts in the evaluation process.
That is why we are releasing Paramount - an Open Source package which incorporates human feedback directly into the evaluation process. By simplifying the step of gathering feedback, ML Engineers can pinpoint and fix accuracy issues (prompts, knowledgebase issues) much faster. Paramount provides a framework for recording LLM function outputs (ground truth data) and facilitates human agent evaluations through a simple UI, reducing the time to identify and correct errors.
Developers can integrate Paramount with a Python decorator that logs LLM interactions into a database, followed by a straightforward UI for expert review. This process aids the debugging and validation phase of launching accurate support bots. We'd love to hear what you think!
- Marietta040 2y ago[dead]
- trothamel 2y ago[flagged]
- deleted 2y ago[deleted]
- simonw 2y agoYour license is very weird: https://github.com/ask-fini/paramount/blob/main/LICENSE https://github.com/ask-fini/paramount/blob/main/LICENSE This file is part of paramount project, licensed under the GNU General Public License (GPL) for companies with fewer than 100 employees or fewer than 1000 invocations/month. For larger companies or higher volume, a commercial license is required. For more information, contact hello@usefini.com. I'm fine with companies not using open source licenses, but this is a very odd way to do it. Licensing something under the GPL doesn't work like this. You should look at one of the existing non-open-source licenses like the Business Source License or https://fsl.software/ https://fsl.software/ rather than modifying the GPL by adding an extra paragraph at the top: https://github.com/ask-fini/paramount/commit/8345edd8f776572d98b3782fbbba447da4d19c6e https://github.com/ask-fini/paramount/commit/8345edd8f776572... Also, with a license like this it's not accurate to say "Paramount - an Open Source package..." - that's a misuse of the term.
- deleted 2y ago[deleted]
- henriquez 2y ago> Also, with a license like this it's not accurate to say "Paramount - an Open Source package..." - that's a misuse of the term. It’s not free and open source software (FOSS) that’s for sure. The GPL can’t be used like this, were it so simple plenty of others like Redis or Elasticsearch would have done so. This license is worse than no license.
- nahikoa 2y agoMaybe work with a specialist attorney on the license if you haven't already?
- nucleardog 2y ago> Licensing something under the GPL doesn't work like this. Sure it does! The GPL covers this exact scenario. Section 7 enumerates the additional restrictions you may include alongside the license which will apply to any further distribution. Those are mostly around indemnification, trademarks, etc. It explicitly says all other non-permissive additions are considered further restrictions and if the program says it is covered by GPL you may remove those terms. (There’s also section 10, but we don’t need it.) Since the README says this is “under GPL license for individuals”, and the GPL license says I can remove those terms… without even getting really far into the mud here, I can download a copy of the software, strip those restrictions, and repost it under the GPL sans restrictions for anyone to use. That all said… it will probably have most of the intended effect. Individuals won’t care about the license much (may limit outside contributors), but no company is going to touch this with a ten foot pole with a hacked up GPL on it, >100 employees or otherwise.
- sltkr 2y ago“Customer support by AI” sounds even worse than the current practice of having customer support provided by utterly clueless and unhelpful “support” from barely-English-speaking-people from the lowest-wage-country-we-could find (stereotypically India). How did you come to hate people so deeply that you made it your life goal to make an already abysmal experience even worse?
- dragonsky67 2y agoWhen the obvious desire of 90% of customer support is to make the customer go away being able to use a box rather than a person to get rid of them seems a perfect solution. Support is a cost centre, the people calling support are not the ones who will be making purchasing decisions so why provide anything close to decent support. Vendor lock in, subscription services and other ways to reduce the chance that the customer will go elsewhere all contribute to the downward spiral in support. Truth is, if they can manage to provide proper feedback to the AI for when the support that is provided is actually useful or successful this may actually learn to be better than employing someone to read off a support flowchart that hasn't been updated in 20 years.
- laborcontract 2y agoI was on a small team of 10 that experienced very strong growth in our product. One of the consequences of success is that we had to take hundreds of calls per day. We eventually had to hire dedicated customer support people but after hours were an issue. There are large outsourced customer service companies but those cost $1 a minute and those people have to deal with CS calls from many different companies, so it took them forever to find scripts relevant to the customer’s request. Most calls failed - that outsourced CS was a glorified voicemail box. Our success rate with outsourced CS was very low, partly because we didn’t have the resources to fly over to Austin to train a workforce that didn’t show to be particularly promising. AI voice bots would have been able to helpfully answer and deflect 90% of our calls, and do so in a much faster and humane way. I understand the sentiment behind your comment, but I cannot agree with your assessment of the outcome.
- stackskipton 2y ago
- chaostheory 2y agoI get that naming is hard but what’s the rational behind the naming because there is a large and well funded brand with the same name?
- hn_version_0023 2y agoJesus H. Fucking Christ. I am infuriated by the very idea. Whenever I have to interact with any of this sort of nonsense I always aggressively hit zero and demand to speak to a HUMAN not a machine while cursing wildly. You are actively making the world a worse place with this “product”. Don’t think about if you can. Think about if you should.
- qeternity 2y agoHow does Fini annual pricing at $0.076 per question work??
- hakimk 2y agoWe typically offer a package price for companies with >1k questions per month with white-glove onboarding and numerous custom optimizations to ensure higher performance. For the fully self-serve pay as you go option, we are able to offer a lower pricing.
- qeternity 2y agoI’m asking how do you blend annual billing and usage based together?
- cachehit 2y agoI am reminded of DHH's recent essay: "It’s easier to forgive a human than a robot" https://world.hey.com/dhh/it-s-easier-to-forgive-a-human-than-a-robot-d4a97b3a https://world.hey.com/dhh/it-s-easier-to-forgive-a-human-tha...
- almazzz 2y ago[dead]
- inetknght 2y agoMany companies use this style of "support" in place of employing real people and don't provide a method to escalate to a real human. What guard do I have, as a consumer, against trash like that? What makes yours unique? When I call for support I expect to receive support. I have never ever had an AI or robot support assistant that actually provides the help that I need when I call. How do you address the possibility of LLM generating absolutely false information? Do you actively tell the user where they can contact a real human to provide feedback about the terrible customer support response that your product will undoubtedly provide? Or is this just another dumpster fire of "it sounds like you're having trouble with our product. here's our FAQ..."? Does your product integrate with internal company APIs? How do you deal with the risk of customers abusing that?
- throwaway-blaze 2y agoHow much do you plan to pay for the product? If you're paying a few dollars per month for a consumer product, be aware that having a human answer a single inbound call or email from you can often wipe out the entire year's revenue (not just profit) the company is getting from you. Either we all need to be willing to pay more in exchange for "premier" support being available or we need to be ok with companies trying to cut down the support load.
- inetknght 2y ago> Either we all need to be willing to pay more in exchange for "premier" support being available or we need to be ok with companies trying to cut down the support load. Third option: if you can't afford to provide human support then you don't deserve to be a business. Too expensive to provide support? Then raise your prices. If your customers leave then your product isn't viable.
- xp84 2y agoOk but you’re still making the same point. You can’t have something cheap with a white glove CS experience for every caller. Either you have to have a screen to filter out the complete idiots who will waste their time all day because their mouse isn’t plugged in, like AI, or some minimum wage tier 1, etc. Or it has to be cost more to pay for that cs.
- smarri 2y agoHey team, nice work. Can you help me understand this better. How does the process work in terms of the human agent evaluations? Is it real time so that the right (maybe a better word is best) answers go to users as they are needed, or is it done asynchronously/batch style so that the humans are training models to be better? Once the best answers are selected, is it fed back into an LLM / AI agent model? Thanks
- smarri 2y agoFollowing up on my own question. I re-read the github, so I can clarify my question better. So the AI agent responses are saved to a database where the human sme can classify responses as good/bad right? Do you intend for the result of this analysis to retrain the AI agent, or is it purely to get a baseline on the as-is AI agent quality?
- hakimk 2y agoIndeed the evaluations are saved to DB. Right now it's possible to use this for regression testing with the help of the Optimize tab. In the optimize tab you can experiment with changing an input parameter (such as prompt or temperature, etc), then rerun recordings and see whether the the LLM response matches all previously accepted recordings or not - to see a similarity score which tells you if your change introduced regressions or not. In the future we are planning to enable a retraining pipeline - most likely we will do this in our core offering at usefini.com
- smarri 2y agoThank you for explaining
- Wazflame 2y agoHi smarri, I hope you're well. Sorry this might not be the best place to ask, but a few months ago you posted on my thread about a a part time, remote, Masters Degree in Software Development in the UK of which businesses recruit directly from - my apologies for the late reply, but could I get in touch with you to have more details?
- zug_zug 2y agoSlightly off-topic, but slightly not. I feel like customer support is one of the worst experiences in all of tech or maybe just modern capitalism. Obviously slashing budgets is a huge part of it, but I think the problem is exacerbated by lying with metrics. One of the great challenges of trusting ANYBODY (even internal teams) with support, is that they have a huge incentive to lie to their management chain about the numbers in every way they can. And their manager or even multiple levels up may turn a blind-eye to these lies if it makes their life easier. I don't know anything about this product, but I would never trust a company with anything of importance if I don't have a complete audit log of their requests and how they handle them which they send to us for review each month (even a 1%-.1% sample of calls will give a very clear picture into how the process is going).