4 ms·
I built an in-house version of this a couple of years ago for where I was working. My concern would be that by excluding observability, you might end up creatin
by gabriel666smith 3mo ago
I built an in-house version of this a couple of years ago for where I was working. My concern would be that by excluding observability, you might end up creating a really selective dataset, whose conclusions you're then asking companies to take seriously when allocating resources to different possible roadmaps.
My guess would be that agent logs would highlight obvious feature requests and bugs for smaller companies - like customers expecting an AI video editor product to be able to add subtitles to a video by itself.
For larger companies who deal with a higher volume of inbound customer support / agent requests, there will probably be big, noisy, already-known-by-the-team query clusters that make up big portions of the dataset - for example, "billing issue with my subscription". After those big clusters you'll likely have a really long tail of different queries, and - without deep observability - no real way to rank their importance. I also think you'd be unlikely to understand the root cause of the product issue in a complex developed product with lots of users solely from agent logs. Most product teams can't make good product decisions consistently, and they're working with a lot more data.
If coupled with staying out of evals (which, btw, I wouldn't find trust-building, if I were a potential customer of yours), I think that it might be difficult to provide genuine value in this space for larger orgs - without evals it's easily dismissed as just fancy & mostly-contextless sentiment analysis.
But I hope I'm wrong! I do think that (though each org's needs probably have to be catered to in a very boutique way) there are huge gains available by rolling LLMs & language analysis into existing product workflows, and that what you're pitching is absolutely a part of what companies should be doing. We are, of course, meant to actually listen to customers - and LLMs/agents should be making that easier, not harder. Absolute best of luck!
- laalshaitaan 3mo agoi like these kinds of critiques, we don’t think conversation logs or analysis on top of it is alone enough to replace observability or evals. imo they answer diff questions for diff use-cases. we're betting that there is a TONNN of product signal buried in conversations that observability misses, esp around like raging, writing in all caps, repeated prompts, frustration loops, and subtle hidden feature demand. thats also why we use per-customer taxonomies instead of a shared one. evals will still be needed. the root cause is harder, especially in more mature agents. we're using this more as a discovery layer for evals or even just whats happening kind of things, then letting teams go deep into the actual conversations and decide what to take action upon
- gabriel666smith 3mo agoThere's definitely a tonne of signal in those, and it's a critique made from a place of strong support of your basic thesis. There's always been a tonne of signal in traditional customer support requests that goes un-used by most orgs, especially b2c orgs. In case it's helpful: I always explained it to people I was training like this: All lean product theory comes from listening to the workers actually assembling the parts at Toyota. Now, most digital products - whether the UI is graphical or linguistic - require a customer to work on an assembly line themselves. An onboarding flow is an assembly line and the user has tasks. Those users complain to agents (whether human or LLM) about their task on the assembly line. The purest implementation of lean philosophy would start with modelling these messages and conversations before it did anything else. If I were you, I'd build a CRM. Intercom and its ilk charge ridiculous money for functionality that the people using it despise. The existing products in the space optimise for 'serve customers quickly' (increasingly irrelevant with LLMs) and not 'learning from your customers' (increasingly relevant as humans talk to customers less day-to-day). They are horrible to try to integrate into an established product development cycle (I've tried). I think this makes the proposition easier to comprehend to a customer, the value-add more obvious, and allows you to undercut on pricing, rather than giving people a new bill for something they don't know if they need. The MVP of a CRM is also perhaps easier to build than it might seem initially. "Serve customers faster, cheaper, and learn from them in a highly configurable & meaningfully better way, giving your product iteration an advantage over your competitors". Building a CRM, crucially, allows you oversight of much more of the data - which then enables significantly more meaningful discovery. This is the unsolved half of the coding agent space: what to actually build, what order to build it in, and why. It's really solvable from your starting point, and is potentially just as important/disruptive as the coding agent has been thus far - especially now that we suddenly have more lines of code than we know what to do with. I'll shut up now - it's a fascinating space to me, so it's easy to get carried away about! Always happy to talk about stuff like this via email (in my profile) on the off-chance any of the above was useful, though :-)
- laalshaitaan 3mo agoits fascinating how the toyota example comes up anywhere, its so good! wdym by modelling the messages and conversations though? i lose you a bit there! for the crm approach, i do think it'll be a problem at some point right now. the replacing budge is an interesting piece, we did not think of it yet, yes let me dm you on x!