i like these kinds of critiques, we don’t think conversation logs or analysis on top of it is alone enough to replace observability or evals. imo they answer diff questions for diff use-cases.
we're betting that there is a TONNN of product signal buried in conversations that observability misses, esp around like raging, writing in all caps, repeated prompts, frustration loops, and subtle hidden feature demand. thats also why we use per-customer taxonomies instead of a shared one. evals will still be needed.
the root cause is harder, especially in more mature agents. we're using this more as a discovery layer for evals or even just whats happening kind of things, then letting teams go deep into the actual conversations and decide what to take action upon
There's definitely a tonne of signal in those, and it's a critique made from a place of strong support of your basic thesis. There's always been a tonne of signal in traditional customer support requests that goes un-used by most orgs, especially b2c orgs.
In case it's helpful: I always explained it to people I was training like this: All lean product theory comes from listening to the workers actually assembling the parts at Toyota.
Now, most digital products - whether the UI is graphical or linguistic - require a customer to work on an assembly line themselves. An onboarding flow is an assembly line and the user has tasks. Those users complain to agents (whether human or LLM) about their task on the assembly line. The purest implementation of lean philosophy would start with modelling these messages and conversations before it did anything else.
If I were you, I'd build a CRM. Intercom and its ilk charge ridiculous money for functionality that the people using it despise. The existing products in the space optimise for 'serve customers quickly' (increasingly irrelevant with LLMs) and not 'learning from your customers' (increasingly relevant as humans talk to customers less day-to-day). They are horrible to try to integrate into an established product development cycle (I've tried).
I think this makes the proposition easier to comprehend to a customer, the value-add more obvious, and allows you to undercut on pricing, rather than giving people a new bill for something they don't know if they need. The MVP of a CRM is also perhaps easier to build than it might seem initially. "Serve customers faster, cheaper, and learn from them in a highly configurable & meaningfully better way, giving your product iteration an advantage over your competitors". Building a CRM, crucially, allows you oversight of much more of the data - which then enables significantly more meaningful discovery.
This is the unsolved half of the coding agent space: what to actually build, what order to build it in, and why. It's really solvable from your starting point, and is potentially just as important/disruptive as the coding agent has been thus far - especially now that we suddenly have more lines of code than we know what to do with.
I'll shut up now - it's a fascinating space to me, so it's easy to get carried away about! Always happy to talk about stuff like this via email (in my profile) on the off-chance any of the above was useful, though :-)
its fascinating how the toyota example comes up anywhere, its so good!
wdym by modelling the messages and conversations though? i lose you a bit there! for the crm approach, i do think it'll be a problem at some point right now.
the replacing budge is an interesting piece, we did not think of it yet, yes let me dm you on x!
Comments
i like these kinds of critiques, we don’t think conversation logs or analysis on top of it is alone enough to replace observability or evals. imo they answer diff questions for diff use-cases.
we're betting that there is a TONNN of product signal buried in conversations that observability misses, esp around like raging, writing in all caps, repeated prompts, frustration loops, and subtle hidden feature demand. thats also why we use per-customer taxonomies instead of a shared one. evals will still be needed.
the root cause is harder, especially in more mature agents. we're using this more as a discovery layer for evals or even just whats happening kind of things, then letting teams go deep into the actual conversations and decide what to take action upon
There's definitely a tonne of signal in those, and it's a critique made from a place of strong support of your basic thesis. There's always been a tonne of signal in traditional customer support requests that goes un-used by most orgs, especially b2c orgs.
In case it's helpful: I always explained it to people I was training like this: All lean product theory comes from listening to the workers actually assembling the parts at Toyota.
Now, most digital products - whether the UI is graphical or linguistic - require a customer to work on an assembly line themselves. An onboarding flow is an assembly line and the user has tasks. Those users complain to agents (whether human or LLM) about their task on the assembly line. The purest implementation of lean philosophy would start with modelling these messages and conversations before it did anything else.
If I were you, I'd build a CRM. Intercom and its ilk charge ridiculous money for functionality that the people using it despise. The existing products in the space optimise for 'serve customers quickly' (increasingly irrelevant with LLMs) and not 'learning from your customers' (increasingly relevant as humans talk to customers less day-to-day). They are horrible to try to integrate into an established product development cycle (I've tried).
I think this makes the proposition easier to comprehend to a customer, the value-add more obvious, and allows you to undercut on pricing, rather than giving people a new bill for something they don't know if they need. The MVP of a CRM is also perhaps easier to build than it might seem initially. "Serve customers faster, cheaper, and learn from them in a highly configurable & meaningfully better way, giving your product iteration an advantage over your competitors". Building a CRM, crucially, allows you oversight of much more of the data - which then enables significantly more meaningful discovery.
This is the unsolved half of the coding agent space: what to actually build, what order to build it in, and why. It's really solvable from your starting point, and is potentially just as important/disruptive as the coding agent has been thus far - especially now that we suddenly have more lines of code than we know what to do with.
I'll shut up now - it's a fascinating space to me, so it's easy to get carried away about! Always happy to talk about stuff like this via email (in my profile) on the off-chance any of the above was useful, though :-)
its fascinating how the toyota example comes up anywhere, its so good!
wdym by modelling the messages and conversations though? i lose you a bit there! for the crm approach, i do think it'll be a problem at some point right now.
the replacing budge is an interesting piece, we did not think of it yet, yes let me dm you on x!