Key Takeaways
- The CRM industry has spent 18 months chasing generative AI's most obvious use case: summarizing customer conversations.
- Sheeran and Webb's 2016 meta-analysis found intentions explain roughly 28 percent of subsequent behavior.
- The DellaVigna-Malmendier gym study showed members projected three weekly visits but averaged four visits per month, paying a 70 percent premium per session versus pay-as-you-go.
- Current AI workflows treat customer interviews as ground truth without cross-referencing them against product telemetry such as API call volumes or completed setup workflows.
The CRM industry has spent the last eighteen months chasing generative AI's most obvious use case: summarizing customer conversations. Call transcripts, survey responses, support tickets, win-loss interviews — feed them to a model, get a tidy synthesis with key themes and sentiment scores. Vendors from Gong to Salesforce to Zoom have baked this into their platforms. It saves hours. It scales qualitative research. It looks impressive in demos.
It also misses the point.
The fundamental flaw in this approach is not technical. It is epistemological. A synthesis of what customers say can only ever surface what customers say. It cannot surface what they do. And in B2B software, as in gym memberships and meditation apps, the gap between the two is where the revenue actually lives.
Behavioral research has quantified this gap. Sheeran and Webb's 2016 meta-analysis found intentions explain roughly 28 percent of subsequent behavior. The rest is circumstance, friction, competing priorities, and the quiet collapse of aspiration into reality. The famous DellaVigna-Malmendier gym study demonstrated the mechanism: people bought monthly memberships projecting three weekly visits. They averaged four visits a month. They paid a 70 percent premium per session versus pay-as-you-go. They were not lying. They were predicting the behavior of their ideal selves. Life intervened.
The same dynamic plays out in every renewal conversation, every expansion forecast, every NPS survey. A champion tells your CSM they "absolutely plan to roll out the advanced analytics module to the Europe team next quarter." They mean it. The budget gets frozen. The Europe team reorganizes. The module sits unused. Churn follows. The post-mortem cites "lack of adoption." The original interview cited "strong intent."
Current AI workflows treat the interview as ground truth. They refine the summary. They tag the sentiment. They file the insight. They do not cross-reference the interview against the telemetry.
This is the strategic error. The highest-value application of AI in customer intelligence is not synthesis. It is triangulation.
Consider what becomes possible when you point models at the say-do delta rather than the say corpus. A customer rates onboarding 9/10 but their admin console shows zero completed setup workflows after 30 days. A buying committee scores "ease of integration" as critical in the RFP evaluation but their API call volume flatlines after authentication. A champion cites "real-time collaboration" as the renewal driver but feature flags show the collaboration module disabled at the org level. None of these contradictions appear in a call summary. All of them appear in the join between conversational data and usage telemetry.
The technology to do this exists today. Vector databases can embed both transcript chunks and event streams in shared semantic space. LLMs can be prompted to detect semantic inconsistency between natural language claims and structured behavioral patterns. The architecture is not the barrier. The product philosophy is.
Most CX platforms still treat qualitative and quantitative data as separate workstreams. Voice-of-customer teams own surveys and interviews. Product analytics teams own event streams and funnel dashboards. RevOps owns the renewal forecast. The AI that summarizes the VOC data sits in the VOC tool. The AI that detects usage anomalies sits in the product analytics tool. The gap between them is organizational before it is technical.
Anthony Ulwick's outcome-driven innovation framework argued decades ago that customers hire products to make progress in specific circumstances — not to satisfy feature wishlists. The say-do gap is the empirical signature of that principle. When a customer describes a desired capability, they are describing the progress they hope to make. When their usage data shows they never activate that capability, they are revealing the circumstances that blocked them. The insight is not in either dataset alone. It is in the friction between them.
For CRM vendors and RevOps leaders, the implication is direct: stop buying AI that summarizes. Start demanding AI that contradicts.
A practical implementation starts small. Pick one high-stakes workflow — renewal risk scoring, expansion qualification, churn root-cause analysis. Ingest the last 12 months of relevant conversations: QBRs, support escalations, executive check-ins. Ingest the parallel behavioral telemetry: feature adoption depth, integration health, license utilization, support ticket taxonomy. Task the model to identify every instance where expressed sentiment or stated intent diverges materially from observed behavior. Cluster the divergences by theme. Rank them by revenue exposure.
The output is not a summary. It is a catalog of broken promises — some made by customers, some made by your product, some made by your onboarding, some made by your pricing. Each cluster is a hypothesis about where value creation fails. Each hypothesis is testable against the next cohort.
This shifts AI from a productivity tool to a strategic sensor. The productivity gain from summarizing 500 calls is real but bounded. The strategic gain from discovering that 40 percent of "committed renewals" in enterprise accounts show zero admin activity in the preceding 90 days is unbounded. It rewrites the forecast. It rewrites the playbook. It rewrites the product roadmap.
The industry's current AI trajectory — faster summaries, richer themes, better sentiment — is a local maximum. The global maximum sits on the other side of the say-do gap. The vendors and operators who cross it first will not have better summaries. They will have better truth.
Frequently Asked Questions
Why does summarizing customer conversations fail to predict churn accurately?
A synthesis of what customers say can only surface what they say, not what they do, and the gap between stated intent and actual behavior is where revenue lives.
What percentage of behavior is explained by customer intentions according to behavioral research?
Intentions explain roughly 28 percent of subsequent behavior, with the rest driven by circumstance, friction, competing priorities, and the collapse of aspiration into reality.
How should AI be applied to customer intelligence instead of summarization?
The highest-value application is triangulation — pointing models at the say-do delta by cross-referencing interview data against behavioral telemetry like admin console activity and API call volumes.