CallLens Coach
Inspiration
Teams have more customer conversations than any manager can realistically review.
A sales manager might have dozens of calls happening every day. A support leader may oversee hundreds. Yet understanding what actually happened still often means listening to recordings, reading transcripts, looking at dashboards, and manually deciding who needs help.
Most conversation-intelligence products make this easier to analyze, but they still leave the final job to a human:
Here are your calls. Here are your scores. Now figure out what to do.
We wanted to ask a different question:
What if an agent could review every conversation, understand the evidence, decide whether anything actually requires attention, and stay quiet when it doesn't?
That became CallLens Coach — an autonomous, evidence-backed conversation coach designed to turn call intelligence into action.
The goal isn't to create more notifications. It's to help humans focus their attention where it matters.
What it does
CallLens Coach continuously evaluates conversations and decides what, if anything, should happen next.
For each call, the agent can:
- Analyze the conversation against configurable communication and performance rubrics.
- Identify specific evidence from the transcript rather than relying on a generic summary.
- Examine deterministic metrics such as talk-time balance alongside semantic behaviors.
- Compare the findings with previous conversations and coaching history.
- Determine whether the issue is meaningful enough to require intervention.
- Generate personalized, evidence-backed coaching when improvement is needed.
- Escalate higher-risk situations to a manager when human judgment is appropriate.
- Take no action when the conversation is healthy.
That final behavior is important.
An autonomous agent shouldn't act simply because it can. Sometimes the correct action is nothing.
Instead of giving a manager another dashboard containing 30 calls to inspect, CallLens Coach aims to surface the two conversations that need coaching and the one conversation that genuinely needs human attention.
How we built it
CallLens Coach combines deterministic conversation analytics with agentic reasoning.
At its core, the system uses Strands Agents SDK as the autonomous orchestration and decision layer.
The Strands agent has access to specialized tools for conversation analysis, evidence retrieval, scoring, historical context, coaching, and escalation. Rather than following a rigid workflow for every call, the agent can reason about the available evidence and determine the appropriate next step.
The architecture is roughly:
Call / Transcript
|
v
CallLens Analysis Tools
|
v
Strands Coach Agent
|
+---- Evidence & Metrics
|
+---- Rubric Evaluation
|
+---- Conversation History
|
v
Agent Decision
/ | \
/ | \
No Action Coach Escalate
Rep to Human
We use Amazon Bedrock as the model layer for agent reasoning.
The existing CallLens analysis engine acts as a specialized toolset rather than forcing the language model to calculate everything itself. Deterministic measurements remain deterministic, while the model is used where contextual and semantic judgment is valuable.
This separation is deliberate.
We don't want an LLM guessing a metric that can be calculated exactly, and we don't want rigid code attempting to understand a nuanced human conversation that requires contextual reasoning.
CallLens also uses evidence verification and confidence gating so that coaching recommendations can be traced back to specific conversation evidence.
From analytics to an agent
One of the most important design decisions was distinguishing analysis from agency.
A traditional pipeline might look like:
transcribe -> score -> summarize -> dashboard
CallLens Coach instead treats analysis as information available to an autonomous agent.
The agent must answer a more consequential question:
Given what happened in this conversation, should I do anything?
For example, imagine three calls.
Call A: The representative handles discovery well and the customer is satisfied.
The agent evaluates it and takes no action.
Call B: The representative dominates the conversation, asks few discovery questions, and repeatedly misses opportunities to understand the customer's needs.
The agent gathers the supporting evidence and generates targeted coaching.
Call C: The customer expresses serious dissatisfaction and signals that they may leave.
Instead of treating this as another coaching opportunity, the agent recognizes that human intervention is appropriate and escalates the conversation to a manager with the relevant evidence.
The same agent therefore produces three different outcomes based on context.
Evidence before advice
A major challenge with AI-generated coaching is trust.
A recommendation like:
"Improve your discovery skills."
isn't particularly useful if nobody knows why the system reached that conclusion.
CallLens Coach therefore emphasizes evidence-backed coaching.
Recommendations can be connected to transcript evidence, behavioral observations, deterministic metrics, confidence information, and the rubric used for evaluation.
The objective is to move from:
AI says you did something wrong
to:
Here is the behavior we observed, here is the evidence, here is why it matters, and here is what you can try on your next conversation.
This makes the agent's reasoning more inspectable for both the person receiving coaching and the manager overseeing the process.
Challenges we faced
Knowing when not to act
The hardest part of building an autonomous coach isn't generating advice.
It's deciding when advice is warranted.
An agent that comments on every conversation quickly becomes noise. We therefore designed the workflow around confidence, evidence sufficiency, and intervention thresholds so that "no action" remains a valid outcome.
Combining deterministic and AI reasoning
Conversation analysis contains problems that require very different approaches.
Talk-time ratios and other measurable statistics should be computed deterministically. Understanding whether a representative genuinely discovered a customer's underlying problem requires semantic reasoning.
We learned that the strongest architecture isn't LLM everywhere. It is giving the agent reliable specialized tools and allowing it to reason over their outputs.
Making autonomous decisions explainable
As the system became more agentic, explainability became more important.
If an agent is going to coach someone — or escalate their conversation to a manager — the user needs to understand why.
This led us to make evidence verification and traceability central to the experience rather than treating them as optional metadata.
Designing human escalation
Autonomy doesn't mean eliminating humans.
Some situations involve ambiguity, sensitive customer relationships, or decisions that shouldn't be delegated to software.
The challenge was therefore not simply determining what the agent can do, but establishing where it should stop and bring a human into the loop.
What we learned
Building CallLens Coach changed how we think about AI agents.
The useful part of an agent isn't that it can call many tools or produce long chains of reasoning. The useful part is that it can absorb repetitive cognitive work and return human attention only when that attention has value.
We also learned that specialized deterministic software and agentic AI complement each other extremely well.
The conversation-analysis engine provides measurable, verifiable signals. Strands provides the reasoning layer that can decide what those signals mean in context and what should happen next.
Together they allow us to move beyond another analytics dashboard toward a system that can actually participate in the coaching workflow.
What's next
CallLens Coach currently focuses on evidence-backed conversation coaching, but the same architecture can extend to sales, customer support, recruiting, healthcare communication, training, and other environments where organizations need to understand large numbers of human conversations.
Future work includes richer longitudinal coaching memory, team-level behavioral trends, additional communication rubrics, integrations with workplace tools, and measuring whether coaching recommendations actually improve future conversations.
Our long-term goal is simple:
Every conversation can be reviewed. Every recommendation should have evidence. And humans should only be interrupted when their judgment is genuinely needed.
Built With
- agentic-ai
- ai-agents
- conversation-intelligence
- explainable-ai
- fastapi
- human-in-the-loop
- mcp
- next.js
- python
- sales-coaching
- strands-agents-sdk
- typescript
Log in or sign up for Devpost to join the conversation.