September 17, 2026
by
Jason Dugdale
Chatbot analytics that connect the score to the conversation
Explore with AI
A chatbot dashboard can tell you that handoffs increased. To judge whether that is a problem, you need to understand why customers needed to speak with a human.
Some handoffs are the correct outcome. Others follow repeated answers, missing instructions, or a request the agent could have handled - that same handoff metric can't really differentiate between the two.
Useful chatbot analytics connect a measure to the question your team needs to answer, then to the conversations that help answer it.
Know what each metric counts
Start with the definitions in your platform. Deflection, containment, and resolution can describe different outcomes. Do not compare them as if they were interchangeable.
| Measure | Question it helps answer | Evidence to inspect |
|---|---|---|
| Conversation volume and topic mix | What work is reaching the agent? | The requests grouped under each topic |
| Resolution or containment rate | How often does a conversation meet the platform's outcome definition? | The customer's full request and the recorded outcome |
| Handoff rate | How often does a person need to take over? | The reason, timing, and context passed to the person |
| Customer satisfaction | How do respondents rate the experience? | Their replies, comments, and the response rate |
Write down the denominator as well as the percentage. A result based on all incoming conversations answers a different question from one based only on conversations the AI answered. Satisfaction feedback also represents the customers who responded, which may differ from the full group.
Check the mix before explaining a change
Here is a hypothetical example. Last week, most conversations concerned delivery dates. This week, a billing change generated more account-specific questions that required human review. Handoffs increased.
That increase does not, by itself, show that the agent got worse. The work changed.
Compare billing conversations with billing conversations. Keep the channel and outcome definition consistent. Check how many conversations each group contains before treating a percentage change as a trend.
Then read examples from the group that moved. Did the agent transfer the right cases? Did it collect the details the next person needed? Did customers have to explain the problem again?
Separate finding problems from estimating their frequency
If you want to find a handoff problem, review conversations with repeated questions or failed transfers. That focused selection helps you locate defects.
If you want to estimate how common those defects are, you need a sample that represents the group you are measuring. A queue selected because it looks problematic will overstate the problem rate if you treat it as representative.
Keep the two purposes clear in reports. "We found eight incomplete transfers in the review queue" is a useful finding. It does not mean eight out of every hundred customers experienced one.
Report a decision, with its evidence
A useful analytics update might read:
In the billing handoffs we reviewed, customers repeatedly had to provide their invoice number after transfer. We will add that field to the handoff context and check new billing transfers after the change.
Attach the supporting threads, the size of the review, and anything you could not verify. That gives the next person enough context to assess the finding.
CraftCX keeps quality findings connected to the conversations behind them. Use Support Quality to investigate those findings, or read our guide to QA when AI handles most of the volume to build a regular review process.