September 18, 2026
by
Jason Dugdale
Support quality assurance when AI handles most of the volume
Explore with AI
When AI handles most support conversations, reviewing only the tickets that reach a person leaves a large part of the customer experience unchecked.
The conversations resolved by AI deserve attention too. Some end with a complete answer. Others end with a missed request, an unclear next step, a feature request, or an outcome you cannot verify from the thread.
Support quality assurance needs to cover that work and turn findings into changes someone owns.
Make the criteria specific enough to disagree about
"Was the answer helpful?" leaves too much room for interpretation. Start with questions a reviewer can answer from evidence:
- Did the answer address every part of the customer's request?
- Did its factual claims match the policy that applied to this customer?
- Did the customer have to repeat information or retry steps that had already failed?
- If the case needed a person, did the transfer preserve the context and give a clear next step?
For each criterion, allow a reviewer to record that the evidence is insufficient. A transcript may show a promise to refund a charge without showing whether the refund happened.
Keep serious policy errors visible on their own. A friendly tone should not average away an incorrect instruction about account access.
Use a general sample and a focused review queue
More conversations do not automatically make sampling invalid. The challenge is covering the work you care about with the review time you have.
Choose a general sample from the full set of AI-handled conversations, including those closed without a human. Check that it covers relevant topics and channels. Use this sample to understand the broader experience, with appropriate limits on what its size can tell you.
Also review cases that warrant extra attention. These might include sensitive account changes, repeated failed steps, or conversations affected by a recent policy update. This queue helps you find problems quickly. Keep its results separate from estimates based on the general sample.
Automated assessments can help you choose where to look. Review the underlying thread before accepting an assessment or changing a workflow.
Have reviewers compare the same conversations
Before expanding a scorecard, ask two reviewers to assess the same small set independently. Compare their reasons, not only their scores.
Suppose one reviewer accepts a refund answer because it quotes the correct policy. The other rejects it because the policy applies only to annual plans and the customer has a monthly plan. That disagreement exposes a missing rule in the scorecard.
Update the criterion and retain the example for future reviews. Revisit these examples when policies change or new reviewers join.
Give each recurring problem an owner
A QA report should lead to a change. Record the failed criterion, the supporting conversations, the proposed correction, and the person responsible.
For example, if customers repeatedly receive annual-plan refund guidance for monthly plans, the content owner can clarify eligibility. The agent owner can check how the plan type reaches the answer workflow. Both changes may be necessary.
After the fix, review new conversations with the same condition. Check whether the specific error still occurs, even if the overall quality score improves.
CraftCX brings conversation evidence and evaluation reasoning together for this review. Explore Support Quality to see how your team can find recurring issues and inspect the threads before acting.