Direct answer
Do not look at AI resolution rate alone. The quality of human-AI collaboration needs at least four groups of metrics: whether transfers are timely (did the ones that should fire actually fire), whether the transfer is smooth (was the customer forced to repeat themselves), whether AI containment is meaningful (is it absorbing simple repetitive questions or difficult complaints), and whether the recheck conclusion closes the loop (was the problem found actually fixed). The first three come from conversation data, the fourth from process records — drop any one and you cannot tell whether collaboration has genuinely improved or the numbers simply look good.
A single number will mislead you
"AI resolution rate 85%" sounds excellent — unless that 85% is made up of "hello" and "what time do you open", while every real complaint is being held back from a transfer. The higher that number climbs, the worse the situation. Measuring collaboration quality means splitting it into four groups.
Group one — transfer timeliness
- The actual trigger rate for the five mandatory automatic-transfer scenarios (which should approach 100%);
- the number of cases that should have transferred but did not (found by targeted sampling);
- the average wait after a customer explicitly asks for an agent.
Group two — transfer smoothness
- The proportion of customers who repeat their description after the transfer (lower is better);
- the first-contact resolution rate after the transfer;
- the proportion asking to escalate again after the transfer.
These three map directly to the service experience requirements in chapter 4 of the standard, and they are where the customer feels the difference most.
Group three — the value of AI containment
Break AI-handled conversations down by type rather than averaging them together:
| Type | What to expect |
|---|---|
| Simple queries (hours, address, process) | A high containment rate is a good sign |
| Policy and terms | Look at whether the answer is accepted by the customer |
| Complaints and disputes | A high containment rate is a danger signal |
The same containment rate means something completely different in each of the three types. Only after splitting it out does "AI resolution rate" say anything.
Group four — recheck closure
- The conclusion of the most recent recheck;
- whether the issues found were remediated and closed;
- whether the same issue recurred after remediation.
This group does not come from system logs; it comes from the register. Without it, there is no way to show that improvements in the first three groups are sustained rather than incidental.
Using the four groups in a monthly review
Ask four questions once a month: did the ones that should transfer actually transfer? was the handover smooth? was what AI contained the right thing to contain? was last month's issue fixed? Answering those four beats reading a screen of data.
One caveat
More metrics is not better. One or two per group is enough, provided you can get the raw sample and trace it to a specific conversation; assembling a dashboard that merely looks impressive only dulls people's judgement.
Key facts
| Metric groups | Four: transfer timeliness / transfer smoothness / AI containment value / recheck closure |
| Target trigger rate | The actual trigger rate for the five mandatory automatic-transfer scenarios should approach 100% |
| Layering rule | AI-handled conversations must be split into simple queries / policy and terms / complaints and disputes before calculating containment |
| Four monthly review questions | Did the right ones transfer / was the handover smooth / was the right thing contained / was last month's issue fixed |
Sources
- GB/T 47746—2026, chapter 4 (general requirements — smooth service experience, service safety and control)
- GB/T 47746—2026, clause 5.2.2.6 (automatic transfer in specific scenarios)
- Process and recheck items among the 59 checklist items
Follow-up questions
What AI resolution rate counts as acceptable?
The standard gives no figure, and this number should not be read on its own. The key is layering: high containment on simple queries is good, high containment on complaints is a warning.
Is a lower transfer rate always better?
No. A very low transfer rate usually means transfers that should happen are not happening. The healthy pattern is a trigger rate approaching 100% on the five mandatory scenarios, with the rest tuned to your business.
Can a small team keep up with these metrics?
Yes. One or two metrics per group, traceable to raw conversations, is enough — it relies on targeted sampling rather than full statistical reporting.
Do these metrics need to be disclosed externally?
Usually not, and it is better not to. They are internal management tools; what you commit to externally should follow your published service commitments.