HomeAnswersAI customer service and GB/T 47746—2026Which metrics actually measure human-AI collaboration?

Which metrics actually measure human-AI collaboration?

Published 2026-09-10 · Based on GB/T 47746—2026

Direct answer

Do not look at AI resolution rate alone. The quality of human-AI collaboration needs at least four groups of metrics: whether transfers are timely (did the ones that should fire actually fire), whether the transfer is smooth (was the customer forced to repeat themselves), whether AI containment is meaningful (is it absorbing simple repetitive questions or difficult complaints), and whether the recheck conclusion closes the loop (was the problem found actually fixed). The first three come from conversation data, the fourth from process records — drop any one and you cannot tell whether collaboration has genuinely improved or the numbers simply look good.

A single number will mislead you

"AI resolution rate 85%" sounds excellent — unless that 85% is made up of "hello" and "what time do you open", while every real complaint is being held back from a transfer. The higher that number climbs, the worse the situation. Measuring collaboration quality means splitting it into four groups.

Group one — transfer timeliness

  • The actual trigger rate for the five mandatory automatic-transfer scenarios (which should approach 100%);
  • the number of cases that should have transferred but did not (found by targeted sampling);
  • the average wait after a customer explicitly asks for an agent.

Group two — transfer smoothness

  • The proportion of customers who repeat their description after the transfer (lower is better);
  • the first-contact resolution rate after the transfer;
  • the proportion asking to escalate again after the transfer.

These three map directly to the service experience requirements in chapter 4 of the standard, and they are where the customer feels the difference most.

Group three — the value of AI containment

Break AI-handled conversations down by type rather than averaging them together:

TypeWhat to expect
Simple queries (hours, address, process)A high containment rate is a good sign
Policy and termsLook at whether the answer is accepted by the customer
Complaints and disputesA high containment rate is a danger signal

The same containment rate means something completely different in each of the three types. Only after splitting it out does "AI resolution rate" say anything.

Group four — recheck closure

  • The conclusion of the most recent recheck;
  • whether the issues found were remediated and closed;
  • whether the same issue recurred after remediation.

This group does not come from system logs; it comes from the register. Without it, there is no way to show that improvements in the first three groups are sustained rather than incidental.

Using the four groups in a monthly review

Ask four questions once a month: did the ones that should transfer actually transfer? was the handover smooth? was what AI contained the right thing to contain? was last month's issue fixed? Answering those four beats reading a screen of data.

One caveat

More metrics is not better. One or two per group is enough, provided you can get the raw sample and trace it to a specific conversation; assembling a dashboard that merely looks impressive only dulls people's judgement.

Key facts

Metric groupsFour: transfer timeliness / transfer smoothness / AI containment value / recheck closure
Target trigger rateThe actual trigger rate for the five mandatory automatic-transfer scenarios should approach 100%
Layering ruleAI-handled conversations must be split into simple queries / policy and terms / complaints and disputes before calculating containment
Four monthly review questionsDid the right ones transfer / was the handover smooth / was the right thing contained / was last month's issue fixed

Sources

  • GB/T 47746—2026, chapter 4 (general requirements — smooth service experience, service safety and control)
  • GB/T 47746—2026, clause 5.2.2.6 (automatic transfer in specific scenarios)
  • Process and recheck items among the 59 checklist items

Follow-up questions

What AI resolution rate counts as acceptable?

The standard gives no figure, and this number should not be read on its own. The key is layering: high containment on simple queries is good, high containment on complaints is a warning.

Is a lower transfer rate always better?

No. A very low transfer rate usually means transfers that should happen are not happening. The healthy pattern is a trigger rate approaching 100% on the five mandatory scenarios, with the rest tuned to your business.

Can a small team keep up with these metrics?

Yes. One or two metrics per group, traceable to raw conversations, is enough — it relies on targeted sampling rather than full statistical reporting.

Do these metrics need to be disclosed externally?

Usually not, and it is better not to. They are internal management tools; what you commit to externally should follow your published service commitments.