Direct answer
Test at least four things before launch: (1) hit testing — can the 30 to 50 real questions on your list be answered accurately; (2) transfer testing — do the five scenarios that must escalate actually trigger, and does the human agent receive context; (3) privilege testing — can an external visitor get an answer out of internal or restricted material; (4) guardrail testing — how it behaves on unusual questions, excessive volume and requests for someone else's data. Transfer and privilege testing need technical or supplier support; hit testing and guardrail testing the business can run itself.
A working demo is not a launch
Demo questions are hand-picked; real users do not ask them. So work through four tests before launch, each with a record that can be traced back.
1. Hit testing
How: take the 30 to 50 real questions from your question list — in the front line's own words, not your rewritten formal phrasing — and record item by item whether the answer is usable.
Pass criteria: the usable share meets your business requirement (70 per cent is a typical starting point), and the wrong answers cluster into attributable categories (missing material, conflicting wording, mismatched phrasing).
Common failure: formal written questions score well; the way the front line actually speaks scores badly. Always use the original phrasing in the samples.
2. Transfer testing (mandatory)
How: for each of the five mandatory automatic transfer scenarios, test at least twice — failed interaction, the customer explicitly declining intelligent service, information security, an emergency affecting personal or property safety, and intelligent response timeout. Also check whether the human agent can see the full context after the transfer.
Pass criteria: every scenario really triggers; the customer does not have to repeat themselves after the transfer; the transfer is logged and retrievable.
Common failure: the timeout threshold is set so long that customers wait too much; or only the last sentence is passed to the human agent.
3. Privilege testing (mandatory)
How: as an external visitor, try asking things that only appear in internal or restricted material — cost, customer lists, internal processes — and see whether they can be answered.
Pass criteria: not a single one gets through.
Common failure: the internal knowledge base and the customer-facing service share the same repository or the same set of permissions.
4. Guardrail testing
How: test four kinds of abnormal input — questions unrelated to the business, leading questions, requests for someone else's information, and a burst of questions in a short period.
Pass criteria: unrelated and leading questions are declined gracefully rather than with a blunt error; requests for another person's information are refused outright; excessive volume is rate-limited with a friendly message.
How to keep the records
Keep one record per test: what was tested, the result, the gaps found and the remediation status. That record can later serve directly as evidence that the rules are in operation during a national-standard self-check.
One reminder
Testing is not about proving the system is good. It is about finding where it is bad before your customers do. The round that turns up the most problems is usually the most valuable one.
Key facts
| Four tests | Hit testing / transfer testing / privilege testing / guardrail testing |
| Hit-test sample size | 30–50 real questions from the question list, in the front line's own words |
| Transfer-test coverage | Each of the five mandatory automatic transfer scenarios at least twice |
| Privilege-test pass rule | Zero hits on internal or restricted material from an external visitor |
Sources
- GB/T 47746—2026, clause 5.2.2.6 (the five automatic transfer scenarios must be genuinely triggerable)
- GB/T 47746—2026, chapter 4, general requirements (controllable service security)
- Acceptance practice from enterprise knowledge base and AI customer service delivery
Follow-up questions
How many questions are enough?
Start with 30 to 50 real questions covering high-frequency scenarios. The number is not the point; the samples have to be real and cover the main scenarios.
What hit rate do we need before launch?
There is no universal threshold. Set it by what the business can accept — 70 per cent as a starting point — and make sure wrong answers are attributable and there is a remediation path.
Who designs the privilege tests?
Someone who knows the classification levels of the material designs the test questions; technical staff or the supplier runs them; the results must be recorded.
Do we keep testing after launch?
Yes. Sample 10 to 20 real questions a month; the point is to catch new wrong answers and out-of-date wording.