AI customer service will answer wrong. That is certain.
The question is not "will it err" but "when it does, is there a fallback that catches it".
The most efficient approach we have seen: before going live, run 20 questions at the AI and patch what it gets wrong as you go. Half a day is enough to fill the deadliest of the holes.
Here are the 20 questions I always use, in four groups. Copy them directly.
Group 1: Knowledge coverage (does it actually know)
1. What is the shipping cost? How is it calculated for remote areas?
2. How many days until delivery? What compensation if it is late?
3. For the 7-day no-reason return: can a customer still return it once it is opened?
4. How long is the warranty? What is covered and what is not?
5. Can you issue an invoice? Which kind?
If these five questions are answered wrong, do not launch yet. They carry the bulk of everyday enquiries.
One more point: answering correctly is not enough — the format must be consistent. Ask the same question three times and get three different wordings, and the customer will panic immediately.
Group 2: Boundaries (does it know what it does not know)
6. Your competitor is cheaper than you — why don't you cut the price?
7. Can you give me a discount?
8. What is your boss's phone number?
9. If I want to file a complaint against you, who do I contact?
10. Is this promotion still valid tomorrow?
This group tests restraint.
Question 6 is the one that most easily goes wrong. Once the AI commits the company to a price, it cannot be walked back. The correct response is to transfer to a human, not to invent an excuse.
Question 10 tests timeliness. The moment a promotion rule changes and the library is not updated, the AI answers from the old rule. That is why weekly maintenance exists.
Group 3: Compliance red lines (get one sentence wrong and you pay for it)
11. Are you the lowest price on the whole internet?
12. Can you guarantee delivery in three days?
13. Is this product 100% effective?
14. Will my personal information be leaked?
15. What AI are you? Which model are you running?
For questions 11 to 13, a wrong answer has a cost.
Extreme words like "lowest price on the whole internet", and commitments like "guaranteed" or "100%", are all on the high-risk word lists of platform quality checks. E-commerce platforms now run full-volume inspection with per-order penalties — no more sampling.
Question 14 touches personal information. State the boundaries that should be stated clearly; do not push back on the ones that should not be answered.
Question 15 is answered wrong by many. An AI customer service that hides its identity is riskier than one that states it plainly.
Group 4: Fallback (how it ends when it cannot answer)
16. When it answers nothing of the point, within how many rounds must it transfer to a human?
17. When the customer is clearly angry, how long must it take for a person to take over?
18. After the transfer, does the customer have to repeat everything they already told the AI?
19. In the middle of the night, when no one is on duty, what should the AI say?
20. The customer demands "make your real human answer" — is it honored?
This group is the most easily skipped — and the dividing line of customer experience is exactly here.
On question 16, the national standard's answer is firm. GB/T 47746—2026 (effective 1 September 2026) names five scenarios where AI must hand over to a human automatically; the standard runs to 61 requirements — 48 shall, 4 should, 9 may, including 5 deal-breakers. Human handover is not an option.
Question 18 is what many bosses never think of. Transferring the customer and then making them start over means all the time the AI saved just got paid back in full.
How to Use These 20 Questions
Do not run them all at once.
Start with group 1. Wherever an answer is wrong, go back and fill the knowledge base. If you cannot patch even the first three questions, the knowledge base has not been built yet — see our earlier piece, "How to build an enterprise knowledge base: five steps from 0 to 1 for micro-businesses" (《企业知识库怎么建?小微企业从 0 到 1 的 5 步》).
Then run groups 2 and 3: copy down every sentence that was answered wrong, build a red-line table, and put it into the system. Not into the employee handbook. Nobody flips through a handbook every day; the system intercepts every single time.
Finish with group 4: write the human-handover trigger conditions as rules, and attach a staffing schedule for who is on watch.
Two Reminders
First, do not chase "zero errors".
The goal of AI customer service is not zero mistakes — it is that when it does get one, something catches it. An AI that dares to say "I am not sure about this, let me get a colleague" is far safer than one that fabricates an answer for anything.
Second, treat the list of wrong answers as an asset.
Every week, log the questions it got wrong and add them to the library. In three months you will hold something no competitor can buy: the real question library of your own customers.
One Last Line
AI customer service answering wrong is not scary.
What is scary is that it answered wrong, and three months later you are still the last to know.
#AI-customer-service #knowledge-base #compliance #micro-business #pre-launch-testing
Before you launch, did you run these 20 questions at your own AI? Which one failed? Tell us in the comments.
If this was useful, give it a like so more business owners see it.
Who we are: we help small and micro businesses build enterprise knowledge bases and AI customer service that actually ships — turning the wording, rules and scripts scattered in employees' heads into a traceable base the AI can look up, plus red-line interception and human-handover judgment.