Let me start with an accident that made no noise.
A company rolled out an AI assistant. The first week went smoothly: no complaints, and responses were 30% faster. In the third week, a manager reconciling the books found a batch of "approved" refunds — under company rules, refunds of that size must go to a human for review. The AI had approved them directly.
The amount was right. The customer was right. The system never threw a single error. The AI even reported back honestly: task completed.
The problem was this: the rule "in which cases a human review is mandatory" had never been written into any document. It lived only in the tacit understanding between two long-serving employees.
This is not a made-up anecdote. It is a failure type pulled apart line by line in an academic evaluation in August 2026. The researchers call it a silent wrong state — the tool call succeeded, nothing errored anywhere, the AI truthfully said "done", and yet the result violated a business rule. It accounted for 78% of all failures.
The AI failure you imagine and the AI failure that happens are not the same thing
When an owner pictures an AI incident, it is usually "it made something up and the customer called to complain". The real distribution of failures does not look like that.
The researchers ran a large number of tasks in an airline ticketing scenario and broke the failures down: the vast majority happened where nobody was looking — an order cancelled, the passenger count changed to zero, a claim accepted without verification. Every step reported "success".
Understanding that difference is worth real money: an error that raises an alarm you can fix; an error that raises nothing means you are the one paying for it.
Why switching to a more expensive model still does not close the gap
This is the most counter-intuitive part.
The researchers did not go and tune a stronger model. They added a few read-only checks — before the actual write, verify the current state to see whether the operation crosses a boundary. That single move, which calls no model and adds almost no latency, lifted the success rate from 29.6% to 42.0% (statistically significant, and reproduced on another set of 15 random seeds).
What stings more: moved to a newer generation of model, under default settings it still attempts the violating write — the same set of checks is what pulls it back in line.
The paper has a line worth taping to your desk: whether an AI is reliable is a question of whether anyone guards the boundary, not of whether it is smart enough.
The same trap has caught 68% of companies — and it traces back to one place
In July 2026, a survey of 101 companies with more than 100 employees produced another number: over the previous six months, 68% of companies had traced a "confidently stated but wrong" AI answer back to the same missing or inconsistent business definition (in June of the same year the figure was 57%).
There is a detail here that is easy to read backwards. Among companies that had already deployed a governed knowledge layer, the share reporting this problem recurring was 50%; among those that had not, only 21%.
Does that mean "governing it makes things worse"? No. Only with a shared, governed reference point can you trace anything at all: this answer was wrong because that rule expired, that definition is inconsistent. Companies with no reference point have the same errors happening all the time — nobody can locate them, so it all ends up vaguely blamed on "the model isn't good enough".
In one sentence: a clean incident log does not prove a healthy system; it may just prove that nobody is looking.
Three things you can do for your own company today
One: pull 20 real conversations and check by hand whether the answers were right. Not whether it errored — whether the answer itself was correct. Pick the high-risk ones: refunds, compensation, account issues. Twenty is enough; an hour will do it.
Two: search your knowledge base for how many versions of the same policy exist. The usual suspects are the clauses that get edited often: return windows, warranty periods, compensation standards. If the old version is still there, retrieval picks by similarity and the AI has no way to know which one to trust — and the more confidently worded one is often the expired one.
Three: for every critical fact, ask "who owns this?" No owner means nobody notices when it expires. It sounds the most old-fashioned, and it works the best.
While we are here, run the numbers: a customer service position starts at 3,000 yuan a month; AI customer service subscriptions commonly run 99 to 500 yuan a month. The difference was never the money — it is that nobody has checked whether it answers correctly.
What we do about this
We are in the same segment — small and micro businesses — so we walked into these pits ourselves, and the answers are built into the product:
- One fact, one source of truth: anyone can submit, but it only takes effect once the owner approves;
- When a new version overwrites an old one, the old entry is deleted with it — no room left for version landmines;
- Questions we cannot answer are not thrown away: whatever the system fails to handle goes onto a list, so what to add next is obvious;
- Every sentence the AI answers can be traced to a source, and if the source does not line up, it is not allowed to make something up.
We do not promise an accuracy rate, and we do not promise "fully automated" — nobody can deliver those two claims in 2026. What we can give you is a process you can verify yourself.
The industry data points the same way: in a Gartner survey of 321 customer service leaders in October 2025, 91% were under pressure to land AI this year, but only about a quarter had actually made self-service work; at the same time 58% planned to train frontline agents into "knowledge management specialists" — not prompt engineers, people who manage knowledge. And across multiple 2026 case analyses, 74.2% of practitioners still insist that human review stays in the loop.
Nobody dares hand over the judgement entirely. The only difference is whether you find out early, or after something has already gone wrong.
SavantCat. We focus on knowledge bases and AI customer service for small and micro businesses: turning the rules that run on tacit understanding inside your company into knowledge a machine can execute and you can verify.