Sistava

How to Make AI Support Answers Accurate and Grounded

How-to — — by Mahmoud Zalt

Make AI support answers accurate: ground every reply in your own docs, orders and policies, and refuse to answer when your content does not cover it.

The fear here is specific and reasonable. Not that the AI sounds bad, but that it sounds excellent while telling a customer your return window is thirty days when it is fourteen. That answer reads perfectly, gets sent instantly, and you find out about it three weeks later during an argument you are going to lose.

That fear is the right instinct, and it points at the right fix. A support AI that guesses is worse than no support AI at all. A slow human reply costs a bit of patience. A confident wrong one about a refund, a warranty or an ingredient costs the customer and sometimes more than that.

Here is how the fix works. With **Sistava** a Customer Support AI Employee answers from the material you give it: your help documents, your written policies, replies you were happy with, and the live record of the actual order or account. When the question falls outside that material it does not stretch. It says so and hands the ticket to a person.

At a Glance

2 rules
trace every answer to your own content, refuse when you cannot
Your docs
the only acceptable source for a policy answer, never general knowledge
30
test questions with known answers is enough to measure real accuracy
0
invented numbers, dates or policies, at any hour, for any reason

What does grounding actually mean?

Grounding means the answer is built from a specific piece of your material, not from general knowledge about how businesses usually work. Your returns page, your shipping table, your written escalation policy, or the actual database row for that customer's order. If you cannot point at where an answer came from, it was not grounded.

This distinction sounds academic until you see it fail. General knowledge about refunds is genuinely useless to your customer. They do not want to know how refunds usually work. They want to know how yours work, and only your policy page knows that. Everything else is a well-written guess.

Why does refusing to answer make it more accurate?

Because the alternative to refusing is inventing. When a question has no answer in your material, there are exactly two possible outputs: an honest escalation or a plausible fabrication. There is no third option where it quietly gets it right. Allowing refusal is what removes the fabrication branch entirely.

This is the setting people are most tempted to loosen, because refusals feel like failures and they show up in your queue. They are not failures. Each one is a precise map of a gap in your written content, delivered free, before a customer got hurt by it. Treat the refusal count as a to-do list rather than a score.

There is a version of this that goes wrong in the other direction. If your content is so thin that half of everything escalates, the AI Employee is not the problem and tightening it further will not help. That is a signal to spend two days writing, then come back. Accuracy work is mostly writing work in disguise.

A supplement brand that audited 30 answers before going live

Nadia sells supplements online. Her support queue is unusual because a wrong answer is not just embarrassing. Questions about allergens, ingredient interactions and what is safe during pregnancy have real consequences, so she refused to send a single automated reply until she had measured accuracy herself.

She built a test set of thirty questions she already knew the correct answers to, taken from real tickets. Ten about shipping and returns, ten about ingredients and allergens, and ten deliberately chosen because her documents did not cover them. That last ten was the important group. She was testing whether it would refuse, not whether it could answer.

Question groupResultWhat it revealed
Shipping and returns (10)9 correct, 1 wrongHer returns window was written as 14 days on one page and 30 on another
Ingredients and allergens (10)10 correctHer ingredient sheets were already precise, because regulation forced them to be
Deliberately uncovered (10)8 refused, 2 answeredThe two answered were both pregnancy questions, which she moved to always escalate
Re-test after fixes (30)28 correct, 2 refused, 0 wrongZero wrong is the only number she cared about, and it took two days of editing

The single wrong shipping answer was the most valuable result of the whole exercise. It was not the AI being unreliable. It was Nadia discovering that her own site had said two different things about returns for eleven months, and that her human replies had been inconsistent for exactly the same reason.

The two pregnancy answers changed her design rather than her content. Some categories should never be answered automatically no matter how good the documentation gets, because the cost of being wrong is not proportional to the cost of waiting. She added a hard rule: anything mentioning pregnancy, breastfeeding, medication or a medical condition goes to a person, full stop.

How do you test accuracy yourself?

Build a list of thirty questions you already know the answer to, and include ten your documents deliberately do not cover. Run them, then score three buckets: correct, refused, and wrong. Only the third bucket matters. Refusals are cheap and fixable. Wrong answers are the thing you are actually testing for.

Repeat the audit whenever you change a policy, and read every real escalation for the first fortnight. Between those two habits you will catch nearly everything. Neither takes long, and both replace worrying with a number you can look at.

Getting from plausible answers to accurate ones

  1. 1. Find every place a policy is written twice — Your site, your terms page, your email templates, your FAQ. Contradictions produce confident wrong answers, so reconcile them before anything else. This alone fixes most accuracy problems.
  2. 2. Write the exceptions, not just the rule — The rule is easy and already written somewhere. The exceptions live in your head, and they are exactly where wrong answers come from. Every document should say when the rule does not apply.
  3. 3. Connect live records for anything time-sensitive — Order status, delivery dates, invoice amounts, account tier. These must be looked up, never inferred, because a typical answer is a wrong answer for the person asking.
  4. 4. Require a source for every answer — The AI Employee should be able to point at the document or record behind each reply. If it cannot, that reply should not go out. This is the whole grounding rule in one sentence.
  5. 5. Make refusal the default, and say so plainly — No grounding, no answer. Write the refusal message yourself so it sounds like your business rather than an error, and make sure it hands the customer to a person with the context attached.
  6. 6. Build a 30 question audit and run it before going live — Twenty you know the answers to, ten your docs do not cover. Score correct, refused and wrong. Do not send a real reply until wrong is zero.
  7. 7. Re-audit after every policy change — Changing a shipping window or a refund term means old wording is now wrong somewhere. Ten minutes of re-testing catches it. This is the maintenance loop and it is the entire ongoing cost.

Comparison

DimensionTraditionalWith Sista
Source of a policy answerGeneral knowledge about how businesses usually handle itYour written policy page, including its exceptions
Source of a status answerA typical delivery estimate stated as factThe live record for that specific order or account
Question nothing coversA fluent paragraph that reads exactly like a correct answerAn honest refusal and a clean handoff to a person
When you change a policyOld wording keeps circulating and nobody noticesUpdate one document, re-run the audit, done in ten minutes
How you find mistakesA customer quotes it back to you during a disputeA test set you run yourself, before anything is sent

Worth saying clearly: this work improves your human support too. Every contradiction you reconcile and every exception you write down was already causing inconsistent answers from people. The audit just made it visible.

The honest limits of accuracy work

You cannot ground your way to an answer that does not exist. If your business has never decided what happens when a customer wants to swap a gift for a different size, no document will save you, and the AI Employee will correctly refuse. That refusal is not a bug, it is a decision you have been avoiding.

Some categories should stay off limits permanently even with perfect documentation. Medical, legal and safety questions, anything where being wrong causes real harm, and anything involving an exception to your own policy. Accuracy work reduces risk. It does not eliminate the need for judgment, and no setting will.

FAQ

How do I stop AI support from making up answers?

Require every reply to come from your own material and make refusal the default when nothing covers the question. Those two rules together remove the invention branch entirely. Then test it with a set of questions you deliberately have no documentation for, and confirm it refuses rather than improvises.

How do I actually test whether the answers are accurate?

Build thirty questions you already know the correct answers to, taken from real tickets, and include ten your documents do not cover. Score three buckets: correct, refused, wrong. Only wrong matters. Do not send a live reply until that number is zero, and re-run the set after any policy change.

What if my own documents contradict each other?

Fix that first, because contradictions are more dangerous than gaps. A missing answer produces a clean escalation, while two conflicting policy pages produce a confident answer that is wrong roughly half the time. Search for every place a number appears twice, pick the true one, and delete the other.

Should the answer show the customer where it came from?

Linking the relevant help page is usually good practice, because it lets the customer verify and read further. More important is that you can see the source internally when reviewing replies. If you cannot tell which document produced an answer, you cannot tell whether it was grounded or invented.

How often do I need to update the content behind it?

Whenever a policy changes, and whenever the same escalation appears twice. Those two triggers cover almost everything. Most businesses settle into a rhythm of a few small edits a month, usually prompted by an escalation pointing straight at the missing paragraph.

Is a better AI model the way to get better accuracy?

Not usually. The failures that hurt come from missing or contradictory content, not from the model misunderstanding a question. A stronger model asked about a policy you never wrote down will produce a more convincing wrong answer, not a right one. Spend the effort on your documents.

Grounding is half the trust equation. The other half is what happens at the boundary: when the AI Employee decides this one is not mine and hands it over cleanly.

Start with the contradiction hunt. Search your site and your saved replies for every number that appears in more than one place, and reconcile them. It takes an afternoon, it makes your human replies better too, and it is the single highest-value hour in this entire subject.