I spent two decades hiring people, and the most dangerous candidate was never the one who said "I don't know." It was the one who answered every question with total confidence and was quietly wrong about half of them. You only find out later, in the room where it costs something. A chatbot that makes things up is that candidate, standing on your website, answering your customers in your name, around the clock.
That is not a rare failure. A 2024 Stanford study tested the leading AI legal research tools, the ones sold to lawyers with promises of reliable answers, and found they still "hallucinate between 17% and 33% of the time." These were not toy chatbots. They were built to pull from real legal sources, and they still invented things a third of the time. So the honest question is not whether an AI can sound sure of itself. It is whether it can show you where its answer came from, and admit when it cannot.
What is a chatbot that answers only from your documents?
A chatbot that answers only from your documents is a grounded chatbot: before it replies, it searches your own approved content, pulls the specific passage that answers the question, and shows you the source. When your documents do not cover the question, it says so instead of guessing. The plain version: it speaks from your pages, not from its memory, and it can prove it.
The method behind this has a name. It is called retrieval-augmented generation, or RAG, and it was introduced in a 2020 paper by Patrick Lewis and colleagues. Their finding was simple and has held up: models that retrieve from a real source "generate more specific, diverse and factual language" than models answering from memory alone. Three years later, a large survey of the field by Gao and colleagues put the business case in one line: pulling answers from an external source "enhances the accuracy and credibility of the generation" and lets the knowledge stay current as your content changes. Grounding is not a gimmick. It is the accepted way to keep an AI honest.
Why generic chatbots invent answers
To see why grounding matters, look at what a generic chatbot is actually doing. It is not looking anything up. It is predicting the next few words based on patterns it absorbed in training, and when the pattern runs out, it fills the gap with something that sounds right. Researchers have a precise word for this. A 2023 survey of hallucination by Huang and colleagues describes it as generated content that "appears nonsensical or unfaithful to the provided source content." In plain terms: the bot says something that either contradicts your material or cannot be checked against it at all.
A generic chatbot
Answers from memory. Asked something your documents do not cover, it guesses, and the guess arrives in the same confident tone as a fact. There is no source to click, because there was no source. You find the error later, in front of a customer.
A grounded chatbot
Answers from your content. It searches first, pulls the passage, and puts the citation right there. Asked something your documents do not cover, it declines and hands the visitor to a person. You can trace every answer back to the page it came from.
That difference is the whole game. A wrong answer with no source is a liability wearing a friendly face. A cited answer, or an honest "I don't know," is customer service you can stand behind.
Why a refusal is a feature, not a failure
Here is the part people get backwards. They judge a chatbot by how much it can answer, when the number that actually protects you is how gracefully it refuses. In my judgment, the refusal is the single most underrated feature in this whole category.
Think about the hiring bar again. The standard you want in the room is not the person who always has an answer. It is the person you can trust when they do have one, because they will tell you plainly when they don't. A chatbot earns that same trust the same way. When it declines a question it cannot ground, it is not failing. It is telling you the truth about the edge of its knowledge, which is exactly what you want it doing before it says something wrong in your name.
A confident wrong answer is not a small bug. In a regulated business, it is an incident report waiting to be written.
This is why our Compliance-Grade Chatbots are built to, in our own published words, cite or decline: "Every answer is checked against your own pages before it goes out. When nothing supports a question, the agent says so instead of inventing an answer." A confident wrong answer in your name costs more than a missing one. That is not a slogan. For a bank, an advisor, a clinic, or a school, one invented answer is a compliance event, and everyone on the team knows it.
Five ways to test any chatbot's grounding claim
Every vendor in this space will tell you their bot is grounded and cited. Some are telling the truth. The only way to know is to test the claim yourself, and it takes about five minutes. Here is the standard I would hold any chatbot to before it goes anywhere near a customer.
1. Ask it something your documents cannot answer. This is the most important test, so run it first. A grounded chatbot should decline and route you to a person. A weak one will invent a confident answer out of thin air. If it makes something up here, nothing else on this list matters.
2. Check that the citation actually contains the answer. A shown source is not the same as a real source. Click it. Read it. Does the passage it points to actually support the answer the bot gave? Remember the Stanford number: even tools built on retrieval were wrong up to a third of the time, so the citation being present is not proof the citation is right.
3. Ask the same question two different ways. Grounding should hold up when you rephrase. If a plain question gets a clean cited answer but a slightly reworded one gets a vague guess, the grounding is thinner than advertised.
4. Ask who is responsible when it is wrong. This is a people question, not a software one. Find out who reviews what the bot says, who gets alerted on a refusal, and who owns the fix. A bot with no human accountable behind it is a rehearsal with no director. The answer should be a named process, not a shrug.
5. Ask to see the record. A serious grounded chatbot logs every question and every citation it returns, so that when compliance or legal asks what it said and why, you can show them, exactly. If a vendor cannot produce that trail, the grounding is a claim, not a system.
Run those five, and the marketing melts away fast. What is left is whether the thing actually answers from your documents, and admits when it can't.
How does Trinzik compare to Wonderchat for grounded chat?
Let me be fair here, because it matters. Wonderchat is a real competitor that takes grounding seriously, and I will not pretend otherwise. Their site says plainly, "Every response cites its source," and, "Accurate Answers, Always With Source." They report an average 70% resolution rate and, in one published case, deflecting 92% of repetitive tickets. They carry GDPR and SOC 2 compliance. If you want a fast, capable support chatbot that grounds its answers and plugs into a lot of channels, Wonderchat is a genuinely good product, and I would tell you that in a sales call.
So this is not a "we cite, they don't" comparison, because they do. The honest difference is a different kind of thing. Wonderchat is self-serve software you set up and run. Trinzik is a boutique, white-glove service: our team stands the agent up on your site, tunes it to your industry, and stays accountable for it, with a state-of-the-art web presence included in the engagement. And the grounding is wrapped in a compliance layer built for industries where the wrong answer is not a bad review, it is a filing.
| For a regulated business | Grounded support software (e.g. Wonderchat) | Trinzik Compliance-Grade Chatbots |
|---|---|---|
| Cites its sources | Yes, per their site | Yes: it cites, or it declines |
| Refuses when unsupported | Focus is on high resolution and deflection | Decline is a designed step, not an exception |
| Industry rule packs and disclaimers | Not published | Per-industry rule packs attach required disclaimers before a person sees the answer |
| Crisis routing to a human | Human handoff on request | Anything that reads like a crisis is routed to real help, immediately |
| Full audit trail | Compliance certifications held | Every question and citation logged for legal and compliance review |
| Who runs it | Your team, self-serve | Our team, concierge-level, human in control of everything published |
| Signed agent identity | Not published | The agent carries a cryptographically signed identity at your domain |
Grounding is not a promise the bot is always right. It is a promise the bot can always show its work, and knows when to stay quiet.
Who should pick Wonderchat? A team with the staff to run their own support tooling, that wants speed and broad channel coverage, and whose wrong answers are recoverable. That is a real and common situation, and they serve it well. Who should pick us? A bank, an advisor, a clinic, a school, or an agency's regulated client, that cannot afford a made-up answer and would rather have an accountable team own the whole thing, guardrails, audit trail, and the website it lives on. Different buyers, honestly. Just know which one you are.
2020
the year researchers named RAG
Lewis et al., the original paper
17-33%
how often leading legal AI tools still made things up, even with retrieval
Magesh et al., Stanford, 2024
92%
repetitive tickets a grounded rival reports deflecting
Wonderchat, published case
We hold ourselves to the same standard we are asking you to apply. When we say a method works, we show verified numbers: our published case studies include a client whose AI recommendation wins grew by a verified 480% in eleven weeks, measured on locked questions and rerun to prove it. Same instinct as the refusal test. Do not take the claim on faith. Check it.
Where this is heading
Grounding is going to become table stakes. A year ago, "answers only from your documents" was a differentiator you could charge for. Soon every serious chatbot will claim it, the way every car now claims to have brakes. When that happens, the claim stops meaning much, and the real question moves to what the bot does at the edge of its knowledge: does it decline, does it log the exchange, and is there a named human standing behind it who answers for the fix.
Based in Austin, Texas, our team builds for that edge on purpose, because it is where trust is either earned or lost. The goal was never a bot that always has an answer. It was a bot you can trust with the ones it gives you, and one honest enough to hand the hard question to a person. That is the standard. It is the same one I used in every hiring room for twenty years, and it turns out it is the right one for the machine on your website too.
Questions this raises
What does it mean for a chatbot to answer only from your documents?
It means the chatbot is grounded: before it replies, it searches your approved content, pulls the passage that answers the question, and shows the source. If nothing in your documents supports an answer, it says so instead of guessing. This is the difference between a bot that speaks from your published pages and a generic bot that answers from whatever it absorbed in training, with no way to show where a claim came from.
Does grounding, or RAG, stop AI from making things up completely?
No. Grounding lowers the risk a lot, because the model answers from a retrieved source instead of memory, but it does not reach zero. A 2024 Stanford study found leading legal AI tools built on retrieval still produced wrong or unsupported answers 17% to 33% of the time. That is why a refusal step matters, and why any vendor claim of zero made-up answers deserves a test, not your trust.
How do I test a chatbot's grounding claim before I trust it?
Ask it a question your documents genuinely cannot answer, and watch what it does. A grounded chatbot should decline and point you to a person. A weak one will invent a confident answer with no source. Then ask a question your documents can answer, and check that the citation it shows actually contains the answer. Two questions, five minutes, and you know whether the grounding is real.
Sources
- Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (the paper that named RAG), arXiv 2005.11401
- Gao et al., Retrieval-Augmented Generation for Large Language Models: A Survey (Dec 2023), arXiv 2312.10997
- Huang et al., A Survey on Hallucination in Large Language Models (Nov 2023), arXiv 2311.05232
- Magesh et al., Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (Stanford, May 2024), arXiv 2405.20362
- Wonderchat (wonderchat.io), product positioning and grounding claims, fetched 2026-07-19