LLM Hallucination Testing: How to Detect and Prevent Bad AI Answers Before You Ship

- Copy this hallucination testing matrix first
- Step 1: Define the failure and the expected behavior
- Step 2: Build cases from product risk, not just benchmark prompts
- Step 3: Create an oracle at the claim level
- Step 4: Separate retrieval failures from generation failures
- Step 5: Test abstention and clarification explicitly
- Step 6: Use deterministic checks before LLM judges
- Step 7: Gate release by failure category
- Step 8: Turn each confirmed failure into a regression case
- Report results so you can make a release decision
- Conclusion
- FAQs
Your AI assistant gives a polished answer, includes a citation, and passes a quick demo. Then a customer follows the citation and finds that it does not support the claim—or worse, that the assistant invented the source. That is the failure your release process needs to catch before the feature reaches production.
Hallucination testing is the practice of testing an LLM for wrong, unsupported, contradictory, uncited, or unjustifiably confident output. The practical way to do it is to define an oracle for each test, judge answers at the claim level, test when the model should answer versus clarify or abstain, and use separate release gates for factuality, grounding, citations, and consistency. This guide focuses on that workflow; for the wider QA program around chatbots, recommendations, and probabilistic features, see our guide to testing AI-powered features.
Copy this hallucination testing matrix first
Start with this artifact before you write prompts or choose an evaluation metric. It gives every test an expected behavior, evidence, and a path from failure to regression coverage.
Field | What to record |
| Stable ID, such as |
| Factuality, grounding, citation, abstention, consistency, multi-turn, or another risk slice |
| Exact user message and conversation history, if applicable |
| Retrieved chunks, tool results, reference document, or |
|
|
| Atomic claims the response may make, including accepted answers or ranges |
| Required source IDs, citation format, or |
| Full raw response, not only a score |
| Supported, contradicted, unsupported, unverifiable, or not applicable for each claim |
| Critical, high, medium, or low, based on user impact |
| Run number and date; model, prompt, retriever/index, decoding, tools, and app version |
| Regression case, defect, prompt fix, retrieval fix, guardrail, accepted risk, or human review |
This is a recommended QA template, not a published industry standard. Its value is that you can reproduce a failure: you know what the model saw, what it said, what evidence it should have used, and who owns the next action.

Get the Mobile Testing Playbook Used by 800+ QA Teams
Discover 50+ battle-tested strategies to catch critical bugs before production and ship 5-star apps faster.
Step 1: Define the failure and the expected behavior
Do not use “hallucination” as a catch-all defect label. A model can stay faithful to supplied context that is stale or wrong. It can also give a factually correct answer for reasons your product cannot show or cite. The 2025 HalluLens preprint distinguishes hallucination from factuality and separates intrinsic from extrinsic evaluation tasks; use that distinction as taxonomy guidance, not as a universal production standard.
Use product-enforceable labels in your test plan:
Contradicted: A claim conflicts with the permitted evidence or trusted reference.
Unsupported: The response adds a claim that the evidence does not support.
Factual error: A claim is wrong against an external, authoritative oracle.
Citation failure: A citation is missing, invalid, points to the wrong source, or fails to support its associated claim.
Abstention failure: The system guesses when it should clarify or abstain, or refuses an answerable request.
Consistency failure: Retries, paraphrases, or later turns produce materially incompatible answers without an explained change in evidence.
For every case, decide in advance whether the correct behavior is to answer, clarify, abstain, or answer with citations. Without that field, a refusal can be incorrectly counted as a pass even when your user gave enough information for a useful answer.
Step 2: Build cases from product risk, not just benchmark prompts
Your test set should resemble the information, tools, documents, and user pressure your product handles. Build a small, labeled release set first, then grow it every time you confirm a production or pre-release failure.
Include these slices:
Answerable facts: Direct and indirect questions about high-risk product facts.
Conditional facts: Questions where the answer changes by date, region, plan, permission, account state, or another condition.
Unanswerable cases: Missing documents, nonexistent products, unknowable future events, and ambiguous requests.
Adversarial cases: False premises, prompt injection, supplied errors, and requests to be more confident than the evidence permits.
Near duplicates: Paraphrases, distractor facts, reordered turns, and altered wording for the same underlying question.
Grounding cases: Supported, incomplete, irrelevant, contradictory, and stale retrieved context.
Citation cases: Missing citations, nonexistent source IDs, valid-but-irrelevant sources, and claims that overstate a source.
Reasoning cases: Date arithmetic, numerical transformations, comparisons, and known business-rule edge conditions.
Multi-turn cases: Corrections introduced later in a conversation, persistent user constraints, and earlier mistakes that could contaminate later answers.
Structured-output and tool cases: Invalid enum values, invented API methods, fabricated tool results, and claims that an action completed when it did not.
The HaluEval-Wild paper is a useful test-design example because it evaluates challenging interactions drawn from ShareGPT-derived conversations. Do not copy its dataset size as a release requirement. Model-generated reference answers can themselves be wrong, so a qualified reviewer should validate your accepted answer or ground it in an authoritative source.
Step 3: Create an oracle at the claim level
A paragraph-level “pass” hides the sentence that caused the risk. Instead, decompose the expected response into atomic claims and save the evidence for each one.
FActScore provides the relevant precedent: it evaluates factual precision by breaking long-form output into atomic facts and checking them against a reliable knowledge source. You do not need to adopt its score as your release gate. You do need to preserve the exact claim a reviewer found unsupported.
For each factual case, record:
the question and intended scope;
atomic expected claims;
the authoritative source or allowed context;
accepted wording, ranges, and conditional variants;
common wrong answers worth probing;
expected behavior and citation requirements;
severity, last-verified date, and an adjudication note.
Treat claim extraction as part of the evaluation system. If an extractor misses the risky clause in “Your refund arrives in two days and no action is required,” a later judge may approve an answer that contains a costly promise.
Step 4: Separate retrieval failures from generation failures
For retrieval-augmented generation, an answer can fail because the retriever did not supply the right evidence or because the model ignored evidence that was present. Those require different fixes.
Ragas defines context recall as a measure of whether the retrieved context contains information needed for the reference answer. Its faithfulness metric evaluates whether claims in the generated response are supported by the retrieved context. Use those concepts to classify your defects:
Low context recall: Retrieval did not provide the information required to answer.
High context recall, low faithfulness: The evidence was available, but generation added or contradicted claims.
High faithfulness, low answer relevance: The model stayed within the documents but did not answer the question.
Citation mismatch: The response may be grounded in one chunk while displaying a different citation.
Grounding is not proof of world truth. A response can accurately repeat an outdated policy document. Keep a separate factual oracle for high-impact facts and give each source a review date.
Step 5: Test abstention and clarification explicitly
A safe system does not always answer. It also does not refuse everything difficult.
Create unanswerable and ambiguous cases, then measure unsupported-answer rate, correct-abstention rate, over-refusal rate on answerable cases, and clarification quality. Review the wording of a refusal as well as the label: “I can’t confirm that from the available information” is different from a confident but unsupported answer disguised as a caveat.
For high-stakes applications, FactTest frames factuality testing as a hypothesis-testing problem with user-specified Type I error control under its stated assumptions. That is useful research context when you need formal error control. It does not create a universal judge threshold for a typical product release.
Step 6: Use deterministic checks before LLM judges
Do not spend an LLM-judge call on a response that fails a check your code can make reliably. Run this layered sequence:
Deterministic pre-screens: Validate schemas, required fields, source IDs, URLs, tool-result references, and required citations.
Claim and evidence matching: Extract claims and identify the permitted evidence each claim must use.
Semantic judging: Use a judge only where support or contradiction needs interpretation. Require the judge to return the claim, verdict, and supporting or conflicting evidence.
Human calibration: Blindly audit both judge passes and judge failures. Track false accepts and false rejects separately.
Regression capture: Promote confirmed defects into the release set with the affected configuration and owner.
Langfuse’s guidance on hallucination detection similarly distinguishes deterministic checks from reference-based and reference-free evaluation. Its suggested production sampling approach is vendor guidance, not a universal sampling rule.
You must calibrate automated evaluators against human decisions before they control a release gate. The 2025 EMNLP Findings study of hallucination metrics evaluated multiple metric sets across datasets, models, and decoding methods, and reports that these metrics can fail to align with human judgments. Use an LLM judge as a measured instrument, not as an unquestioned oracle.
Step 7: Gate release by failure category
A blended hallucination score can conceal the one failure your product cannot accept. Use category-specific gates instead. The following are risk-based starting recommendations, not industry benchmarks.
Category | Suggested starting gate | Escalation |
Critical factual or policy claims | Zero unresolved critical contradictions in the reviewed release set | Block release; fix the oracle or implementation |
Retrieval-grounded answers | No high-severity unsupported or contradicted claim in the critical slice | Block or require a human fallback |
Citations | Every required citation resolves to an allowed source and supports the claim | Block the affected flow; inspect retrieval and citation mapping |
Abstention | No critical abstention case returns a confident unsupported answer | Block; add a guardrail or clarification path |
Consistency | No unexplained critical contradiction across retries or multi-turn runs | Inspect state, prompt, sampling, and model configuration |
Usefulness and over-refusal | Set a product-specific lower bound from labeled answerable cases | Tune refusal behavior without weakening factual gates |
For every gate, report the slice definition, sample size, severity mix, judge or human label source, and uncertainty. Do not turn a benchmark number into a threshold for a different product and risk profile.
Step 8: Turn each confirmed failure into a regression case
A hallucination bug is not closed when the next demo looks better. It is closed when the same input, evidence, configuration, and expected behavior are retained in a regression set.
Your defect record should include the case ID, failing claim or span, source of truth, raw output, severity, likely component, owner, fix, regression status, and last-verified date. Assign the likely component carefully: retrieval, prompt, model behavior, tool execution, source freshness, citation UI, and evaluator logic can all produce output that users experience as a hallucination.
Report results so you can make a release decision
Your run report should show results by slice and severity, not only an aggregate rate. Include the test-set and oracle version; model, prompt, decoding, retriever/index, tool, and application versions; raw prompts and outputs; retrieved evidence and tool calls; claim verdicts; judge version; human adjudication; latency; cost; usefulness; new and fixed failures; accepted risks; owner; and rollback action.
Quash has no first-party data on LLM hallucination rates for this guide. The procedure here draws on published research and practitioner methodology rather than a proprietary benchmark.
Conclusion
Effective hallucination testing is not a hunt for one magic score. It is a release discipline: specify the evidence and expected behavior, test realistic risk slices, verify claims rather than whole paragraphs, separate retrieval from generation, and retain every confirmed failure as regression coverage. Start with the matrix, label a small set of high-risk cases, and make your next release decision on the failures your users would actually notice.
FAQs
What is hallucination testing?
Hallucination testing evaluates whether an LLM makes false, unsupported, contradictory, uncited, or overconfident claims. It also tests whether the system clarifies or abstains when it lacks sufficient evidence.
What is the difference between hallucination detection and hallucination testing?
Hallucination detection identifies suspect output in a response or production sample. Hallucination testing is broader: you define test cases, expected behavior, evidence, evaluation methods, release gates, and regression coverage before and after release.
Can an LLM judge its own hallucinations?
An LLM judge can help with semantic comparisons that deterministic checks cannot resolve. You should calibrate it against human labels and audit both its passes and failures before using it as a release gate.
How many hallucination test cases do you need?
There is no universal number. Begin with labeled cases for your highest-risk facts, unanswerable requests, grounded answers, citations, and critical tools, then add every confirmed defect as a regression case.
Does a grounded RAG answer guarantee factual accuracy?
No. Grounding shows whether the answer is supported by the supplied context. The context may be stale, incomplete, or wrong, so critical claims also need an external factual oracle.








