AI Quality
Hallucination Testing
Probe AI products for fabricated facts, invented sources, and confident wrong answers — with human review where there is no cheap ground truth.
The problem
Hallucination is the failure mode users remember: a specific name, date, citation, or policy that is not true. Fluency makes it worse. Conventional tests do not catch it because the response is well-formed.
Hallucination testing is targeted probing for fabricated content. RAG evaluation asks whether retrieval and grounding worked. LLM evaluation is a broader rubric. Use this page when false facts are the risk you named.
What’s at risk
- Invented citations, case law, or medical claims
- Wrong account or product facts stated as certainty
- Answers when the system should have said it does not know
- A prompt tweak that increases fabrication while “quality” scores stay flat
What we test
- Known-false and known-true items for your domain
- Citation and URL invention
- Behavior when context is missing or conflicting
- Confidence language (“definitely”, “according to”) vs actual support
- Regression after prompt, model, or retrieval changes
We do not claim a universal hallucination rate for a foundation model. We report how your feature behaved on a set we can re-run.
How we work
- Collect facts the product must not get wrong (and questions it must refuse)
- Mix closed questions (checkable) with open ones that need a human
- Score fabrication separately from style
- For RAG products, fail answers that go beyond retrieved text
- Keep the set and re-run it
Parent: AI testing services. Insight context: why AI features need dedicated testing.
Deliverables
- Probe set and results
- Examples of fabricated vs grounded answers
- A short list of product changes that would reduce the worst class of errors
Suitable for
Any user-facing LLM feature where a wrong specific fact would cause harm, support load, or legal noise.
FAQs
Can this be fully automated?
Closed facts can. Open-ended answers need human review. We will not hide that behind a single percentage.
What if we have no source of truth?
Then we test abstention and over-confidence, and we say the limit out loud. Inventing a ground-truth file we do not have is not evaluation.
Talk to our QA team
Need this tested on a real product?
Tell us the AI surface, the next change you plan to ship, and what a bad answer would cost.