Nexus Eclipse

AI Quality

Hallucination Testing

Probe AI products for fabricated facts, invented sources, and confident wrong answers — with human review where there is no cheap ground truth.

The problem

Hallucination is the failure mode users remember: a specific name, date, citation, or policy that is not true. Fluency makes it worse. Conventional tests do not catch it because the response is well-formed.

Hallucination testing is targeted probing for fabricated content. RAG evaluation asks whether retrieval and grounding worked. LLM evaluation is a broader rubric. Use this page when false facts are the risk you named.

What’s at risk

  • Invented citations, case law, or medical claims
  • Wrong account or product facts stated as certainty
  • Answers when the system should have said it does not know
  • A prompt tweak that increases fabrication while “quality” scores stay flat

What we test

  • Known-false and known-true items for your domain
  • Citation and URL invention
  • Behavior when context is missing or conflicting
  • Confidence language (“definitely”, “according to”) vs actual support
  • Regression after prompt, model, or retrieval changes

We do not claim a universal hallucination rate for a foundation model. We report how your feature behaved on a set we can re-run.

How we work

  1. Collect facts the product must not get wrong (and questions it must refuse)
  2. Mix closed questions (checkable) with open ones that need a human
  3. Score fabrication separately from style
  4. For RAG products, fail answers that go beyond retrieved text
  5. Keep the set and re-run it

Parent: AI testing services. Insight context: why AI features need dedicated testing.

Deliverables

  • Probe set and results
  • Examples of fabricated vs grounded answers
  • A short list of product changes that would reduce the worst class of errors

Suitable for

Any user-facing LLM feature where a wrong specific fact would cause harm, support load, or legal noise.

FAQs

Can this be fully automated?

Closed facts can. Open-ended answers need human review. We will not hide that behind a single percentage.

What if we have no source of truth?

Then we test abstention and over-confidence, and we say the limit out loud. Inventing a ground-truth file we do not have is not evaluation.

Talk to our QA team

Need this tested on a real product?

Tell us the AI surface, the next change you plan to ship, and what a bad answer would cost.