Gurpreet Singh
QA Lead, QualiteSoft Solutions
- Published
- Reading time
- 1 min read
Teams often ship AI features with no test plan because outputs vary. Variation is not an excuse; it is the reason to test properties instead of exact strings.
Write acceptance rules you can observe
On a career-coaching platform we made this rule explicit: an insight that cannot point to the member's own answers does not ship. Another: two members answering differently must see different dashboards. Both are testable.
Build a golden dataset
Collect representative inputs — real-shaped, anonymised — with the properties each output must have: grounded, on-topic, no invented facts, correct tone, correct hand-off when unsure.
Score automatically and by review
- Deterministic checks: schema validity, forbidden claims, citation presence.
- LLM-as-judge scoring for quality dimensions, calibrated against human review.
- Human review on a sample every cycle.
Regression on every change
Prompt edits and model upgrades are code changes. Run the evaluation suite in CI and block releases that regress.
Report evidence
Every stage should end with a checklist of what was tested and what was found — the same discipline as any other release.
Share this article
About the author
Gurpreet Singh
QA Lead, QualiteSoft Solutions
Top Rated Plus QA lead on Upwork. Writes about test automation, release quality and testing AI features.
Related services