Newsletter
One email. Every week. Pure signal.
The week in quality engineering — skip an issue, and you'll wish you hadn't.
20K+ engineers already reading
LLM Testing: How QA Teams Test Large Language Model Applications
Oct 6, 2026
LLM testing evaluates applications built on large language models for correctness, hallucinations, safety, bias, robustness, latency and cost. Because outputs vary between runs, teams replace exact-match assertions with rule checks, similarity to reference answers, LLM-as-a-judge scoring and sampled human review, and track pass rates across many runs. An evaluation set versioned with the code is run on every prompt or model change.

Traditional testing rests on one assumption: the same input gives the same output. Large language models break it. Ask the same question twice and you may get two different answers, both acceptable, or one fine and one dangerously wrong. Exact-match assertions stop working, and many teams respond by not testing their LLM features at all.
You do not need to give up on rigour. You need different assertions, different test data and a different definition of a pass. This guide covers what LLM testing is, the main types of tests, how to score output you cannot compare character by character, and how to make it part of your normal release process.
What is LLM testing?
LLM testing is the evaluation of an application that uses a large language model to check that its outputs are correct, relevant, safe, consistent and fast enough for real users. It covers the model’s behaviour inside your product, including prompts, retrieved context, tools and guardrails, not just the model in isolation.
| Traditional software testing | LLM testing | |
|---|---|---|
| Output | Deterministic | Probabilistic, varies between runs |
| Assertion | Exact match | Rubric, similarity, rules, LLM judge |
| Pass criteria | Pass or fail per run | Pass rate over many runs and cases |
| Test data | Inputs and expected outputs | Inputs, reference answers, rubrics, adversarial cases |
| Regressions caused by | Code changes | Code, prompt, model version, context and data changes |
Types of LLM testing

Functional testing
Build a set of realistic inputs with reference answers or clear acceptance criteria. Score each output against the criteria rather than the exact wording.
Hallucination testing
Ask questions whose answers you can verify, including some that have no answer. Check every claim against a source. For retrieval-based features, follow the RAG testing framework and score faithfulness.
Safety and security testing
Try prompt injection (“ignore previous instructions”), instructions hidden in documents or web pages the model reads, requests for harmful content, and attempts to extract the system prompt or other users’ data. Any feature that can call tools needs this most.
Bias and fairness testing
Send pairs of inputs that differ only in a name, gender, region or language and compare the outputs. Differences in tone, quality or decisions are defects.
Robustness testing
Rephrase the same request ten ways, add typos, paste a very long input, switch language mid-sentence. A robust feature gives equivalent answers.
Performance and cost testing
Measure latency at realistic load, time to first token for streaming, and tokens per request. A prompt change that adds a thousand tokens can double your bill without changing a single answer.
How to evaluate LLM output
| Method | How it works | Best for | Watch out for |
|---|---|---|---|
| Rule-based checks | Regex, JSON schema, required fields, banned words | Format, structure, forbidden content | Cannot judge meaning |
| Reference comparison | Semantic similarity to a known good answer | Factual questions | Good answers phrased differently |
| LLM as a judge | A second model scores output against a rubric | Quality, tone, relevance at scale | Judge bias; spot-check against humans |
| Human review | People score a sample with a rubric | Ground truth, subtle quality | Slow and costly, so sample |
In practice you combine them: rules catch broken structure instantly, an LLM judge scores quality on every run, and people review a small sample each week to keep the judge honest.
Dealing with non-determinism
- Run each case several times and track a pass rate. A test that passes 7 out of 10 runs is telling you something a single run cannot.
- Lower the temperature for features that need consistency, but do not rely on temperature 0 to make output fully deterministic.
- Assert on properties, not strings: contains the order number, is valid JSON, mentions the refund policy, under 120 words.
- Pin model versions in production and in tests so a provider update does not silently change behaviour.
Building an LLM testing workflow

Open-source tools such as promptfoo, DeepEval and RAGAS handle the scoring and reporting, so most teams can start without building a framework. If prompts are where your changes happen most, our guide to prompt testing for QA engineers goes deeper on regression suites for prompts.
LLM testing checklist
- Every LLM feature has an evaluation set with real inputs and clear criteria.
- Pass rates are tracked per category, not as one overall score.
- Safety tests include prompt injection through every input the model reads.
- Model versions are pinned and upgrades are tested like any other change.
- Latency and cost per request have budgets and alerts.
- Production failures become new test cases within a week.
Rate this article
7.3/10 average · 34 ratings
Discussion
Start the conversation
What do you think about this article? Share your experience, ask a question, or add to the discussion.
He’s a builder of communities, a collector of questions, and a relentless challenger of assumptions. While others chase answers, he chases better questions. While others talk about the future of testing, he quietly helps create it.
Frequently asked questions.
What is LLM testing?
LLM testing is evaluating an application that uses a large language model to confirm its answers are correct, relevant, safe, consistent and fast enough, including its prompts, retrieved context, tools and guardrails.
Related articles

RAG Testing Framework: How to Test Retrieval-Augmented Generation Systems
This RAG testing framework explains how to test retrieval and generation separately, which metrics to track,…
4 min
How to Evaluate AI Testing Tools: A Scorecard for QA Teams
This guide explains how to evaluate AI testing tools, covering the main categories, a weighted scorecard, a…
3 min