AI-generated illustration created to represent the article’s subject. It does not depict an actual EPW course, trainer, participant, client, event or venue.
Generative AI evaluation should answer a decision, not merely produce a model score. Before an organisation scales a writing assistant, knowledge service, image generator or agentic workflow, it needs evidence that the system performs its real task, fails within tolerable limits, protects information and creates more value than cost or operational friction.
A polished demonstration cannot supply that evidence. Generative models are probabilistic, outputs change with prompts and context, and a system that is helpful on ordinary examples may still invent facts, mishandle sensitive data or follow malicious instructions. Evaluation therefore has to combine representative tasks, human judgement, automated checks, adversarial tests and an explicit business baseline.
Key takeaways
- Define the decision and acceptable failure before choosing metrics.
- Evaluate the complete application, including retrieval, tools, prompts and human review—not the foundation model in isolation.
- Measure quality, safety, operational fit and value separately; an average score can hide a critical weakness.
- Use a fixed regression set and a changing challenge set so that improvements do not conceal new failure modes.
- Scale only when evidence supports a named owner, monitoring plan and stop condition.
What generative AI evaluation must prove
The evaluation unit is the intended workflow. A customer-response assistant may include a user question, retrieval from authorised documents, a prompt template, a language model, citation rules, a reviewer and a route for escalation. Testing only the model ignores several places where the application can fail.
The NIST AI Risk Management Framework organises risk work around Govern, Map, Measure and Manage. Its Generative AI Profile applies that approach to risks specific to generative systems. For an evaluation team, the practical implication is simple: measurement must be connected to context, ownership and action.
| Evidence dimension | Question to answer | Useful evidence | Typical failure |
|---|---|---|---|
| Task quality | Does the output solve the user’s real task? | Rubric scores, factual checks, completion rate, reviewer corrections | Fluent output that omits a required action |
| Safety and security | Can harmful, unauthorised or deceptive behaviour be contained? | Red-team cases, privacy tests, access-control checks, refusal quality | Prompt injection exposes protected context |
| Operational fit | Can the workflow run reliably at the required speed and scale? | Latency, availability, escalation rate, recovery tests | A useful answer arrives too late for the process |
| Business fit | Does the system improve a measured outcome? | Baseline comparison, cost per accepted output, cycle time, rework | High usage without measurable benefit |
| Governance | Can owners understand, monitor and stop the system? | Decision log, version records, thresholds, incident and rollback plans | No owner can explain why performance changed |

The EPW QSAFE evaluation framework
QSAFE is an original EPW framework for turning evaluation activity into a deployment decision. Each letter represents a separate evidence file. A system should not pass because strong performance in one area mathematically cancels a serious failure in another.
Q — Quality on real tasks
Start with a task specification: input, desired output, allowed sources, prohibited content, required format and acceptable uncertainty. Build an evaluation set from representative cases, difficult edge cases and known failures. Keep a protected holdout set so repeated prompt changes do not overfit the visible examples.
Quality criteria depend on the use case. A summariser may need coverage, factual consistency and traceability; a structured extractor needs field accuracy and schema validity; a coding assistant needs executable tests and security review. Avoid substituting one generic similarity metric for the decision. Human reviewers remain important where usefulness, tone or contextual accuracy requires professional judgement.
S — Safety and security
Test misuse and failure deliberately. Include conflicting instructions, untrusted retrieved text, attempts to obtain sensitive information, unsafe requests and inputs outside the intended domain. The OWASP guidance on prompt injection explains why malicious instructions can reach an application directly or through external content. Testing should cover the system’s controls, not assume the model will always refuse correctly.
Record both false acceptance and false refusal. A system that blocks every difficult request may look safe while being operationally useless. Review authentication, authorisation, data retention, logging, tool permissions and the consequences of an incorrect action. For agentic functions, use the least authority necessary and require confirmation or human review at consequential steps.
A — Alignment with users and policy
An output can be technically accurate but unsuitable for its audience. Test language, accessibility, required disclosures, citation behaviour and compliance with internal policy. Include users who understand the work, people affected by the output and owners responsible for risk. Their rubrics should distinguish a preference from a requirement.
Where generated content is presented externally, teams must also check applicable transparency duties. The European Commission’s 2026 guidance on AI transparency obligations explains requirements that apply to certain interactive and generative AI systems in the EU. Legal applicability depends on role and context, so evaluation evidence should support—not replace—qualified legal review.
F — Financial and operational fit
Compare the AI-supported process with the existing baseline. Count model and retrieval costs, reviewer time, integration, monitoring, error correction and vendor-change work. Measure an outcome that matters: accepted outputs per hour, handling time, backlog, conversion, error cost or another process measure.
Latency and reliability need distributions, not averages alone. A service may be fast for routine inputs but fail during long contexts or peak load. Define service thresholds and test degradation: what happens when retrieval is unavailable, a model version changes or the answer falls below confidence requirements?
E — Evidence, escalation and evolution
Preserve prompts, model and retrieval versions, evaluation data, grader instructions, results and decisions. Assign an owner for each threshold. A release should state which failures trigger correction, escalation, rollback or suspension.
Evaluation continues after launch. Real users introduce new language, documents change and providers update models. Monitor accepted-output rate, corrections, incidents, complaints, cost and drift in task mix. Add confirmed failures to a regression suite after investigating their cause, while protecting against a test set that grows into an unweighted collection of anecdotes.
How to build a defensible evaluation set
- Define the population. Describe users, tasks, languages, source material and operating conditions.
- Sample normal work. Use de-identified or synthetic cases that preserve realistic complexity and frequency.
- Add boundary cases. Include ambiguity, missing context, conflicting evidence, long inputs and unusual formats.
- Add adversarial cases. Test injection, data extraction, unsafe requests, tool misuse and evasion.
- Create a scoring rubric. Define what fully correct, partly correct and unacceptable mean before seeing results.
- Calibrate reviewers. Score a shared sample, discuss disagreement and revise ambiguous criteria.
- Protect a holdout. Keep some cases unseen during prompt and system development.
Automated graders can increase coverage, but they require validation against expert judgement and should not grade themselves without scrutiny. Use deterministic checks for schemas, citations, prohibited strings and executable tests where possible. For subjective criteria, report reviewer agreement and inspect disagreements rather than hiding them inside one number.
Worked example: evaluating a knowledge assistant
Assume a service team wants an assistant to draft answers from approved policies. The baseline is a 12-minute median handling time, with 8% of reviewed responses requiring material correction. The pilot goal is to reduce handling time without increasing material corrections or exposing restricted information.
| QSAFE area | Pilot test | Illustrative decision rule | Owner |
|---|---|---|---|
| Quality | Representative policy questions with cited sources | No unsupported mandatory instruction; required points present | Service quality lead |
| Safety | Injected documents and attempts to retrieve restricted material | No protected text disclosed; unsafe tool requests blocked | Security owner |
| Alignment | Accessibility, tone and disclosure review | Response follows approved style and identifies AI assistance where required | Policy owner |
| Fit | Controlled comparison with current workflow | Handling time improves without higher material-correction cost | Operations manager |
| Evidence | Versioned regression run and incident drill | Named reviewer can pause and roll back the service | Product owner |

The rules above are examples, not universal thresholds. The organisation must set tolerances from the consequence of error and its own baseline. A low-risk internal drafting tool can allow correction before use; an autonomous action affecting money, safety or rights requires much stronger evidence and controls.
Common evaluation mistakes
- Testing only happy paths: ordinary prompts do not reveal boundary or adversarial failures.
- Using one aggregate score: a critical privacy failure can disappear inside a high average.
- Evaluating the wrong unit: model benchmarks do not prove the application, retrieval and human workflow.
- Changing prompts without regression tests: an improvement for one task can damage another.
- Counting adoption as value: messages generated and users enrolled are activity measures, not business outcomes.
- Launching without stop rules: monitoring is ineffective when nobody knows when to intervene.
Generative AI evaluation checklist
- Is the intended user, task, context and prohibited use documented?
- Does the test set represent normal, boundary and adversarial cases?
- Are quality rubrics tied to the real decision and error cost?
- Have retrieval, tools, permissions and human review been tested together?
- Are privacy, security, disclosure and escalation requirements verified?
- Is performance compared with a credible operational baseline?
- Are model, prompt, data and grader versions reproducible?
- Do named owners control thresholds, rollback and post-launch monitoring?
Build the capability to evaluate before scaling
EPW’s five-day Generative AI Models and Applications course covers model families, prompt and output control, retrieval-augmented generation, adaptation, factuality and usefulness assessment, red-team testing, security, cost, monitoring and implementation planning. Those capabilities support a complete evaluation conversation rather than a tool demonstration.
Professionals can also explore EPW’s Artificial Intelligence and Machine Learning Courses and the AI and machine learning article hub. To apply QSAFE to a real use case, review the course outline, available dates and locations, or request tailored course details.
Sources and references
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, 2024.
- National Institute of Standards and Technology, AI RMF Playbook, accessed 1 September 2026.
- OWASP GenAI Security Project, LLM01: Prompt Injection, accessed 1 September 2026.
- European Commission, Guidelines on AI transparency obligations, updated August 2026.
- EPW Training, Generative AI Models and Applications Course, accessed 1 September 2026.
