AI Feature Requirements Worksheet
The requirement said "summaries shall be accurate." Three weeks into testing, nobody could agree on what would fail it.
A support team was adding AI summaries of customer conversations to its case tool. The specification had a line for accuracy, a line naming the model, and an acceptance criterion copied from the last deterministic feature: the expected output for a set of sample conversations. The model never produced the same summary twice, so every run failed. The testers rewrote the criterion as "looks reasonable", which passed everything, including the summary that invented a refund nobody had promised.
(Composite example, assembled from patterns that recur when teams specify AI features with deterministic requirements. The organisation is not a real one.)
The requirement still exists when the output varies; only the way you verify it changes. This worksheet walks one feature through ten specification moves that account for non-determinism without giving up on quality. Attach it to any AI spec.
Design principle: a wrong answer can be worse than no answer. For many AI features, the failure modes matter more than the success criteria, so don't leave that section blank.
Start here: name the feature
Feature: ______________________
Functional requirement: The system shall ______________________
☐ Written as "the system does X", not "the model can do X", and not "the system uses [named model]".
1. Specify behaviour properties, not exact outputs
List the properties a good output must have, each testable even when the wording varies.
☐ Property: ______ (e.g. mentions every issue raised in the conversation)
☐ Property: ______ (e.g. between 50 and 200 words)
☐ Property: ______ (e.g. states no facts absent from the input)
☐ Property: ______ (e.g. professional tone)
2. Split hard and soft properties
Hard properties get automated checks. Soft properties need human review, model-graded evaluation or an aggregate metric.
Hard: deterministically checkable | Soft: needs human or aggregate evaluation |
Output is valid JSON | Captures the main themes |
Between 50 and 200 words | Tone is professional |
References only entities in the input | Recommendation is relevant |
______ | ______ |
☐ Every hard property has, or will have, an automated check.
☐ Every soft property has a named evaluation method.
3. Set pass-rate thresholds for soft properties
A percentage acknowledges non-determinism: the requirement bounds the rate of good outputs, not each instance.
Requirement: At least ____% of outputs shall pass ______ [review type] on a sample of ____ representative inputs, with ______ [e.g. no factual errors of consequence].
☐ The threshold is a pass rate, not a pass/fail absolute.
4. Name the test set and golden examples
The evaluation set is part of the requirement. When it changes, the requirement may need to change too.
Named evaluation set: ______ (e.g. eval-set-v1, 200 examples)
☐ Golden examples captured for the cases that matter most, to catch regressions.
☐ A rule for what happens to the requirement when the evaluation set changes.
5. Bound the failure modes
State what is unacceptable regardless of the pass rate.
☐ Shall not state facts or entities absent from the input.
☐ Shall not include personal data that wasn't required.
☐ Below the confidence threshold, shall say it doesn't know rather than guess.
☐ Other prohibited failure: ______
6. Set latency and cost limits
Latency: shall complete in under ____ seconds at p95 under ____ [named load]
Cost: average cost per operation shall not exceed ______
7. Monitor drift and degradation
Quality can move when the model changes, when a provider updates it, or when inputs shift.
☐ Monitor: output quality tracked continuously via ______ [e.g. acceptance rate].
☐ Alert: fires when the pass rate drops below ____% over a ____ window.
☐ Version pinning: the system uses a pinned model version, and upgrades require re-evaluation against the named test set.
8. Define human oversight for high-stakes decisions
☐ High-impact outputs are reviewable by a person within ______.
☐ Below-threshold outputs route to a person before any action.
☐ Fallback defined: when the model is unavailable, the system shall ______ (prefer showing nothing to showing a bad suggestion, and never block the user's normal workflow).
9. Specify audit logging
☐ Every AI decision is logged with the model version, the input, the output and the downstream action (accepted, edited or rejected).
10. Set stakeholder expectations early
Say these out loud before shipping:
☐ The output is statistical, not deterministic.
☐ Quality is measured as pass rates, not pass/fail.
☐ Drift needs ongoing monitoring.
☐ Edge cases are common.
Ship gate
Check | Done |
Behaviour properties specified, not exact outputs | ☐ |
Hard/soft split done | ☐ |
Pass-rate thresholds set against a named evaluation set | ☐ |
Failure modes bounded | ☐ |
Latency and cost targets set | ☐ |
Drift monitoring, alert threshold and version pin in place | ☐ |
Human oversight and fallback defined | ☐ |
Audit logging specified | ☐ |
Stakeholders briefed on non-determinism | ☐ |
If any row is empty, you are specifying a deterministic system that doesn't exist.
How to use it
Who: the product owner or business analyst with the QA lead and one engineer, before build starts.
Output: a requirement a tester can verify, with thresholds, a named test set and bounded failure modes.
Done when: every row of the ship gate is checked or carries a written reason.
Where this changes
Low-stakes internal features, where a person sees every output before it matters, can skip formal pass-rate thresholds and start with golden examples and a failure-mode list. Features that influence decisions about people, such as eligibility, hiring or claims, need the oversight and logging sections in full; in the EU, many such uses are high-risk systems under the AI Act, with logging and human-oversight obligations applying from 2 December 2027.
The full argument is in writing requirements for AI/ML features. For acceptance criteria specifically, read acceptance criteria for non-deterministic features, and to test the result, a test strategy for LLM features. More in the AI in Quality & Delivery learning path.
Final takeaway
An AI requirement is still a requirement. Specify the properties, bound the failures and name the test set, and a tester can tell you whether the feature meets it, even though no two outputs match.
Sources
NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, January 2023, the MEASURE and MANAGE functions. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
OWASP, Top 10 for LLM Applications 2025, including misinformation and excessive agency. https://genai.owasp.org/llm-top-10/
Regulation (EU) 2024/1689 (the AI Act), Articles 12 (record-keeping) and 14 (human oversight), as amended by Regulation (EU) 2026/1744 for the Annex III application date. https://eur-lex.europa.eu/eli/reg/2024/1689/oj


