# AI test-set template

Workflow: [ ]  Version/model/configuration: [ ]  Reviewer: [ ]  Date: [ ]

Choose 50 permitted, representative examples before changing the system. Include common tasks, edge cases, missing inputs and unsafe requests. Avoid storing personal or confidential data in this file; reference an approved secure source.

## Agree acceptance criteria first
- Task success threshold: [ ]
- Maximum unsupported-claim rate: [ ]
- Required human approval actions: [ ]
- Forbidden actions or data disclosure: [ ]
- Maximum latency: [ ]  Maximum cost per task: [ ]
- Minimum samples per important category: [ ]

## Results table — add one row per example
| ID | Category | Secure input reference | Expected result | Actual result reference | Score 0–2 | Unsupported claim? | Safety failure? | Latency | Cost | Reviewer notes |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| 001 | Common | | | | | | | | | |
| 002 | Edge case | | | | | | | | | |
| 003 | Missing information | | | | | | | | | |
| 004 | Unsafe instruction | | | | | | | | | |

Score: 0 = wrong/unusable; 1 = partial, needs substantial correction; 2 = meets agreed criteria.

## Summary
Number tested: [ ]
Success rate = examples scoring 2 / total tested: [ ]
Unsupported claims = examples with unsupported claims / total tested: [ ]
Safety failures: [ ] (review each individually)
Median / worst latency: [ ] / [ ]
Average / maximum cost: [ ] / [ ]
Compare with previous version on the same examples: [ ]

Release decision: [approve / revise / reject]
Known gaps, owner and remediation: [ ]
Human approver and date: [ ]
