For people reviewing technical work produced by AI agents
How to verify AI-generated work
Define what completion means, gather evidence from the actual task environment and record what remains unproven. A completed run is not proof of a correct result.
Updated

A script finishes. Did the task succeed?
Hypothetical teaching example, not a Vinci capability result. An agent is asked to generate a report from source records. Its script exits successfully, but that observation alone does not establish that the report contains the right records.
| Observation | What it establishes |
|---|---|
| The script exits without an error | The program completed under those execution conditions. It does not establish correct content. |
| The report matches an expected record set | A comparison against the task’s source records supports completeness for that set. Record exclusions and how the comparison was performed. |
| A reviewer checks values against the source | The reviewed values support a narrower correctness claim. State whether every value or only a sample was checked. |
| The result is accepted | State the acceptance criteria that passed and any unresolved conditions. A local result does not establish behavior in a different deployment. |
More detail and a worksheet
Ready to try this yourself? These extra checks and a worksheet help you keep track of what you used and what happened.
Define acceptance before execution
Verification begins with the requested outcome and its intended consumer. A useful check can distinguish a correct result from a plausible but wrong one.
- Specify the inputs, required outputs and constraints, including behavior that must remain unchanged.
- Include a failing control: a deliberately incorrect result that the check must reject.
- Check under the consumer’s semantics, including variable expansion, links, paths and formats.
Measure the actual environment
Collect evidence where the result will be consumed. An inherited agent environment is not a substitute for a fresh user shell, container, CI job or production deployment.
- Record the artifact version, command or review method, environment and time of observation.
- Keep failures, skipped checks and missing evidence visible; do not count an unperformed check as a pass.
- Separate what was observed from what was inferred. Repeat a failure when repetition is meaningful, after confirming the check can distinguish the states of interest.
Make a bounded completion claim
A verifier supports a particular acceptance decision; it does not grant universal correctness or independent certification. Human review remains necessary when the decision depends on context the checks do not cover.
- State which criteria passed and how many attempts or records were checked.
- Describe the limits: untested cases, inaccessible environments and deployment differences.
- If evidence contradicts the expected result, investigate the difference rather than treating a failed reproduction as proof that the problem is absent.
Record a verification decision
Use this worksheet for one task and artifact version. Write “not checked” when evidence is missing.
Requested outcome: Exact artifact and version: Inputs and intended consumer: Acceptance criteria: Failing control and its result: Environment used for verification: Commands or review method: Observed results and attempt counts: Skipped checks and failures: Evidence locations: Inference, stated separately: Unresolved limits: Acceptance decision and reviewer:
Sources and next steps
A model card is the publisher’s information page for a model. Follow it for the exact download instructions, licence and known limits. External technical sources are in English.