Skip to main content
All resources

For people reviewing technical work produced by AI agents

How to verify AI-generated work

Define what completion means, gather evidence from the actual task environment and record what remains unproven. A completed run is not proof of a correct result.

Updated

A script finishes. Did the task succeed?

Hypothetical teaching example, not a Vinci capability result. An agent is asked to generate a report from source records. Its script exits successfully, but that observation alone does not establish that the report contains the right records.

ObservationWhat it establishes
The script exits without an errorThe program completed under those execution conditions. It does not establish correct content.
The report matches an expected record setA comparison against the task’s source records supports completeness for that set. Record exclusions and how the comparison was performed.
A reviewer checks values against the sourceThe reviewed values support a narrower correctness claim. State whether every value or only a sample was checked.
The result is acceptedState the acceptance criteria that passed and any unresolved conditions. A local result does not establish behavior in a different deployment.
More detail and a worksheet

Ready to try this yourself? These extra checks and a worksheet help you keep track of what you used and what happened.

Define acceptance before execution

Verification begins with the requested outcome and its intended consumer. A useful check can distinguish a correct result from a plausible but wrong one.

  • Specify the inputs, required outputs and constraints, including behavior that must remain unchanged.
  • Include a failing control: a deliberately incorrect result that the check must reject.
  • Check under the consumer’s semantics, including variable expansion, links, paths and formats.

Measure the actual environment

Collect evidence where the result will be consumed. An inherited agent environment is not a substitute for a fresh user shell, container, CI job or production deployment.

  • Record the artifact version, command or review method, environment and time of observation.
  • Keep failures, skipped checks and missing evidence visible; do not count an unperformed check as a pass.
  • Separate what was observed from what was inferred. Repeat a failure when repetition is meaningful, after confirming the check can distinguish the states of interest.

Make a bounded completion claim

A verifier supports a particular acceptance decision; it does not grant universal correctness or independent certification. Human review remains necessary when the decision depends on context the checks do not cover.

  • State which criteria passed and how many attempts or records were checked.
  • Describe the limits: untested cases, inaccessible environments and deployment differences.
  • If evidence contradicts the expected result, investigate the difference rather than treating a failed reproduction as proof that the problem is absent.

Record a verification decision

Use this worksheet for one task and artifact version. Write “not checked” when evidence is missing.

Requested outcome:
Exact artifact and version:
Inputs and intended consumer:
Acceptance criteria:
Failing control and its result:
Environment used for verification:
Commands or review method:
Observed results and attempt counts:
Skipped checks and failures:
Evidence locations:
Inference, stated separately:
Unresolved limits:
Acceptance decision and reviewer:

Sources and next steps

A model card is the publisher’s information page for a model. Follow it for the exact download instructions, licence and known limits. External technical sources are in English.