Skip to main content
All resources

For anyone reading an AI claim or comparison

What do model test results actually tell us?

See how a model can make fewer mistakes yet answer fewer questions correctly. Learn what to ask before trusting a headline score.

Updated

Fewer wrong answers can hide another change

Hypothetical teaching example — these are not Vinci results. Two models answer the same 100 questions. Each answer is marked correct, incorrect, or declined: the model chose not to answer. For Vinci MLE 1.0, a valid test on ML-engineering tasks set aside for assessment was not completed; this example is not evidence of its capabilities.

What happenedHow to read it
Model A: 70 correct, 20 incorrect, 10 declined70 out of 100 answers were correct: 70/100 = 70%. Keep all 100 questions in the comparison, including the ones the model declined.
Model B: 68 correct, 8 incorrect, 24 declinedWrong answers fell from 20% to 8%: 20 − 8 = 12 percentage points. But correct answers also fell from 70 to 68, while declined answers rose from 10 to 24.
A headline says “Model B is more accurate”That misses part of the story. Model B made fewer mistakes, but also gave fewer correct answers. Whether that trade-off helps depends on what you need it to do.
A fuller description“On these 100 questions, Model B gave fewer wrong answers and declined more often, while correct answers fell from 70 to 68.” We still do not know how it would handle different questions.
More detail and a worksheet

Ready to try this yourself? These extra checks and a worksheet help you keep track of what you used and what happened.

Find the exact model and test

A model family can contain several sizes and versions. A result belongs to the version that was tested, not automatically to everything with the same family name.

  • Find the exact model name, version and downloaded file type in the report.
  • Check how many tasks were attempted, who judged the answers and what counted as correct.
  • Keep results for each Cyber size separate. Testing one size does not establish the results for another.

Ask whether the comparison is fair

The starting system used for comparison is often called the baseline. Check that both systems faced the same questions under the same rules.

  • Compare the time allowed, tools available and number of attempts.
  • Look for counts as well as percentages, including skipped, unfinished and invalid attempts.
  • Check the actual work. A program finishing without an error does not prove that its answer is right; Vinci Technical Report No. 3 examines this distinction.

Look for the unanswered question

A coding test can show how a model handled those coding tasks. It does not, by itself, prove the model can plan and complete a whole machine-learning project.

  • For Vinci MLE 1.0, a valid test on reserved ML-engineering tasks was not completed. These are tasks set aside for assessment rather than used to develop the model; the technical term is a held-out evaluation.
  • Before relying on a result for your own work, try a small task similar to what you actually need. Decide what success means before seeing the answer.
  • When sharing a result, include the model version, the test and its limits, with a link to the report.

Keep a note about a test result

Use one note per model and test. Write “not reported” when the source leaves something out.

What question was this test meant to answer?
Exact model and version:
Report link and date:
What was it compared with?
How many tasks and attempts were there?
Who checked the answers, and how?
Correct / incorrect / declined counts:
Skipped, unfinished or invalid attempts:
Did both systems have the same time and tools?
What improved, and by how much?
What got worse?
What does the test leave unproven?
How similar is it to my intended use?
My one-sentence summary:

Sources and next steps

A model card is the publisher’s information page for a model. Follow it for the exact download instructions, licence and known limits. External technical sources are in English.