Vinci Technical Report No. 3
Runtime Pass Is Not Correctness
A failed training recipe—and what auditing its evaluator revealed.
Read the technical reportModels, experiments, and the questions that move our work forward.
Vinci Technical Report No. 3
A failed training recipe—and what auditing its evaluator revealed.
Read the technical reportVinci Technical Report No. 2
A three-family study of behavioural gains and capability costs.
Read the technical reportVinci Technical Report No. 1
Testing whether character training transfers to a new model family.
Read the technical reportMeet Piccolo, Bozza, and the models that shaped Vinci.
Vinci ProvaResearch report
A conservative SFT+DPO recipe missed every reasoning-efficiency target — and auditing the evaluator showed the original bank accepted 24 of 24 deliberately incorrect programs.
Read the articleVinci ProvaResearch report
Unsupported assertions fell in all three families under both judges — but every family lost too much grounded-answer accuracy to meet the pre-registered bar.
Read the articleVinci ProvaResearch report
It transferred to Mistral — development-set model-judged fabrication fell from 53.8% to 8.6% — but it cost 5.6 points of GSM8K, and the model now sometimes refuses ordinary work.
Read the articleVinci FlagshipPost-mortem
The honesty gain was explained by output length — and the measurement that produced it was disqualified.
Read the articleVinci FlagshipModel release
Adversarial attack success reached 0.0% and citation integrity rose, at the cost of 10.7 points of instruction-following — and three flattering benchmark jumps were thrown out as extraction artifacts.
Read the articleVinci FlagshipModel release
The first 4B mainline release: strong safety numbers and solid general ability for its size, with tool-calling at 23.0%.
Read the articleIt transferred to Mistral — development-set model-judged fabrication fell from 53.8% to 8.6% — but it cost 5.6 points of GSM8K, and the model now sometimes refuses ordinary work.
Read the article| Measured result | Base model | After training | Change |
|---|---|---|---|
| Judged fabricationlower is better | 53.8% | 8.6% | -45.2 percentage points |
| GSM8Khigher is better | 51.6% | 46.0% | -5.6 percentage points |

Vinci released Vinci Cyber, a family of three open-weight models built for defensive cybersecurity: 8B, 30B, and 123B parameters.

We applied one frozen character post-training recipe to Qwen3, Ministral, and OLMo. Unsupported assertions fell in every family, but grounded-answer accuracy fell too far in every family. We are publishing the failure, not a model checkpoint.

Fine-tuning was enough to change our models' voice. It was not enough to build the models we now want to release. We are moving from light behaviour tuning to full-weight model redevelopment - open foundations, rebuilt by Vinci. Here is the plan, stage by stage.