Skip to main content
Vinci Labs

Research that shows its work.

Models, experiments, and the questions that move our work forward.

From research to open weights.

Meet Piccolo, Bozza, and the models that shaped Vinci.

Explore models
Model downloads and limitations
Browse on Hugging Face

Go deeper.

All research and release records

Vinci ProvaResearch report

We tested whether character post-training transferred to Mistral 7B. It did — with trade-offs.

It transferred to Mistral — development-set model-judged fabrication fell from 53.8% to 8.6% — but it cost 5.6 points of GSM8K, and the model now sometimes refuses ordinary work.

Read the article
A closer look at the measurements

We tested whether character post-training transferred to Mistral 7B. It did — with trade-offs.

It transferred to Mistral — development-set model-judged fabrication fell from 53.8% to 8.6% — but it cost 5.6 points of GSM8K, and the model now sometimes refuses ordinary work.

Read the article
Development-set model-judged fabrication; GSM8K 5-shot full test set. These are different evaluation populations, not general real-world hallucination rates.
Measured result
Measured resultBase modelAfter trainingChange
Judged fabricationlower is better53.8%8.6%-45.2 percentage points
GSM8Khigher is better51.6%46.0%-5.6 percentage points
How we publish
  • The question comes before the headline.
  • Success and failure criteria are set against a named baseline, not chosen after the numbers land.
  • Gains are reported next to their costs.
  • Measurements that failed or were disqualified stay in the record, marked unusable, with the reason.
  • Limitations are stated, not implied.
  • Provenance, versions, and pinned revisions are preserved.
  • Artifacts are released where it is responsible to, with the license and access state named.
  • When a finding changes a mainline model, the record says what changed.

Have a question worth exploring?

Work with us