Skip to main content
Research

Why Vinci started with post-training

George PuSeptember 30, 20266 min read
Conceptual illustration of two ivory-and-sage block structures on a stone workbench, with loose foundation blocks beside them.

Since the beginning of Vinci, I’ve wanted us to pre-train our own models. There is something deeply compelling about choosing the data, building the training process, and taking responsibility for a model’s foundations.

We plan to pre-train 4B and 10B models. That ambition is part of our direction. Our first public releases, though, have focused on post-training existing open-weight models. I want to explain why.

Pre-training builds a model’s foundation by learning from a large body of material. Post-training adapts an existing foundation toward particular tasks and behaviours. The distinction matters when assessing what a Canadian AI company has built, what it controls, and what further investment would support.

Open weights are a starting point

Open weights—the downloadable numerical values that define a model—let us run, adapt, and evaluate it ourselves. That is a substantial advantage. But access to weights does not necessarily give us a complete account of the data or decisions behind them. Disclosure varies across model families.

For enterprises and public institutions, I understand the appeal of working with a company that controls more of that process: where the data came from, what permissions apply, and how the model was developed and tested. Those questions help buyers assess accountability, privacy, and suitability for a particular use.

Pre-training gives us more responsibility and more control. It does not automatically make a model safe. Data governance, evaluations, security, and deployment controls still have to earn that confidence.

The economics are real

Even relatively small models involve substantial compute. The 4B and 10B labels describe billions of parameters: the adjustable values a model learns. Tokens are the pieces of text it processes during training. Model size is only one part of the bill; how much material it learns from matters just as much.

For a dense transformer, a useful first approximation is training compute ≈ 6 × parameters × training tokens, as used in the Chinchilla research. Holding model size and effective hardware performance constant, three times the training tokens means roughly three times the main training compute.

Here are hypothetical examples, assuming effective throughput of 300 TFLOPS per GPU and CA$7 per GPU-hour. Throughput means the useful training work each graphics processor sustains, rather than its advertised peak speed.

Training tokens per model 4B main training compute 10B main training compute
1 trillion About CA$156,000 About CA$389,000
3 trillion About CA$467,000 About CA$1.17 million
10 trillion About CA$1.56 million About CA$3.89 million

One trillion tokens is a comparison point, not a complete development budget or a claim that training should stop there. At ten trillion tokens each, the two models together would require about CA$5.44 million for main training compute alone under these assumptions.

These are calculated illustrations, not measured Vinci throughput, vendor quotations, or our execution budget. None of these token counts is an announced Vinci training target. The approximation does not capture every architecture or long-context cost, and lower effective performance raises the bill: at 200 rather than 300 TFLOPS per GPU, the ten-trillion-token 10B example rises to about CA$5.83 million.

The figures also exclude data preparation, separate experiments, evaluation, post-training, reruns, and people. A full programme costs more than its successful main run.

Why we may need more training tokens

A first training milestone can establish that the pipeline works without producing the model we ultimately want. Broad language ability, French and English coverage, code, mathematics, and technical knowledge all place demands on the training material and recipe. We may need substantially more exposure to suitable data before a small model meets our intended quality targets.

There is a practical reason to explore longer training at a fixed size: a compact model may be easier and cheaper to run repeatedly once deployed. Research on training and inference economics examines that tradeoff between greater upfront training and lower ongoing serving costs.

This is not hypothetical across the field. Hugging Face reports training its 3B SmolLM3 on 11.2 trillion tokens, with further stages afterward. That is a documented example of a small model with a large training budget, not a requirement that Vinci copy its recipe or a direct comparison of costs.

More tokens are not a guarantee of better results. Quality, permissions, diversity, and the learning recipe matter. Training exposures also differ from unique data: processing the same material twice consumes compute twice without creating new material. Evaluations must tell us whether further training is helping.

That is before considering larger dense models or ambitious mixture-of-experts programmes. Pre-training is achievable. Choosing where to begin takes discipline.

Why post-training came first

Post-training lets us build on foundations that already exist and focus our work on the behaviour and tasks we want Vinci to support. We can improve our data, evaluation methods, and engineering workflows while putting models into developers’ hands.

Our Vinci MLE releases are part of that work. Their cards identify their parent models and report the evidence and its limits. They are adaptations of existing foundations; they are not models we pre-trained from scratch.

For me, this is a sensible sequence. Keep delivering useful work while developing the expertise to make our own foundations. Native pre-training should improve the product enough to justify its cost.

A Canadian ambition, on our own timetable

I believe pre-training belongs in Canada’s AI sovereignty story. For me, sovereignty means having meaningful choices: who controls the model, who can examine and change it, where it can operate, and whether Canadian teams can maintain and improve it.

Developing models here, with Canadian ownership of the process and the expertise to repeat it, builds a capability we can direct ourselves. The public value includes skilled people, accountable data practices, and an ability to evaluate systems for Canadian needs—not just ownership of a downloadable file.

That does not remove dependence on chips, software, or other parts of the global supply chain. It does give us more agency over an important part of it.

At Vinci, we have chosen to remain independent rather than take foreign investment at this stage. That reflects how we want to set our technical priorities and make long-term decisions. I believe we can finance an initial, modest pre-training effort ourselves; larger ambitions bring much larger economics.

I’m excited to get started. We’ll share concrete milestones as the work earns them, rather than promise a release date today. In the meantime, post-training remains a productive part of our research and our releases.

4B and 10B are on our roadmap. Stay tuned.

Cover: AI-generated conceptual illustration, not a depiction of a training facility or a completed run.

George Pu

George Pu is the founder and CEO of SimpleDirect, an independent Canadian AI lab developing research, models, technologies, and products under Vinci.

SimpleDirect® is an independent Canadian AI lab in Toronto. Vinci names the AI research, models, technologies, and products developed by SimpleDirect.

Share