Skip to main content
AI Research

Why we’re building Vinci MLE

George PuSeptember 28, 20263 min read
Why we’re building Vinci MLE

Today we're releasing Vinci MLE 1.0: two open-weight research models for machine-learning engineering, in 8B and 30B sizes.

They build on IBM Granite 4.1, with our specialization developed in Canada, and we're releasing the weights under Apache-2.0.

The reason we're building them is personal and practical. I want AI to become a useful collaborator in the work of building AI.

Not just explaining a concept or writing a training script. I want help inspecting what actually happened in an experiment, understanding what the evidence supports, and deciding what to do next.

That's the direction of Vinci MLE.

Why this work

Consider the questions around an ML experiment.

  • Why did a training run fail to resume?
  • Is a better validation score evidence of a better model, or a problem in the data split?
  • Is a numerical failure caused by the configuration, or do we need more evidence before touching anything?

These are the kinds of questions I want our models to become useful at answering. They connect directly to the work we want help with at Vinci.

The longer-term ambition is to give researchers and small teams more capacity to do that work.

If AI can become a dependable collaborator in experiments and engineering, it could help people build better AI systems themselves.

That is an ambition we need to earn through results. This release gives the program a concrete starting point.

What I like about 1.0

The training is organised around ML workspaces, rather than only question-and-answer pairs. Examples combine file inspection, tool feedback, diagnosis, bounded changes, and a final report of what is known, what changed, and what remains uncertain.

The specialization corpus contains 91 trajectories across 64 distinct task roots in eight families.

Those families include data quality, experiment interpretation, failed resumes, memory and batching, evaluator defects, numerical failures, packaging and serving, and tokenizer or model configuration.

The part that matters most to me is the decision the model is being trained to make.

Sometimes the right response is to act. Sometimes it is to leave a correct setup alone. Sometimes it is to stop and ask for missing evidence.

All three are represented in the training. I don't want an engineering assistant that treats every request as permission to change something. I want one that can justify whether a change belongs there in the first place.

There is also a useful measured result to build on: on our recorded BFCL tool-calling evaluation, both sizes finished within 0.6 percentage points of their respective Granite parents. That is a narrow result about retaining general tool-calling capability after specialization, not proof of better ML-engineering performance.

And the weights are open. Researchers can download a complete merged checkpoint, run it on infrastructure they control, and evaluate it in their own setting. The model cards include the training details, evaluation results, loading examples, and limitations.

What is still missing

The main missing result is a valid held-out ML-engineering task evaluation. We have measured general capabilities, but those tests do not establish whether the specialization improves performance on new MLE tasks.

That's why we're releasing 1.0 as a research baseline. The weights are available for experimentation; they are not a complete autonomous engineering system. Anyone connecting them to tools needs to provide and validate the workspace, permissions, execution environment, and human review.

For MLE 1.1, we intend to make the held-out MLE evaluation and external MLE benchmark suite release-gating measurements, alongside checks on general capability retention. The next release should answer more of the question this one is setting up.

Why we're building it here

I want Canada to contribute more of the AI that people build with.

For Vinci, that means developing models and the systems around them, testing the work, and releasing things other people can use and examine.

It also means building on the contributions of others. IBM developed Granite; our contribution here is the ML-engineering specialization and its release.

I don't see those ideas as being in conflict. Building capability in Canada and participating in a global open-model ecosystem belong together.

The goal is to make something useful beyond our own team and beyond our own country.

Vinci MLE 1.0 is our first release in that direction. The 8B and 30B model cards and weights are available on Hugging Face. I would especially value feedback from people willing to test them on bounded ML workflows and share examples we can reproduce.

We're going to keep building, measuring, and improving from there.

— George

Explore Vinci MLE · 8B model card and weights · 30B model card and weights

George Pu

George Pu

George Pu is the founder and CEO of SimpleDirect, an independent Canadian AI lab developing research, models, technologies, and products under Vinci.

SimpleDirect® is an independent Canadian AI lab in Toronto. Vinci names the AI research, models, technologies, and products developed by SimpleDirect.

Share