Vinci MLE 1.0 · Open-weight release
MLE 123B 1.0
ML-engineering specialization on Devstral 2. Developed in Canada by SimpleDirect®, with merged BF16 weights and three GGUF quantizations for use in your own environment.

What it is trained to help with
It is trained to read files and experiment results, look for problems, and explain what to do next.
The training uses 91 examples of multi-step work across 64 tasks. It covers eight areas: data quality; understanding experiments; restarting failed training; managing memory and batches; checking how results are scored; unstable training; packaging and running models; and matching model settings.
Change it. Leave it. Ask for more.
- Make a change
- Suggest a limited fix when the evidence supports it.
- Leave it alone
- Keep a working setup as it is when a change is not justified.
- Ask for more evidence
- Explain what is missing instead of guessing at a repair.
This describes the training, not proven results on new tasks.
What we have—and haven’t—tested
Evaluation coverage differs by model. Valid held-out ML-engineering and external MLE-bench evaluations were not completed for 1.0. We plan to require those evaluations before releasing MLE 1.1.
Each model card compares measured general capabilities with its own parent model and documents gains, tradeoffs, and missing evaluations. These results do not establish performance on new ML-engineering tasks.
In the BFCL tool-use evaluation, 123B scored within the preregistered retention margin of its Devstral 2 parent. Descriptive general-capability results include GPQA at 63.64% versus 57.07% for the parent, alongside a 4.8-percentage-point decline on MATH-500 across 500 questions. Coding evaluation was not completed; no score is reported or inferred. A fixed BF16 tool-format smoke test passed 12 of 12 cases, which is not a guarantee of reliable tool use.
See the full results and limitationsRun it on your own setup
Model formats: merged BF16 weights (250.05 GB), Q4_K_M GGUF (74.9 GB), Q5_K_M GGUF (88.3 GB), and Q8_0 GGUF (132.9 GB). Runtime needs extra memory. Each GGUF produced non-empty responses for six fixed neutral prompts; the single tool prompt did not produce a parseable tool call. These checks do not establish task-level parity with BF16.
123B uses Mistral’s Modified MIT licence, unlike the Apache-2.0 8B and 30B models. It grants no rights when your company’s or employer’s global consolidated revenue exceeded US$20 million in the preceding month. Review the full licence before use.
These are downloadable research models, not a ready-made app or hosted assistant. To let a model work with files or run commands, your application needs to provide the tools, a controlled workspace, and permissions. Review and test any changes it suggests.
8B offers a smaller download and lower memory needs. 30B scored higher than 8B in our coding and tool-use tests. 123B uses Devstral 2 and a different licence. These results do not rank the models on ML-engineering tasks.