How useful is a small AI model on hardware people already have?
We’re trying it on ordinary computers and asking fairly ordinary questions: does it fit, how long do you wait, and does the answer help?
The model is Vinci Piccolo 1.0, which we released at SimpleDirect® on June 29, 2026.
It has about four billion parameters, the learned numbers inside a model, and is fine-tuned from Qwen3.5-4B for conversation and drafting.
You can download the model under Apache-2.0 to run and adapt it on your own hardware.
This project asks what that experience actually looks like.
We’re working with Conestoga College to find out. The partnership is ongoing through at least the end of November 2026, and these posts share the first findings while the work continues.
Anamika Rawat and Tulsi Patel are full-time co-op student researchers on Conestoga’s team, supervised by Dushyant Puri. They did the benchmarking behind this series.
I’m writing about the results and what I think we should try next; we haven’t independently rerun the hardware tests.
At SimpleDirect®, we work with Conestoga’s Centre for Commercialization (C4C). The benchmarking comes through its SMART Centre, where Dushyant is a senior researcher.
C4C helps entrepreneurs and smaller businesses with intellectual property and commercialization.
SMART works on applied research, technical services and training, and collaborates with C4C on IP services. They’re distinct parts of the college. [1, 2]
Dushyant’s work includes software development, IoT, data processing and business solution integration.
For us, that’s a useful mix: running a model means getting the hardware, software and application to work together. [3]
What we’ve tried so far
The team tested Piccolo on two Macs, Jetson and Raspberry Pi boards, and a 6 GB Android phone. [5]
The Mac mini generated text faster than the Air on all 24 matched prompts. Some requests were also marked truncated, so we still need to look at the answers.
Jetson was faster than Pi in four comparisons, but neither device supplied the price in a simple discount-and-tax question: one observed failure on each.
The phone generated text, with waits measured in minutes and a memory warning.
That’s already enough to change what I’d test next. I want full responses beside the timing numbers, and checks that tell us whether someone actually got what they asked for.
A result from one setup also won’t automatically carry over to another device or task.
For Anamika and Tulsi, this gives the research a real company problem to work on. Comparing systems, explaining unexpected answers and deciding what a measurement can tell you are useful skills beyond this model.
Conestoga’s research program offers paid co-op and part-time opportunities around practical problems like these. [4]
I’m glad we can share their work.
If you’re trying local AI yourself, I hope the results help you choose what to try—and what to check—on your own hardware.
Benchmarking by Anamika Rawat and Tulsi Patel, full-time co-op student researchers on Conestoga College’s team, supervised by Dushyant Puri.
Read the first findings
- We tested a 4B AI model on a Mac mini and MacBook Air
- We tested a 4B AI model on Jetson and Raspberry Pi
- We tested a 4B AI model on a 6 GB Android phone
Sources
- Conestoga College — Centre for Commercialization
- Conestoga College — SMART Centre
- Conestoga College — Dushyant Puri, senior researcher
- Conestoga College — Students in research
SimpleDirect Vinci-Piccolo Benchmarking.pptx, slides 2–37 — phone, edge-board and Mac observations credited to Conestoga SMART Centre. Accompanying Mac CSV exports and the prompt bank underpin the Mac article, where those files are available to download. The deck is cited as a supplied source record and is not offered as a download.
George Pu is the founder and CEO of SimpleDirect, an independent Canadian AI lab developing research, models, technologies, and products under Vinci.