I like running a small AI model on a computer you already own. Before recommending a setup, though, I want to know how long you wait, and whether the answer is worth it.
For this test, we used Vinci Piccolo 1.0, a chat model we released at SimpleDirect® on June 29, 2026.
It has about four billion parameters, the numbers a model learns during training. That’s what “4B” means; it isn’t a 4 GB memory requirement.
We fine-tuned it from Qwen3.5-4B for conversation and drafting. The model is downloadable under Apache-2.0, so you can run and adapt it yourself.
The Conestoga team tried Vinci-Piccolo on a 2023 M2 Pro Mac mini and a 2024 M3 MacBook Air with 8 GB of RAM.
In September’s 24-prompt comparison, the Mini generated text faster on every matched prompt. Its mean rate was 45.30 tokens per second, against 20.29 on the Air: 2.23 times as fast. [1, 2]
I’d choose the Mini for the next round of testing.
There’s a catch, though: all 48 requests were marked completed, and nine were also marked truncated. We still need to check the answers.
Our Conestoga collaboration continues through at least the end of November 2026; these are the first findings from that work.
How the two Macs did
Both Macs used llama.cpp to run the model locally, with Apple’s Metal graphics acceleration.
They ran the same compressed download: Vinci-Piccolo-1.0-GGUF, in Q4_K_M quantization.
Each profile had six different prompts covering writing, coding or reasoning, with one run per device per prompt. These averages include truncated runs.
Tokens are the pieces of text the model generates; “thinking” refers to reasoning text generated before an answer. [1–4]
| Workload profile | Mac mini, tokens/s | MacBook Air, tokens/s |
|---|---|---|
| Thinking – General | 44.20 | 21.24 |
| Thinking – Precise Coding | 47.01 | 19.46 |
| Instruct – General | 44.90 | 19.06 |
| Instruct – Reasoning | 45.11 | 21.40 |
| All 24 prompts | 45.30 | 20.29 |
The first generated token arrived after an average 640 milliseconds on the Mini, versus 1,415 on the Air. The Mini was quicker on 23 of 24 prompts.
First-token timing doesn’t tell us when a useful final answer appeared, especially when the model is producing reasoning text. [1, 2]
Request times added up to 20.81 minutes on the Mini and 48.18 on the Air, excluding time between requests. They generated 55,750 and 57,582 output tokens, respectively.
Some individual answers differed substantially in length, so the total wait reflects both speed and how much text each produced. [1, 2]
Did it finish?
| Recorded status | Mac mini | MacBook Air |
|---|---|---|
| Requests marked completed | 24 of 24 | 24 of 24 |
| Runs marked truncated | 3 of 24 | 6 of 24 |
| Runs without recorded truncation | 21 of 24 | 18 of 24 |
All nine truncated runs were in thinking profiles, each with exactly 4,096 tokens. That looks worth checking against the output settings, but the exports don’t include the cap or exact finish reason.
One reasoning task produced just 18 tokens on the Mini; a different task produced 30 on the Air.
Short or long, we need the response text to judge it. There are no answer-quality scores here. [1, 2]
Next, I want to repeat the comparison with full answers and clear checks for each task.
We should save the model file, runtime build, launch and sampling settings, and starting conditions.
Thinking and instruct used different prompts here; comparing that setting fairly means giving both the same task.
Test notes
These figures come from the team’s exports; we haven’t independently rerun the hardware tests.
After normalizing whitespace, all 24 prompt texts and input-token counts matched. Inputs were 100–212 tokens; we haven’t tested very long prompts here.
Each prompt was run once on each Mac, so we don’t yet know how much repeated runs vary.
The Air’s largest recorded swap value was about 11,115 MB, versus 1,223 MB on the Mini; the Air already showed 5,975 MB in its first observation.
Swap and RAM readings include the rest of the system, not just Piccolo.
The deck names psutil and powermetrics as monitoring tools, and its Mini memory label still needs verification.
We can’t separate chip, RAM and cooling effects here, or infer wall-plug energy, battery life or throttling from these readings. [1–3]
Benchmarking by Anamika Rawat and Tulsi Patel, full-time co-op student researchers on Conestoga College’s team, supervised by Dushyant Puri.
Sources and calculation notes
You can download both CSV exports and the prompt bank below. The CSVs don’t include full answers or quality scores.
- Mac mini CSV export (
macmini_20260929_141518_all_profiles.csv) — 24 prompt observations. TheTOTAL PROMPT TIMEfooter is a summary, excluded from row-level calculations. - MacBook Air CSV export (
macair_20260929_131947_Vinci-Piccolo-1.0-GGUF_Q4_K_M.csv) — 24 prompt observations, paired with the Mini after whitespace normalization. SimpleDirect Vinci-Piccolo Benchmarking.pptx, slides 27–37 — Mac configuration and results, credited to Conestoga SMART Centre.- Mac prompt bank (PDF) (
Prompts MacAir & Mac Mini.pdf), pages 1–7 — four workload profiles with six prompts each.
Speed and first-token figures are arithmetic means of the per-request fields. The 2.23× ratio uses unrounded means; duration totals sum each device’s 24 request rows.
Continue the series: Jetson and Raspberry Pi, the Android phone test, and our ongoing Conestoga collaboration.
George Pu is the founder and CEO of SimpleDirect, an independent Canadian AI lab developing research, models, technologies, and products under Vinci.