Can you run a useful AI model on a small computer that fits into a DIY or embedded project?
We tried two compact computers: NVIDIA’s Jetson Orin Nano and the Raspberry Pi 5.
The model was Vinci Piccolo 1.0, which we released at SimpleDirect® on June 29, 2026.
It’s a chat model fine-tuned from Qwen3.5-4B, with about four billion parameters—the learned numbers inside the model.
You can download the model under Apache-2.0 to run and adapt it yourself. We wanted to see how that small model handled these two setups.
Conestoga’s first tests give us a starting point: roughly 13–15 generated tokens per second on Jetson, against about 3 on Pi. Tokens are the pieces of text a model generates. [1]
Jetson looks like the setup I’d try first if someone needs an interactive response.
But one of the answers is also a good example of why I don’t want to pick hardware from a speed table alone.
Our Conestoga collaboration continues through at least the end of November 2026.
This is an early look at the work, based on the team’s supplied deck, not a new set of hardware runs.
Four prompts on two boards
Both used llama.cpp to run the model locally.
They ran the same compressed download: Vinci-Piccolo-1.0-GGUF, in Q4_K_M quantization.
Jetson Orin Nano had approximately 7.4 GiB of usable memory and Ubuntu 22.04.5; Raspberry Pi 5 had about 15 GiB and Debian 13. [1, slide 13]
Jetson used NVIDIA’s CUDA GPU acceleration, while Pi ran on its CPU. The Pi screenshots show its neural accelerator, or NPU, as unavailable.
There was an attempt to install a Hailo-8L accelerator, but these results don’t show Hailo running the model. The software setup is part of this comparison too.
The tasks were renewable energy, a TypeScript algorithm, photosynthesis and a discount-and-tax calculation.
Each row is one prompt, run once on each device. Ratios divide the displayed rates; there are no repeated-run averages or uncertainty intervals. [1, slides 15–25]
| Prompt profile | Jetson, tokens/s | Raspberry Pi, tokens/s | Jetson/Pi ratio |
|---|---|---|---|
| Thinking – General | 13.98 | 2.95 | 4.74× |
| Thinking – Precise Coding | 14.77 | 3.21 | 4.60× |
| Instruct – General | 13.41 | 3.19 | 4.20× |
| Instruct – Reasoning | 13.19 | 3.00 | 4.40× |
That’s a 4.2–4.7× generation-speed difference. The first token arrived in about 327–375 milliseconds on Jetson, versus 2.30–3.66 seconds on Pi.
In thinking mode, the model generates reasoning text before its answer, so the first token may not be the information someone wants.
For renewable energy, the full request took 171.06 seconds on Jetson and 826.31 on Pi: about two minutes and 51 seconds versus 13 minutes and 46 seconds.
Outputs were 2,387 and 2,433 tokens. Response length matters alongside speed. [1, slide 16]
The $226 that didn’t arrive
The last prompt asked for a $250 item after a 20 percent discount and 13 percent sales tax:
$250 × 0.80 × 1.13 = $226
Jetson returned a JSON object resembling a call to calculate_price_with_discount_and_tax, with the right arguments, but no calculated price. Pi declined to do the calculation.
Both requests were marked completed, without truncation or a runtime error. That’s one observed failure on each device to answer this question. [1, slide 26]
It doesn’t mean Piccolo can’t do arithmetic. We need to check the chat template, format, settings and whether the application was supposed to execute a tool call.
If I ask an app for a price, I expect a price.
Next, I want full interactions, answer checks and repeated prompts in fresh contexts.
We also need the exact model file, llama.cpp build, launch commands, sampling settings, output limits, board power mode, cooling and starting conditions.
Pi’s CPU results give us something to compare with a verified accelerator setup later.
Test notes
Substantial RAM was occupied before requests started; these readings don’t isolate Piccolo’s memory use or show it needing twice as much on one board.
Temperature sensors differ, so GPU/junction and CPU readings can’t rank whole-system cooling.
Pi lacks power and energy readings comparable to Jetson’s. There’s no energy-efficiency ranking, broad accuracy score or repeat-run reliability result here. [1, slides 16–26]
Benchmarking by Anamika Rawat and Tulsi Patel, full-time co-op student researchers on Conestoga College’s team, supervised by Dushyant Puri.
Sources
[1] SimpleDirect Vinci-Piccolo Benchmarking.pptx, slides 13–26, credited to Conestoga SMART Centre. Configuration: slide 13. Prompts: slides 15, 18, 21, 24. Metric tables: slides 16, 19, 22, 25. Output and monitoring screenshots: slides 17, 20, 23, 26.
The supplied SimpleDirect Vinci-Piccolo Benchmarking (1).pptx contains the same substantive edge-board observations; it is not an independent replication.
Continue the series: the Mac comparison, the Android phone test, and our ongoing Conestoga collaboration.
George Pu is the founder and CEO of SimpleDirect, an independent Canadian AI lab developing research, models, technologies, and products under Vinci.