Skip to main content
Researchpiccolo

We tested a 4B AI model on a 6 GB Android phone

George PuOctober 5, 20264 min read

Can a phone you already own run an AI model locally?

The Conestoga team tried a Samsung S20 FE with 6 GB of RAM, and got it generating text.

I’m glad to see output on a phone that isn’t the latest flagship. I also want to understand the wait, a memory warning, and whether the answers were any good.

We used Vinci Piccolo 1.0, a chat model we released at SimpleDirect® on June 29, 2026.

It has about four billion parameters, the numbers learned during training, and is fine-tuned from Qwen3.5-4B. That’s the “4B” in the title, separate from the phone’s 6 GB of memory.

You can download the model under Apache-2.0 to run and adapt it on your own devices.

The team ran the model in Off Grid AI, an app for running AI locally. The download was Vinci-Piccolo-1.0-GGUF.

A separate app monitored CPU, GPU, battery and memory.

This was one phone, one prompt and two sequential observations.

The figures below are from the supplied deck and app screenshots; we haven’t independently repeated the test. [1]

Our collaboration with Conestoga continues through at least the end of November 2026, so this is a first look, with more to investigate.

How long did it take?

The prompt was:

Write me a 200-word summary on aquatic life.

The team tried it with thinking enabled, then disabled. Thinking mode lets the model generate reasoning text before its answer.

Here’s what the app reported. Tokens are pieces of generated text, not a count of words in the final answer. [1, slides 6–11]

App-reported measurement Thinking enabled Thinking disabled
Generation rate 5.1 tokens/s 2.7 tokens/s
Generated tokens 1,024 262
Generation duration 201.8 seconds 96.7 seconds

That’s about three minutes and 22 seconds with thinking and one minute and 37 seconds without it.

The thinking run generated tokens faster, but produced more of them and took longer.

I wouldn’t read that as “thinking makes it faster.”

The runs happened in a fixed order, with starting CPU readings of 31°C and 42°C. Both prompts appear in the same conversation, so the second may have carried extra context.

We need fresh conversations and repeated runs to compare that setting fairly. [1, slides 6–9]

The exact 1,024-token count is worth checking against a possible output limit. We don’t have a confirmed cap or finish reason, so we can’t say it was cut short.

Nor do we know whether either answer met the requested length or quality.

For a quick question, I’d find these waits long. A background task could be different.

Streaming might let someone start reading sooner, but this test didn’t measure when useful information appeared.

I want to try a few real phone tasks, check the answers, and see which ones make sense at this speed.

The memory warning matters too

A warning appeared during setup. It tells us to look more closely at this setup; it doesn’t give us a minimum RAM requirement for Piccolo.

We need the exact model file, app settings and available memory first. The records also don’t identify the mobile model’s compression settings or the engine the app used to run it. [1, slide 7]

The monitor reported 84°C CPU readings during or after both responses. Those are internal sensor values, not surface temperatures or a battery-safety assessment.

Two snapshots don’t tell us about sustained heat or prove throttling, and can’t explain the difference in speed. [1, slides 8–11]

Next, I’d record the missing settings and starting conditions, repeat a small set of tasks, and save the full answers.

We’ve seen output on a 6 GB phone. I want to find out what I’d actually use it for.

Benchmarking by Anamika Rawat and Tulsi Patel, full-time co-op student researchers on Conestoga College’s team, supervised by Dushyant Puri.

Sources

[1] SimpleDirect Vinci-Piccolo Benchmarking.pptx, slides 2–11: device and applications, slides 2–3; prompt and setup, slides 6–7; generation figures and monitor readings, slides 8–11. The deck credits Conestoga SMART Centre.

The supplied SimpleDirect Vinci-Piccolo Benchmarking (1).pptx contains the same substantive phone observations, rather than an independent replication. There are no repeated-trial averages or uncertainty estimates.

Continue the series: the Mac comparison, Jetson and Raspberry Pi, and our ongoing Conestoga collaboration.

George Pu

George Pu is the founder and CEO of SimpleDirect, an independent Canadian AI lab developing research, models, technologies, and products under Vinci.

SimpleDirect® is an independent Canadian AI lab in Toronto. Vinci names the AI research, models, technologies, and products developed by SimpleDirect.

Share