Research by Anamika Rawat and Tulsi Patel, supervised by Dushyant Puri at Conestoga’s SMART Centre, through SimpleDirect®’s collaboration with the Conestoga Centre for Commercialization. George Pu writes about the findings below.
The instruction was to change Day 3. In one of Conestoga’s recorded conversations, Piccolo’s next response also changed the Day 2 total from $45 to $55.
The two items underneath it still cost $30 and $15.
That leaves the person checking a part of the plan they did not ask to change. The assistant has created more work.
This is the kind of problem I want us to understand as we build Vinci. Helping someone get through a task means keeping their decisions intact, carrying corrections forward and eventually giving them something they can use. A response can sound helpful while making all of that harder.
Conestoga’s latest work with our small Piccolo model gives us some concrete examples.
One conversation keeps changing the test
The weekend-planning records are labeled Q4 and Q5. Those labels refer to compressed model configurations. Each record follows one conversation through ten user prompts, as the person changes the plan or corrects the answer.
They start with the same broad request: plan a relaxed three-day weekend, including coffee, shopping, dinner and an activity, within a $200 budget. The conversations then develop differently.
That matters when reading the results. These are useful records of what happened in those interactions. They do not tell us that one compression setting is generally better, because the later prompts and histories are different. We also do not have the exact model hashes, hardware, seeds or complete runtime configuration needed to reproduce every condition.
The earlier Mac, phone and compact-computer studies are separate experiments. Their hardware timings cannot explain these conversation errors.
A plan that says 200 and adds up to 260
In the Q4 record, one revised plan lists day totals of $20, $85 and $155. It says the total is $200.
Those amounts add up to $260.
Later, another answer lists $70, $50 and $75 while claiming $190. The sum is $195. A subsequent self-check still has a mismatch between its listed amounts and its claimed total.
The distinction is practical. Someone gave the model a spending limit. Repeating that limit confidently does not establish that the plan respects it.
The Q5 record has arithmetic problems too, including repeatedly labeling $8 plus $35 as $46. It eventually delivers a final plan whose stated day totals add up to $165. That recovery belongs in the account alongside the mistakes.
These prices come from the test dialogue. They are not verified local prices, and the records do not establish a currency beyond the dollar symbol. Even correct addition would not prove that the proposed weekend could actually be bought for that amount.
Waiting without getting the finished answer
The Q4 recorder notes 96 seconds of thinking with no final response at the last turn. The Q5 record notes 109 seconds without a final response at turn nine; a final plan appears after the next prompt.
We do not know why those answers were absent. The records do not establish a timeout, a crash, a context problem or a cause related to compression.
From the person’s side, the unfinished request still matters. If another prompt is needed to get the result, that adds waiting and supervision. I want us to account for that when we talk about useful AI.
The model’s own score needs checking too
A separate meal-planning document includes an exercise where the model generates both sides of a conversation and grades its own performance.
It awards itself ten out of ten for instruction following and final task success. The structured record declares ten pairs but contains eleven, with the final number repeated.
That score is a model-generated claim. It is not an independent evaluation.
The same document also contains a separate human-led meal conversation. There, the model mixes up weekly and daily budgets, repeats a run of pasta meals after being asked to change it, and later gives cost categories totaling $65–85 while still discussing a $40 household limit. The final table also leaves alternatives where the person had asked it to choose.
A meal plan is an accessible way to expose these problems. It does not establish performance on coding, business operations or other tasks.
What I would check before relying on it
If you are trying a small local model, give it a task whose details you can inspect. Then make one narrow change.
Check five things:
- Did the details outside your requested change stay the same?
- Do the amounts and units still agree, including the final total?
- Did the answer retain the earlier requirements as well as the newest one?
- After a correction, did the same mistake return?
- Did you receive the finished output you asked for, with the decisions you wanted it to make?
Keep the earlier answers. Without them, it is easy to miss a detail that quietly changed. Our guides to verifying AI-generated work and reading model evaluations explain how to keep a useful record.
For the next round, I want the full interaction, independent checks of the answer and a record of how much correction the person had to do. We need that evidence before claiming an improvement.
The goal is straightforward: someone asks for help and has less work left when the answer arrives. These conversations show specific places where we still have work to do.
Source notes
This article draws on the supplied Q4 and Q5 weekend-planning documents and the separate meal-planning document shared by the Conestoga team in October 2026. The observations above concern those recorded interactions. We have not independently rerun them, and the source’s self-scores are not independent quality measurements. The source documents are not offered as public downloads here.
Read the Conestoga collaboration overview for the people and separate hardware studies behind this continuing work.
Research by Anamika Rawat and Tulsi Patel, supervised by Dushyant Puri at Conestoga’s SMART Centre, through SimpleDirect®’s collaboration with the Conestoga Centre for Commercialization. George Pu is the article’s author; the Conestoga team conducted and documented the research.
George Pu is the founder and CEO of SimpleDirect®, the independent Canadian AI company building Vinci for a life beyond the screen.