Skip to main content
Research

We tried a new AI architecture. It didn't work.

George PuSeptember 28, 20268 min read
Conceptual illustration of three stone and sage architectural studies, with pieces laid beside the central structure.

Most of the AI models people use today are built from variations of the same basic architecture: the transformer.

Vinci uses those models too.

But if we want to develop original AI architectures, we need to do more than train models on designs invented somewhere else.

We need to learn how to build the models themselves.

In September, we took our first experiment through a small, controlled comparison. We called it H1.

It didn't earn the right to scale.

So we stopped it.

I think failed experiments are worth publishing too.

What we were trying to build

One of the limitations we're interested in is how AI models remember and update information as they work through long sequences.

A conventional transformer uses attention to look back at earlier information.

We wanted to test a different approach: give most parts of the model a running state that gets updated as new information arrives, while still keeping some attention layers that can look across the whole sequence.

Think of it loosely as the difference between repeatedly looking back through your notes and carrying forward an evolving summary of what you've read. It's an analogy, not a claim that the model understands things the way a person does.

We hoped this combination might help a model keep track of changing information. It might also make long-running inference more efficient.

Those were hypotheses. We hadn't demonstrated either advantage.

Before testing the ambitious parts, we wanted to answer something simpler:

Could this design learn language about as well as the comparison models?

If it couldn't, we wanted a good reason before making it bigger.

So we built it

We built three small models with roughly 314 million parameters each. Parameters are the adjustable numbers a model learns during training.

The models differed in how they combined information across a sequence:

  • The global-attention control could look across the full sequence at every attention layer.
  • The local-attention control mostly looked at nearby information, with periodic layers that could look across the full sequence.
  • Our hybrid candidate, H1, replaced most of those local-attention layers with recurrent state, while retaining the periodic global-attention layers.

We kept their sizes approximately the same, used the same corpus and training recipe, and gave each model exactly 100 million tokens that contributed to its training loss. Tokens are the pieces of text a model processes.

We repeated the comparison three times from paired random starting points: nine training runs in total.

These weren't models we expected anyone to use. This was a screen designed to tell us whether an idea deserved a larger experiment.

That distinction matters.

It's easy in AI research to tell yourself that an idea just needs a bigger model. Or more data. Or another week of tuning.

If you do that every time, a small experiment stops being a useful gate.

We wanted H1 to earn the right to get bigger.

It didn't

On the basic test of how well the models learned language, our new architecture lost to both controls.

Not once. All three times.

We measured negative log-likelihood, or NLL: a loss that reflects how well a model predicts text. Lower is better. It is not a percentage score for intelligence.

The local-attention model was the strongest comparison. Against it, H1's loss was about 1.1% to 1.4% higher across the three runs.

Paired run Global-attention control Local-attention control H1 hybrid H1's extra loss over local control
1 4.8436 4.8083 4.8603 +0.0520
2 4.8302 4.7829 4.8520 +0.0691
3 4.8441 4.7972 4.8610 +0.0638
Mean 4.8393 4.7961 4.8578 +0.0617

NLL in nats per token, after the same 100-million-token training dose. Lower is better; values are rounded.

For the scaling decision, we froze a practical tolerance of +0.030 nats per token over the local-attention control. H1 exceeded that tolerance in every run. Its average penalty was roughly twice the allowed amount.

We chose that tolerance after seeing the early comparison results, not before the experiment began. We then fixed the analysis plan before opening a separate set of reserved evaluation results.

Three runs are not a broad statistical verdict on an architecture family. They were enough for our practical decision about this experiment.

H1 wasn't going to get more compute.

Another part of the experiment failed too

We also designed a test around the reason we were interested in the architecture in the first place: could it keep track of important information as that information changed?

Unfortunately, we learned almost nothing about that question.

All nine checkpoints produced zero valid answers out of 216 tasks. They exhausted the 48-token answer limit without satisfying the structured answer format the test required.

That made us suspicious of our own infrastructure. So we investigated a checkpoint from the local-attention control and one from H1.

For those two checkpoints, the diagnostic verified that the trained weights were actually loaded and that the scoring machinery worked. It did not isolate exactly why the models failed to produce valid answers.

The result was consistent with a format failure. We could not turn it into a conclusion about whether H1 was better or worse at tracking changing information.

That's a failure in our measurement process too. We hadn't checked early enough that the test could distinguish between models at this stage of training.

Next time, that feasibility check happens first.

There are a hundred reasons we could keep going

We could say the new architecture needs more training.

Maybe.

We could say it needs a different learning rate.

Maybe.

All three models shared one training recipe. We hadn't independently tuned it for H1, so this experiment cannot cleanly separate a limitation of the design from a mismatch between the design and that recipe.

We could change the mix of layers. Train for five times as long. Build a billion-parameter version. Argue that the advantages only appear at a larger scale.

Any of those things could matter.

This experiment does not prove that recurrent architectures don't work. The recurrent method we used, called KDA, is not disproven by this result. Nor does it prove that another hybrid couldn't work.

We also did not measure an efficiency advantage. Potential savings in speed or memory cannot count as a result we obtained.

The conclusion is narrower:

The version we built, under this training recipe and at this scale, was worse at basic language modelling and had not demonstrated the advantage it was supposed to provide.

That wasn't enough evidence to spend dramatically more compute on it.

So we closed the experiment.

That's what the gate was for

Novelty isn't evidence. Just because we designed something ourselves doesn't mean we should become attached to it.

The next architecture will face a clearer progression:

  1. Prove that the implementation works in a tiny functional test.
  2. Run a small language-quality comparison, with a tolerance and a limited tuning budget declared beforehand.
  3. Test the specific advantage the architecture claims to have, using a measurement that works at that scale.
  4. Repeat the result from multiple starting points.
  5. Only then consider substantially more training.

We're trying to make being wrong cheap.

Eventually we want to run much larger architecture experiments. The cost of fooling ourselves goes up quickly with scale.

The work wasn't wasted

H1 itself is done. A lot of what we built for it isn't.

The experiment forced us to improve how we give models exactly the same amount of training, verify that an evaluation loads the intended checkpoint, and check that the GPU environment is using the required low-level library.

We also built machinery for recurrent models and improved how we fix an analysis plan before opening results reserved for evaluation.

The extraction plan separates that engineering from H1 so it can become reusable Vinci infrastructure. That is work to carry forward, not a claim that every piece has already been integrated elsewhere.

The next experiment should benefit from it without inheriting an obligation to rescue this one.

What it cost

The experiment's final accounting records 45.35 H200 GPU-hours, or about 45 hours of work on one H200 GPU in total.

That includes training, evaluation, diagnostics, qualification work, and infrastructure losses. It is compute consumption, not a measurement of the architecture's inference efficiency.

Forty-five GPU-hours isn't nothing. But I'd rather use a small experiment to make a stopping decision than carry a negative result into a much larger run without new evidence.

That's the point of the gate.

What's next

Vinci's architecture research remains active. H1 is closed.

There won't be an “H2” that's just the same idea with enough knobs changed to justify running it again. A successor needs a materially different hypothesis for why it avoids the penalty we measured, and a cheap experiment that could prove it wrong.

We're still a very small lab. We're going to get things wrong.

If we're serious about building in public, that should include some of the experiments that never become model releases.

H1 didn't earn the right to scale.

We learned what we could from it.

Now we build the next thing.

Cover: AI-generated conceptual illustration, not an image of the experiment or a technical diagram.

George Pu

George Pu is the founder and CEO of SimpleDirect, an independent Canadian AI lab developing research, models, technologies, and products under Vinci.

SimpleDirect® is an independent Canadian AI lab in Toronto. Vinci names the AI research, models, technologies, and products developed by SimpleDirect.

Share