← All projects
in progress Started Sep 1, 2026

Small Model Research

An ongoing program asking how far small, cheap models can be pushed under one fixed testing protocol. Findings will be published here as they land.

BenchmarksArchitecturesLocal AI

This is the home base for our core research program. It is a living document: as findings firm up, they get published here, and the posts get linked.

The question

The industry puts all its money and attention into models that cost hundreds of millions to train. The question we are asking is quieter: how much of what people actually use AI for really needs that scale… and how much of it can a small model running on ordinary hardware handle just as well?

Nobody answers this honestly at large scale, because the answer is expensive to publish. So we answer it at small scale, where the experiments are cheap enough to run properly and repeat.

The method

A few rules keep the work honest:

  • One protocol. Every model and architecture is compared on the same data, the same number of training steps, and the same schedule. The harness controls everything except the thing under test.
  • Hypothesis first. Every run states what it expects to find before it trains. No post-hoc storytelling.
  • Everything is journaled. Each run logs its full curve, config, and provenance. Results are reproducible or they do not count.
  • Null results count. A failure list is part of the paper. Most experiments do not work, and we report that.

Findings

Nothing publishable yet. The first phase of work is done: a fixed testing protocol is built, and around eighteen architectures have run through it … recurrent trunks, attention hybrids, state-space variants, looped and slimmable designs, each journaled against the same baselines. Most of what we tried did not win, which is exactly what the protocol exists to record.

What the runs did produce is a leader: a hybrid recurrent recipe that beat every architecture we tested under the protocol. The program is now testing that leader at the one-billion-parameter scale … a step up from the small models the protocol was built around, and the first real check of whether what we learned at small scale survives growth. That campaign is running now, and its findings will land here.

Check back, or read the blog for the thinking as it develops.

Discuss a project like this →