On Sub-Billion Models and Why I Refuse to Train Monoliths
Rethinking sparsity, Grouped-Subspace Latent Attention, and running intelligence on local NPU silicon.
There is an unspoken sickness in contemporary machine learning: the belief that intelligence is purely a function of linear parameter scaling. If your model fails at multi-step deductive routing, the industry answer is almost comical — throw another thirty billion parameters into the feed-forward layers, rent eight more racks in Ohio, and bill the venture fund. I reject this premise entirely.
I. The Monolith Trap
When you train a monolithic dense transformer, every single token pays the exact same thermodynamic tax. Whether the model is predicting the next letter in an arithmetic proof or outputting a generic punctuation mark, all layers activate simultaneously. This is not how biological synapses operate, nor is it how economical computing survives.
In my experiments with small language models (SLMs) in the 250M to 750M range, I noticed something remarkable: when you constrain capacity, the network is forced to learn structural representations rather than surface-level token memorization. It stops being a search engine in disguise and starts acting like an algorithmic engine.
“Constraints are not limitations; they are the only reason elegance exists in engineering.”
II. Grouped-Subspace Latent Attention
The primary bottleneck on consumer chips isn't FLOPs — it is the memory bus. During autoregressive decoding, loading the KV-cache across memory channels throttles the processing units. Standard Multi-Head Attention is recklessly wasteful here.
By decomposing the attention projections into grouped subspaces with low-rank factorization, we can compress the active memory footprint down to a fraction of standard MHA. The query vectors only look into specific latent subspaces conditioned on the routing token.
[Token In] ──► [Sparse Subspace Projector]
│
┌───────────┼───────────┐
▼ ▼ ▼
Subspace 0 Subspace 1 Subspace N
│ │ │
└───────────┼───────────┘
▼
[Orthogonal Recomposition] ──► [Next Layer]class SubspaceLatentRouter(nn.Module):
def __init__(self, d_model=512, n_subspaces=8, rank=32):
super().__init__()
self.n_subspaces = n_subspaces
self.down_proj = nn.Linear(d_model, n_subspaces * rank, bias=False)
self.up_proj = nn.Linear(rank, d_model, bias=False)
self.gate = nn.Parameter(torch.randn(n_subspaces))
def forward(self, x):
# Route dynamically without dense activation penalties
b, s, d = x.shape
subspaces = self.down_proj(x).view(b, s, self.n_subspaces, -1)
weights = F.softmax(self.gate, dim=-1)
routed = torch.einsum('bsnr,n->bsr', subspaces, weights)
return self.up_proj(routed)III. Local NPU or Bust
I want my AI agents to live on my desk, not in a server room in North Virginia. When an agent runs locally — wired to a local serial port, monitoring system logs, listening to an audio buffer — the latency drops from 450 milliseconds down to 14 milliseconds.
At 14ms, interaction ceases to be an RPC request and becomes a continuous biological feedback loop. You feel the machine reacting before your fingertips have fully lifted from the mechanical switches.
IV. The Virtue of Constraints
We have traded intellectual discipline for brute force compute. Building Small Language Models forces you to inspect every layer, question every tensor allocation, and optimize every activation function.
This isn't about nostalgia for slow hardware. It's about self-reliance. An intelligence that fits on a USB drive and runs on battery power is sovereign; an intelligence tethered to a subscription token is just an API lease.