A recurrent language model whose state is moved by Householder reflections. Open architecture research, built in India.
Website · What the research found · Why it is built this way · Results · Quickstart · Paper A · Paper B · Evidence · Status
Each layer keeps one small vector state per head. Every token moves that state with two input-dependent Householder transforms, then a gate mixes in new content:
h_t = (1 − g_t) · H_2 H_1 h_{t−1} + g_t c_t H_i = I − β_i u_i u_iᵀ, β_i = 1 − cos θ_i ∈ [0, 2]
The state has a fixed size, so writing a token costs the same whether the context is ten characters or ten million. The 25M model carries 11,904 numbers of state in total.
Independent open research, built in India.
What can a fixed-size recurrent state actually represent, and what does it do when it cannot?
A vector state carries order and position well, and this project has group-theoretic results saying exactly how much. It does not carry many independent facts. A matrix memory read by contraction does — a faithful DeltaNet solves the recall task at 1.000 on three seeds where this recurrence sits at 0.32 — so it is now part of the design rather than a planned extension. Whether a wide vector with proper binding would do as well is untested and openly recorded as such.
Two technical reports. Paper A is the architecture; Paper B is what the same transport does when it fails, which turned out to be the more interesting half.
Failure is quantised. Train this transport on the word problem of a finite group and it either
solves it exactly at any length, or it lands on 1/|N| for a normal subgroup N — it has learned
the quotient G/N, gets the coset right, and guesses inside it. Across Q₈, S₄ and A₅, twelve runs
land on a rung of their own group, worst deviation 0.021. The lattice is a property of the group,
and the model never sees it (evidence/word_problem.py).
The errors name the subgroup, not just its size. Q₈ has three different normal subgroups of
order 4, so accuracy alone cannot tell them apart — all three give the same 0.250. Coset consistency
picks exactly one, and scores the other two at the predicted 0.5
(evidence/which_subgroup.py).
Binding capacity is a character. For an orthogonal map, the expected overlap of a vector with
its image is the normalised trace, E[v·gv] = tr(g)/d. A single Householder reflection has
tr = d−2 — the least hiding non-trivial element of O(d) — and k of them give overlap
(1−2/d)^k. That is why one reflection cannot bind, and why it gets worse as heads get wider
(evidence/trace_law.py).
...but more reflections do not buy this architecture a memory — we checked, and it fails
backwards. The obvious consequence was that raising n_h should turn tracking heads into storage
heads. Swept in a trained model, the trace follows the law closely (0.818, 0.611, 0.364, 0.122 at
n_h = 2, 4, 8, 16) while recall falls monotonically, 0.113 → 0.004, ending below chance
(evidence/nh_sweep.py). A bind/unbind store recovers a value by applying
its key's reflections a second time, and assumes the stored pair is untouched in between. A
recurrence does neither: it never applies a query's inverse, and every later token transports the
whole state. More reflections are therefore a faster scrambler, not more capacity. The trace law
bounds what a product of reflections could store; it does not say this recurrence can reach it.
The recall wall is the read, and it took a baseline we had never run to see it. Associative
recall held at ~0.32 for four key–value pairs across 170 runs and 21 interventions, every one of
which changed the transport, the write or the training order. A 290,370-parameter transformer
solves the identical task at 1.000, so the task was solvable all along and the wall was ours
(evidence/baseline_transformer.py). Replacing the read with a
contraction S q does solve it: a faithful DeltaNet scores 1.000 on three seeds
(evidence/delta_reference.py, runs/dref-*). That is consistent
with the common-factor derivation below, and it is consistency, not
proof — the alternative, a wide vector with proper binding, has never been run at the same standard,
so "the read is the wall" is the best available reading rather than a demonstrated result.
An earlier version of this section stated it as demonstrated, citing a comparison table that had no
run behind it. Those figures are retracted; see evidence/CLAIMS.md #25. The
surrounding claims — a 290,370-parameter transformer at 1.000, this recurrence at ~0.32, seven
falsified counting arguments, five constructions that predicted ~0.97 and trained at 0.25–0.31 —
are all backed by logged runs and stand. The standing conclusion that constructions here have no
demonstrated predictive value for what gradient descent finds stands with them, and #25 is the
sharpest example of why.
A trained transport can be read as a representation. Its character norm ⟨χ,χ⟩, computed from
traces alone with no group table and no labels, identifies which quotient it learned. Over sixteen
seeds it agrees with the independent error-based reading on all seven runs that are genuine
homomorphisms (evidence/character_table.py).
Every derived number in Paper B — the bounds, the rung sets, the attainable character norms — is
recomputed from the group definitions by
evidence/verify_paperB.py, which also re-reads the measured tables
from the raw logs rather than trusting a transcription.
Every choice below exists for a stated reason, and most were settled by a small experiment in
evidence/ before the language model was trained. paper/main.pdf
has the full arguments.
Recurrent, not attention. A transformer re-reads a cache that grows with the context. A recurrent state does not grow, so generation cost and memory per token are constant. The question a recurrent design has to answer is what that fixed state can represent; the choices below are that answer.
How much that is worth, stated honestly. KV-cache compression has moved a long way. DeepSeek-V4.1
-Flash reports 890 bytes per token of global cache — 437× below DeepSeek-V1 — and single-token
decode FLOPs that rise by only about a quarter when context is extended 256-fold from 4K to 1M. So
the gap is not the four orders of magnitude a naive uncompressed-cache calculation suggests. What
remains is a difference in kind rather than degree: their footprint still grows linearly with
context and ours does not grow at all. At a million tokens that is a few hundred megabytes against
a few tens of kilobytes — unremarkable in a datacentre, impossible on a phone. This project's
efficiency case therefore rests on the on-device and memory-constrained regime, and not on the
claim that long-context serving is otherwise unaffordable
(results/literature_kv_compression.txt).
Reflections, not a diagonal recurrence. Diagonal (element-wise) transitions commute, so a pure
diagonal transport ends in the same state for every ordering of the same tokens. It can only
compute functions of the multiset of inputs. On the word problem of S₃ that caps accuracy at a
provable 0.385; two commuting transports sit on that bound (0.388, 0.389) while a non-commuting one
goes above it (0.779) (evidence/ceiling.py,
results/v2.txt). Two reflections about different mirrors do not
commute, and each costs only a dot product and a vector update.
β = 1 − cos θ, not a sigmoid and not a fixed 2. The transform scales the component along u
by 1 − β = cos θ. This form keeps β in [0, 2] for every input, reaches the exact reflection β = 2
(a sigmoid only approaches it), and has zero slope there, so a reflection is a stable place to
sit. β stays learnable because a fixed reflection can never forget along a direction and fixes the
sign of the determinant for every token (evidence/parity.py). In a toy
sweep the sigmoid solved 0 of 16 state-tracking runs and the cosine form 9 of 16
(evidence/beta_activation.py). The trained 25M model does not
use the whole range, which an earlier version of this section claimed: β sits at a pure reflection
in every layer (medians 2.000, 1.998, 1.998, 1.988, nothing anywhere near the singular β = 1), and
the learned input dependence varies β close to 2 rather than across the range. What does change with
depth is the gate — retention 1/g runs 11–17 tokens in the first two layers and reaches 86 and 132
in layer 2 (results/trained_geometry.txt).
A convex, scalar gate. Each transform has spectral norm at most one and the write is a convex
mix, so the state is bounded by its inputs for every sequence (a one-line proof in the paper). The
gate is one scalar per head because a scalar commutes with the reflections, and that is what makes
this kernel's derivation work (evidence/chunkwise.py).
An earlier version said flatly that "a per-dimension gate breaks it." That is too strong.
Channel-wise gating is compatible with an exact chunk-parallel algorithm in general — KDA does it
via a diagonal-plus-low-rank formulation, and the matrix memory in this repo now does it too, with
a decay-weighted Gram G = K̂K̃ᵀ and log-space centring, checked exact to 8e-15. Whether the
level-1 Householder kernel admits the same treatment is untested; note that with a per-channel
gate its transition (I−D)·H₂H₁ is itself diagonal-plus-low-rank, which is precisely the structure
KDA handles. The gate is capped at 0.9 so the kernel's
rescaling factors stay below e^8.1.
An exact parallel kernel, not an approximation. Training processes chunks of 8 tokens with one
unit lower-triangular solve per chunk and no matrix inverse. It matches the step-by-step
recurrence to float precision, and tests/ check that on every run.
A small vector state for tracking — and now a matrix for storage. A vector per head is cheap
and suits tracking where a sequence is. At n_h = 2 it holds almost no independent facts. An
earlier version of this section argued the fix was more reflections — n_h ≈ d_h/2, the same
mechanism at a different setting rather than a second memory. That was wrong, and it
contradicted the paragraph above, which had already measured recall falling backwards as n_h
rises. The real cause is narrower than either account. Unrolling the recurrence gives
h_T = a_T · T_T · [ Σ_t M_{t→T−1} b_t ] + b_T
so the query's transport is a common factor over every stored item: it reorients all of them
together and cannot pick one out of a sum. Selection is left to a diagonal output gate, which has no
item axis. A contraction read S q = Σ_i v_i (k_i · q) has one. At equal state size on the same
task: a faithful DeltaNet — matched component by component against the reference implementation —
solves it at 1.000 on three seeds, where the shipped vector with a diagonal gate sits at 0.32
(evidence/delta_reference.py, runs/dref-*).
Retraction. This passage previously cited 0.984 against 0.32 from results/matrix_decision.txt,
along with a matrix-versus-vector comparison, a crossover table and a head-structure table. None of
them had a run behind them; the logged runs of that experiment scored 0.133–0.164. The direction
survived and has been re-earned from new logged runs, but the comparison has not: the vector arm
has never been run with the same care, so "the matrix beats a wide vector" is an open question here,
not a result. Saryu is adopting the matrix
state with a targeted rank-one erase — a known design that DeltaNet and the models built on it
already ship, not something invented past it. The Householder-product transition stays, and that is
the part that is ours.
The block around the transport. Normalisation, a width-4 causal convolution, per-head RMSNorm, a SiLU output gate and a SwiGLU feed-forward surround the recurrence. This surrounding anatomy matters a lot: a GRU placed in the same block does not merely come close to Saryu at 5M, it ties it on final loss (table below). So comparisons keep the block fixed and change only the transport -- and the transport has to be argued for on parallel trainability rather than on quality, because on quality it is a draw.
Small, cheap and falsifiable first. Everything so far runs on one T4 GPU or a laptop CPU, on
character-level enwik8, so each design question is settled in hours before anything is scaled.
Predictions are committed before a run and reported whichever way they come out; the same-order
control in evidence/same_order.py is recorded as falsified, because one
of its 28 scored failures missed the registered tolerance, and it says so.
evidence/CLAIMS.md is the ledger of every claim this project has made — 22
so far, 13 of them retracted — with what caught each one. A test that decided nothing is recorded as
void rather than negative, so it is not later miscounted as evidence.
Bits per character on the last 128 characters of 64 evaluation windows. Every window ends at the same position, so each context length scores identical text with more or less of it in front. Every model was trained at context 128. Lower is better.
| model | params | train steps | ctx 128 | ctx 512 | ctx 2048 | ctx 8192 |
|---|---|---|---|---|---|---|
| Saryu 25M | 25.2M | 10,000 | 1.535 | 1.433 | 1.433 | 1.433 |
| Saryu 5M (released checkpoint, value embedding) | 5.0M | 5,000 | 1.671 | 1.586 | 1.586 | 1.586 |
Reading it honestly:
- The model reads about 256 characters and no more. Past that the recurrent state entering the
scored text is identical to six decimals in every layer and the predictions do not change
(
evidence/results/effective_context.txt). Measuring it directly — loss at the final position with only the lastkcharacters visible, 512 samples perk— saturation is at 256, not the 512 stated here earlier:k= 256 costs +0.0006 bits against full context, and 256 → 512 is worth nothing (results/trained_geometry.txt). Context 128 is worse only because the scored characters then have nothing in front of them. - An earlier version of this table claimed the model kept improving out to context 8192. That was an artefact of drawing window positions separately per length, so each column scored different text. These numbers score identical text at every length; they are better than the old ones everywhere, over a much shorter span than was claimed.
- Scale helps: 25M beats 5M by about 0.15 bpc at every length.
- These rows were marked superseded pending a rerun with several seeds. That rerun is below.
- These are small models on one corpus, with our own 95/5 character split, so they are not comparable with published enwik8 numbers and say nothing yet about LLM scale.
Done on 2026-09-22 (evidence/efficiency.py): four arms matched within
2% of 5M parameters, two seeds, loss against tokens seen rather than steps, enwik8 at
context 128.
| arm | tokens to bpc 3.0 | best bpc |
|---|---|---|
| Saryu (plain) | 204,800 | 2.176 / 2.189 |
| Saryu 3:1 hybrid | 204,800 | 2.176 / 2.293 |
| Saryu + matrix memory | 409,600 | 2.201 |
| GRU (pre-norm, residual) | 409,600 | 2.190 / 2.193 |
| transformer | 1,638,400 | 2.571 / 2.751 |
Three of the four registered predictions failed.
- Against a transformer the gap is large: ~8x fewer tokens, and it never reaches bpc 2.5.
- Against a GRU it is a tie on final loss. Saryu is ahead early and ahead on wall-clock, but a 1990s architecture given the same pre-norm and residuals lands in the same place.
- The matrix memory and gate lower bound cost 2x here. Both were built for recall capacity on MQAR, both work there, and on language modelling they are a penalty.
- The 3:1 hybrid shows no effect at context 128 -- a fault in the experiment, not the hybrid. At that length there is nothing to retrieve, so the harness cannot see what a hybrid is for.
The state-tracking result needs the same correction. Saryu solves the S3 word problem at 1.000
where a matched transformer reaches 0.727 -- but a GRU also solves it at 1.000
(evidence/state_tracking_gru.py). That is what the theory
predicts: fixed-depth attention is in TC0 and the word problem is NC1-complete, so the transformer
is the arm the argument is about; a GRU is a nonlinear recurrence and was never in that class. The
Householder and delta-rule literature is about recovering this ability in linear RNNs, which
train in parallel -- not about exceeding classical ones, which have it and cannot be parallelised.
What is actually distinct is the parallel kernel, and that is measured. Step time against
sequence length at constant tokens per step, against a parameter-matched GRU
(evidence/parallel_advantage.py):
| context | 128 | 256 | 512 | 1024 | 2048 |
|---|---|---|---|---|---|
| GRU / Saryu step time | 1.36x | 1.80x | 2.28x | 2.95x | 3.70x |
Monotone, and it had a live falsifier: nn.GRU dispatches to a fused C++ kernel while our chunk
path is a Python loop over many small operations, and the same measurement against a transformer
went the other way. The honest position is GRU-quality, parallel-trainable -- narrower than
what this page previously implied, and supported on both halves.
Requires Python 3.10+.
pip install -r requirements.txt
gh release download v0.1 --repo varun29ankuS/Saryu-RNN -D checkpoints
# or without the GitHub CLI:
mkdir -p checkpoints && for f in saryu_25m.pt saryu_v4b_last.pt; do curl -L -o checkpoints/$f https://github.com/varun29ankuS/Saryu-RNN/releases/download/v0.1/$f; done
python -m pytest tests
python scripts/talk.py "The history of India begins with "
python scripts/serve.py # web UI at http://127.0.0.1:8471 that shows the state while it writes
Training needs the corpus:
mkdir -p corpus && curl -L http://mattmahoney.net/dc/enwik8.zip -o corpus/enwik8.zip && unzip corpus/enwik8.zip -d corpus
CHARS=400000 STEPS=5 TARGET_PARAMS=300000 EVAL_LENS=128 SAVE= python scripts/train.py # smoke run
python scripts/train.py # the 25M recipe
saryu/model.py reproduces the original training code's logits bit-for-bit on both checkpoints.
saryu/model.py the model: kernel, block, LM, checkpoint loader
scripts/train.py the 25M training recipe
scripts/talk.py text completion from the 5M checkpoint
scripts/serve.py streaming web UI that shows the state while it writes (--selftest)
tests/ kernel exactness, norm preservation, order sensitivity, checkpoint loading
evidence/ the small experiments behind each design choice, with their result logs
word_problem.py group word problems: the rung structure of trained transports
which_subgroup.py Q_8, where three normal subgroups share one rung and the errors must choose
character_table.py reading a trained transport as a representation, from traces alone
trace_law.py binding capacity is a character: overlap ~ (1-2/d)^k
verify_paperB.py recomputes every derived number in paper B from the group definitions
baseline_transformer.py a matched transformer on the recall task: the control 170 runs lacked
cell_shootout.py matrix-with-contraction vs wide-vector-with-unbinding, at equal state
CLAIMS.md every claim made here, 13 of 22 retracted, and what caught each
experimental/memory/ the matrix-memory extension (see its README)
paper/main.pdf report A: the architecture, the kernel, the trained models
paper/paperB.pdf report B: what the transport does when it cannot solve a group
checkpoints/ saryu_25m.pt, saryu_v4b_last.pt (release v0.1, not in git)
corpus/ enwik8 (not in git)
- Works: the reflection recurrence at 5M and 25M parameters, the exact kernel, the demos, and the group-theoretic results in Paper B.
- Next: a matrix state with a contracting read inside the existing block, keeping the Householder-product transition and the chunk-parallel kernel. Then matched comparisons against other recurrent and attention models at larger scale, with several seeds and a second corpus — the single biggest gap, and the reason no competitive claim is made here.
- Open: a hysteretic write gate, recorded void rather than negative — both arms scored below
chance because the cell was deliberately given the weak diagonal read, so there was no working
baseline for an improvement to appear against
(
results/matrix_decision.txt). - Retired: the heterogeneous layer at
n_h ≈ d_h/2, falsified by the sweep and by the derivation above. A fused GPU kernel is also no longer next: choosing the chunk size properly already buys 1.1–5.4×, which drops the recurrence to 17–53% of a step, and a hand-written kernel would be competing for the remainder without a fused backward (results/kernel_profile.txt). - Known limits: the trained models stop using context at about 256 characters; two of the four rows of Paper B's character table have no qualifying run behind them; A₅'s rung coincides with chance, so those runs confirm the lattice law without demonstrating it.
@misc{sharma2026saryu,
author = {Varun Sharma},
title = {Saryu: a recurrent language model whose state is carried by Householder reflections},
year = {2026},
note = {Technical report A, draft v1},
url = {https://github.com/varun29ankuS/Saryu-RNN}
}
@misc{sharma2026quantised,
author = {Varun Sharma},
title = {Quantised failure: trained reflection transports collapse onto exact group quotients},
year = {2026},
note = {Technical report B, draft v1},
url = {https://github.com/varun29ankuS/Saryu-RNN}
}Apache License 2.0; see LICENSE. The weights in the v0.1 release are under the same
license.
- Grazzi et al. 2024. Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues. arXiv 2411.12537.
- Siems et al. 2025. DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products. arXiv 2502.10297.
- Yang et al. 2024a. Parallelizing Linear Transformers with the Delta Rule over Sequence Length. arXiv 2406.06484.
- Marathe et al. 2026. Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models. arXiv 2609.12303.