01 The Big Picture
The transformer is not one idea. It is a stack of assumptions bundled together in 2017 — and each newcomer in this doc is an attack on exactly one of them.
Doc 02 built the baseline: an autoregressive model that generates one token at a time, each token attending to every previous token through softmax attention, with the KV cache as its memory. Doc 21 showed the incumbent evolving inside that frame — MoE sparsity, long context, better caches. This doc is the outside view: architectures that don't tune the transformer but replace parts of its skeleton.
The assumptions under attack, one per family:
| Transformer assumption | Who drops it | What they gain |
|---|---|---|
| Attention must be O(s²) and exact over all history | Mamba / SSMs, RWKV, Gated DeltaNet | O(s) time, fixed-size state |
| Every layer must be attention | Jamba, Zamba, Griffin/Hawk hybrids | Recall where needed, speed everywhere else |
| Weights need 16 bits | BitNet b1.58 | ~9× less memory per weight |
| Text must be chopped into tokens first | Byte Latent Transformer | No tokenizer, no token tax |
| Output must be emitted left-to-right | Diffusion LMs (LLaDA, Mercury) | Many tokens per step (bridges doc 24) |
| The model must generate text to answer | JEPA family, JEV decision model | Latents or typed decisions — no decoding loop at all |
02 What — The Taxonomy Up Front
Here is the whole landscape on one card. Generative quality is a judgment about frontier-class text benchmarks, not a moral ranking.
| Architecture | Time / token | State size | Generative quality | Representative models |
|---|---|---|---|---|
| Softmax attention (baseline) | O(s²) train, O(s) per new token | KV cache — grows O(s) | ★ reference | GPT, Llama, Claude, Gemini |
| State-space model (selective) | O(s) | Fixed — one h per layer | Near-parity, weaker recall | Mamba, Mamba-2, Codestral Mamba |
| Attention–SSM hybrid | O(s²) + O(s) mixed | Grows only at attention layers | Parity (best of both) | Jamba, Zamba, Griffin/Hawk |
| Linear attention / DeltaNet | O(s) | Fixed matrix S (d×d) | Good LM, weak multi-query recall | Gated DeltaNet, Kimi Linear |
| Data-dependent linear RNN | O(s) | Fixed per-token state | Good LM, weak recall | RWKV v5/v6 |
| Exponential-gate LSTM | O(s) | Fixed cell + matrix memory | Competitive at small–mid scale | xLSTM (sLSTM/mLSTM) |
| Ternary-weight LM | Same O(s²) — cheaper ops | Same KV cache | Parity at matched scale | BitNet b1.58, b1.58 2B4T |
| Byte-level patching LM | O(bytes + patches²) | KV over patches | Parity at large scale | Byte Latent Transformer |
| Diffusion LM | O(s²) per refinement pass | Full-sequence activations | Closing gap; weaker long-form | LLaDA, Mercury (Inception), Grok-fast style |
| Joint-embedding predictive | O(context) — no decode | Latent representation | N/A — not generative | I-JEPA, V-JEPA (LeCun lineage) |
| Decision model (typed readout) | O(S) prefill, O(1) per question | Shared-context KV, reused | N/A — calibrated decisions, not text | JEV (TypeSafe) |
| Memory-augmented transformer | O(s²) + NN lookup | Weights + product-key store | Parity + long-tail facts | Product-key memory layers, Titan |
MoE is deliberately absent: it changes which parameters fire, not the sequence math, and got its full survey in doc 21. The rows above change the sequence math itself.
03 Why These Arise Now
04 How — Computing a 4-Token Answer, Five Ways
Same task for every architecture: read 4 context tokens, emit a 4-token answer. Watch where the state lives at each step.
The visual tells the whole story: attention replays the transcript (state grows), SSM and linear attention keep a fixed notebook (state compresses), diffusion works on the whole answer at once, and the decision model never generates at all. Compression is the theme — and every compression is a bet about what the task won't need later.
05 The Time-Complexity Table
For s sequence length, w window size, d state dimension, n generated tokens:
| Regime | Training | Decode step | Total for n tokens | Where state lives |
|---|---|---|---|---|
| Full attention | O(s²·d) | O(s·d) — read whole KV | O(n·s·d) | KV cache (grows) |
| Sliding window | O(s·w·d) | O(w·d) | O(n·w·d) | window KV (capped) |
| SSM / linear attention | O(s·d²) | O(d²) — one state update | O(n·d²) | fixed h or S |
| Decision readout (JEV) | one O(S) prefill | ~O(1) per question branch | ~O(S) for Q questions | shared KV, reused |
Read the last row against row one: Q questions over S context tokens cost ~Q·S in decode-heavy form versus ~S with a single shared prefill. That asymmetry — encode once, branch many — is the entire JEV pitch, and it's the same asymmetry that makes prefix caching cheap for chat.
06 The Families, In Turn
6.1 State-Space Models — Mamba
A linear recurrence that replaces attention's global lookup with a compressed rolling state. The selective trick: the transition matrix depends on the input, so the model can choose per-token what to remember and what to forget.
Δ is the gate. Large Δ → Ā ≈ 0 → previous state erased, current token written sharply (a "reset" on a new section or a fact worth anchoring). Small Δ → state persists (drift over filler). Compression becomes inference: deciding what h keeps is a learned judgment call about the future, made at read time.
Mamba-2's SSD duality shows the recurrence can be rewritten as a form of (masked, decayed) linear attention — the two families are two views of one algebra. That matters because it means hybrid layers (next subsection) are mixing siblings, not strangers.
Why it may win
Linear time, constant memory per layer, real length extrapolation, throughput that crushes attention on long sequences. On language modeling perplexity it matches transformers at matched scale.
Why it won't (yet)
The multi-query recall anchor problem: tasks like MQAR — "what were the exact values associated with keys K₁, K₃, K₇?" — require retrieving specific past tokens, and a fixed-size h has thrown away exactly that precision. Compression destroys addressable memory.
6.2 Attention–SSM Hybrids — Jamba, Zamba, Griffin/Hawk
The engineering compromise: interleave. A few full-attention layers (exact recall, exact copy) among many SSM/linear layers (cheap drift), so recall anchors exist without paying O(s²) everywhere.
Why hybrids exist — the per-layer tradeoff: each layer independently chooses between compute/memory spend (attention: O(s) state, exact recall) and bandwidth thrift (SSM: fixed state, lossy recall). Recall behavior is task-level, but the cost is layer-level — so spend attention only where the representation actually needs addressable memory. Jamba (AI21), Zamba (Zyphra), and Griffin/Hawk (DeepMind) all landed near-attention quality at SSM-class cost on long context.
6.3 RWKV v5/v6
A linear-attention RNN wearing an LSTM's coat, famous for community training at real scale. v5/v6 make the recurrence data-dependent (like Mamba's Δ):
The bytes-at-runtime framing: a deployed RWKV runs as a tiny stateful kernel — weights + one small state vector per layer. No KV cache file, no context-length allocator. That makes it the natural architecture for on-device and edge inference, where the KV cache is the memory problem (doc 22).
6.4 Gated DeltaNet & the Linear-Attention Family
Linear attention (Katharopoulos et al.) replaced softmax(q·kᵀ)·v with φ(q)·(Σ φ(k)ᵀ v) — associative, O(s), but a sum that only grows. DeltaNet adds a delta rule: retrieve-then-correct the state, so newer keys can overwrite stale ones instead of just adding:
The MQAR lesson: associative recall (multi-query associative recall) separates from perplexity. A linear-attention model can match a transformer's next-token loss while failing MQAR, because perplexity is dominated by local, drift-y prediction where a compressed state suffices — while MQAR demands exact key→value lookup that compression can't preserve. Benchmark on recall, not just perplexity, when evaluating any fixed-state model.
6.5 xLSTM — Exponential Gating, LSTM Form
Beck et al. (2024) asked: what if the LSTM's problem was never gating, just capacity and parallelism? xLSTM recovers exponential memory decay (like Mamba's Ā = exp(Δ·A)) inside LSTM form: sLSTM adds scalar exponential gates with new mixing, mLSTM swaps the scalar cell for a matrix memory cell with a covariance-style update — effectively a linear-attention state with LSTM-style gating, fully parallelizable.
6.6 BitNet b1.58 — The Ternary Bet
Not a sequence-model change — an arithmetic change to the whole stack. Every weight is forced to −1, 0, or +1; activations stay INT8:
Why it matters for a hardware-minded reader: a matmul over {−1,0,1} weights has no multiplies — it's adds and sign flips — so the FLOP cost per weight drops ~9× versus FP16, and so does the memory traffic that doc 10 identified as the decode bottleneck. Training needs tricks (this is the hard part): quantization-aware training from scratch, and spot precision — keeping critical accumulation steps in higher precision so gradients stay stable through the rounding.
6.7 Byte-Level Models — Byte Latent Transformer
BLT (Meta, 2024) deletes the tokenizer. Raw bytes flow in; a small local model groups them into patches whose size tracks local entropy; a big global transformer attends over patches; a local decoder expands back to bytes.
The trade: compute per byte moves up front — you pay a small-model pass over every byte before the big model sees anything. But you gain: no OOV, no tokenization artifacts in code/math, no "why is my French prompt 30% more expensive in tokens" tax. At scale, BLT matched Llama-3 tokenizer quality at better byte-level compute parity.
6.8 Diffusion LMs — LLaDA, Mercury/Grok-Fast Style
Autoregressive models factor text as a chain: p ∝ Πₜ p(xₜ | x<t) — one token at a time, forever. Diffusion LMs factor it as a denoising problem: corrupt the whole sequence with mask noise, learn to restore it, refine all positions in parallel:
This is the bridge to doc 24: speculative decoding and multi-token heads try to squeeze parallelism into an autoregressive loop; diffusion LMs are born parallel — several refinement rounds emit a whole block. Mercury (Inception Labs) reportedly runs 5–10× faster than size-matched autoregressive models precisely because each forward pass advances many tokens. The cost: refinement rounds re-run full attention over the whole sequence (wasted compute on already-finalized tokens), and quality on long, tightly-ordered text still trails the best autoregressive models.
6.9 JEPA — Predict Latents, Not Tokens
The LeCun-lineage answer to a different question: what if generation is the wrong objective? Joint-Embedding Predictive Architectures (I-JEPA on images, V-JEPA on video) predict abstract representations of missing content, never pixels or tokens:
By predicting in representation space, the model is never forced to spend capacity modeling unpredictable detail — the exact noise that makes pixel/token generators hallucinate texture. Why it may matter for LLMs: "token detokenization" — the waste of modeling surface word choice rather than meaning — is plausibly the same failure; a JEPA-style LLM would score candidate thoughts, not spellings. Why it's not here: without a generative head you can't sample text from it — the interface itself has to change, which is a product problem, not just a research one.
6.10 JEV (TypeSafe) — The Decision Model ★
The newest and strangest entry — a typed-decision architecture. Instead of generating an explanation ending in "Yes", it outputs the decision itself with a calibrated probability: Yes 80% / No 20% — read directly from hidden states over the allowed answers. There is no autoregressive decoding loop at all.
Honest framing: this is a new, poorly documented design — an illness of an ecosystem where closed internals force black-box reconstruction from ~10K API calls. Everything below is a public-article reconstruction; treat "reportedly" as attached to every claim.
Calibration as the training target. Reportedly trained with RLCD (Reinforcement Learning for Calibrated Decisions) — plausibly a log-loss or Brier-style objective — so that a stated "80%" event fires about 80% of the time. Reported MMLU calibration error ~0.031, with most predictions concentrated at high confidence. A transformer's sampled text has no such guarantee: "I'm 80% sure" in prose is theater; a calibrated readout is a contract.
Expected-cost routing — why agents care. Given a decision with probability p of the bad branch, escalate-to-human (or to a bigger model) when the expected cost of trusting it exceeds the escape cost:
With calibrated p, routing thresholds become arithmetic instead of vibes. A fleet of agent decisions — approve, retry, escalate — becomes a budget line you can actually compute. This connects directly to the doc-21 theme: spend expensive compute only where the expected cost justifies it.
6.11 Three Footnotes That Could Become Chapters
Multi-token heads
Doc 02's trick: parallel output heads predicting tokens t+1…t+k in one pass. Cheap 2–3× decode speedup; full survey of the strategy space in doc 24.
Memory layers
Product-key memory: a huge parameter store consulted by nearest-neighbor lookup — a few keys in, one value out. Adds billions of "lookup parameters" without quadratic attention; recent work shows it beating MoE at matched FLOPs for fact-heavy tasks.
Titan-style test-time memory
Memorize at inference: an online surprise gradient trains a small memory module while the model runs — loss Δ = Σ ‖ f(x) − g(h, x) ‖² updated per token. Learning at test time; caching at inference time (doc 07) as its learned cousin.
07 Will They Replace the Transformer?
Probably not by revolution. By infiltration.
The incumbent's moat is not mathematical — Mamba and DeltaNet are legitimately better sequence math in several regimes. The moat is infrastructural:
08 Mental Models
Transformer = full transcript replay — every word is still on the table, reread each step. SSM = rolling notebook — you keep one page of distilled notes and update it as you read; fast, but you can't quote the original verbatim. JEV = typed form over the transcript — read the document once, then tick pre-printed checkboxes with confidence scores; you never write prose at all.
Every fixed-state model (Mamba, RWKV, DeltaNet, xLSTM) is a learned compression algorithm running online: h is a zip of history, and the update rule decides what's worth keeping. Compression ratio is the whole ballgame — too lossy and recall breaks (MQAR), too faithful and you've reinvented the KV cache.
BLT spends compute ∝ surprise; Mamba's Δ gates memory ∝ surprise; Titan memorizes ∝ surprise. Three architectures, one instinct: predictable bytes deserve cheap handling, surprising bytes deserve expensive handling — the same instinct as doc 07's cached-vs-fresh pricing.
09 Common Misconceptions
"SSMs are just worse transformers." They're different points on the recall-vs-cost curve. For pure sequence drift they match or beat attention; they lose specifically on exact multi-query recall. "Worse" without naming the task is meaningless.
"Linear time means faster responses for users." Not automatically: decode speed is bandwidth-bound (doc 10), and a fixed-size state must be read and rewritten every step. The O(s) win shows up at long context and in training/throughput, not necessarily in time-to-first-token on a 2K prompt.
"BitNet means you can quantize your existing model to 1.58 bits." No — b1.58 models are trained from scratch with quantization-aware objectives. Post-training-rounding an FP16 model to three values destroys it.
"Diffusion LMs are non-autoregressive, so they dodge the quality penalty." They trade it: parallelism for refinement waste and weaker long-range ordering. Speed ≠ quality; Mercury's 5–10× is real, but so is the frontier gap on long-form coherence.
"JEPA/decision models can replace LLMs." They abandon the generative interface — no free-form text comes out. They're complements: representation scorers and decision readouts beside a generative core, not instead of one.