The Architecture Atlas

Fourteen ways past the Transformer.

Every architecture below answers a specific failure of the one before it: quadratic attention and its growing KV-cache pushed the field to fixed-state recurrence Phase 1, then to adaptive test-time compute inside a latent loop Phase 2, and finally away from causal autoregression altogether Phase 3. Click any node to open its dossier. This map is the living companion of the survey in preparation.

1997 → 2026, three lanes

Node position = year of the anchor paper. The axis is compressed between 1997 and 2018.

↑ Select an architecture on the timeline to read its dossier.

Master comparison

All fourteen, side by side.

Click a column header to sort. “Inference / token” is the marginal cost of generating one more token at context length N.

Architecture Year Phase Training parallelism Inference / token State / memory Generation paradigm

Compiled from the primary papers linked in each dossier; the survey in preparation gives the full derivations under a unified notation.