The Architecture Atlas
Fourteen ways past the Transformer.
Every architecture below answers a specific failure of the one before it: quadratic attention and its growing KV-cache pushed the field to fixed-state recurrence Phase 1, then to adaptive test-time compute inside a latent loop Phase 2, and finally away from causal autoregression altogether Phase 3. Click any node to open its dossier. This map is the living companion of the survey in preparation.
1997 → 2026, three lanes
Node position = year of the anchor paper. The axis is compressed between 1997 and 2018.
↑ Select an architecture on the timeline to read its dossier.
Master comparison
All fourteen, side by side.
Click a column header to sort. “Inference / token” is the marginal cost of generating one more token at context length N.
| Architecture | Year | Phase | Training parallelism | Inference / token | State / memory | Generation paradigm |
|---|
Compiled from the primary papers linked in each dossier; the survey in preparation gives the full derivations under a unified notation.