A 12-minute interactive explainer for people who know what a Transformer's KV cache is and want to understand how "attention as synaptic memory" works in Pathway's Dragon Hatchling (BDH). Prerequisite: matrix–vector products.
The claim you can test below: a fixed-size synaptic matrix written by a Hebbian outer-product rule stores in-context associations without allocating a slot per token, and it recalls them exactly only while the stored keys stay nearly non-overlapping. As the number of stored pairs or the overlap between keys grows, recall degrades through interference, not by "running out of slots".
If this were false, you would be able to store many overlapping pairs in the matrix and still recall every one exactly. Try.
64 neurons. Each stored pair is a sparse, non-negative key pattern x and value pattern y (k active neurons each). Writing pair i adds the outer product yixiT to the synaptic matrix σ. Reading with a probe key xq computes σ·xq, then keeps the k largest entries as the recalled value.
σ, 64×64 = 4,096 synapses. Same size at 1 stored pair or 40. Purple outline: columns read by the probe key.
Recall accuracy averaged over all stored pairs, as the number of pairs grows, at the current k and decay. This whole curve is recomputed live; the dot is where you are.
True value yq (k active neurons)
Recalled top-k. Teal = correct, orange = wrong neuron, pale = missed.
Reading is linear. With m pairs stored and decay weights wi, the readout for probe xq is
With random k-sparse keys in n = 64 neurons, two keys share about k²/n neurons on average, so a typical stored pair leaks about k²/n of its value into every read. Crosstalk on a given wrong neuron is then roughly (m−1)·(k³/n²) for u = 0, and the memory stays exact only while the largest such leak is below the signal, k. That is why the failure point moves earlier when you raise k, and why nothing about the matrix's size changes as you approach it. The "capacity" is set by key overlap and count, not by slot exhaustion.
Decay adds a second forgetting channel: an old pair's signal shrinks as (1−u)age·k while newer pairs keep writing on top of it. Set u = 0.2 and probe pair 1 with m = 20; the memory has not been overwritten, it has been faded under. That is the trade a fixed-size state makes to stay bounded.
Contrast with a Transformer's KV cache: it appends one key and one value per token, so memory grows as 2·n·m (at m = 40 here, 5,120 cells versus σ's constant 4,096) but softmax attention with a sharp temperature retrieves each stored value essentially exactly regardless of how many are stored. The two designs fail in opposite directions: one runs out of memory, the other runs out of separability.
What you just manipulated is the attention state of BDH, written in its graph form. Kosowski, Uznański, Chorowski, Stamirowska and Bartoszkiewicz (Pathway, 2025) define BDH's per-layer dynamics by three equations (their eq. 6; layer index l, token index t):
Map it line by line to the toy:
| Toy | BDH (paper, eq. 6 and 8) | What is changing |
|---|---|---|
| σ ← (1−u)σ + y xᵀ | σ ← (σ + y xᵀ ⊙ G_s)·U. G_s is a fixed synapse mask (all-ones in the BDH-GPU correspondence, eq. 9); U is a damping/rotation matrix such as ALiBi decay or RoPE rotation. | State, not parameters. σ is reset for a new context; the trained weights G_x, G_y (or E, D_x, D_y) do not move during inference. |
| read: σ·x_q, then top-k | y = ((G_y σ x)⁺ ⊙ x). The synaptic read σ·x is passed through a trained linear map, a ReLU, and is gated by x itself. | Activity. The paper reports ~5% of neurons active per token in trained BDH-GPU models.developer-reported |
| k-sparse random 0/1 patterns | x, y are non-negative, learned, and sparse; not random. | Training presumably shapes keys toward separability. The paper does not quantify interference capacity. |
| n = 64, one σ | n = 32,768 neurons at 25M parameters (d = 256); one σ per layer, L layers, weights shared across layers. | State per layer is n×d in BDH-GPU, comparable to the parameter count; the paper argues this ~1:1 state-to-parameter ratio matters. |
The GPU formulation (BDH-GPU, eq. 8) never materialises σ. It keeps a compressed state ρ = E·σ, where E is the trained d×n encoder, and updates it as ρ ← (ρ + LN(E y) xᵀ)·U. Two consequences matter here. First, the write is still an outer product accumulated into a fixed-size state, so the interference mechanism you saw is inherited, not removed. Second, the compression through E adds a further lossy step that this toy does not model. The paper notes that the same state can equivalently be realised in the graph form with a sparse G_s of O(nd) edges (Claim 4).
The paper's own framing (its section 1.2) is the one the toy reproduces: a fixed ruleset G (parameters) plus an evolving ruleset σ (fast weights) that is potentiated whenever neuron i in y fires just before neuron j in x. They call the resulting attention "synaptic plasticity with Hebbian learning" operating at "potentiation scales of minutes for the brain (up to hundreds of tokens)", and report that specific individual synapses strengthen consistently when the model processes a specific concept (their section 6.3, monosemantic synapses).developer-reported, small models
BDH-CQ (Engdahl, Kosowski, Chorowski, Stamirowska, Uznański, Jiang, Phadke, Kinas, Zhong, 2026) is the system built on BDH layers that learns an ARC task from its demonstrations at inference time. Its technical report writes the contextual memory as S_t = U_θ(S_{t−1}, D_t) with fixed θ, and states that linear attention is the simplest standalone realisation of this, capturing the special case S_t = S_{t−1} + U_θ(D_t). That special case is exactly the toy at u = 0: state accumulates additively, one demonstration at a time, in a state that does not grow with the number of demonstrations. The report explicitly relates this to attention, fast-weight memory and linear-attention views of contextual association.
Evidence discipline. The report says the exact update rules and dimensions of BDH-CQ are proprietary, so the toy is an illustration of the stated special case, not a model of BDH-CQ. The result that is relevant to associative memory is the controlled color-permutation test: BDH-CQ applied a demonstration-defined mapping of 2 to 8 simultaneous color bindings correctly on 96/96 held-out outputs, which is contextual binding, the thing σ is for. Its headline number, 29.5% pass@2 on public ARC-AGI-1 at a computed $0.0007 per task, is a developer-reported benchmark that was reproduced in a black-box audit by co-authors from Bielik AI and NYU without weight access; that is not an independent reimplementation, and the same report shows a hard failure on composing a color swap with relocation (0/72).benchmark, not deployment
Keep k = 8 and u = 0. Before touching the slider, predict the smallest number of stored pairs m at which the average recall curve first drops below 100%. Then find it.
Then press "New random patterns" a few times. The threshold moves. That is not a bug: the claim is about a mechanism, and the exact number depends on which keys happened to overlap. If the threshold ever stopped existing, the claim would be false.
Limitations of the toy, stated plainly. Random 0/1 patterns instead of learned sparse activations; no encoder/decoder (E, D), no ReLU-lowrank step, no LayerNorm, no per-neuron rotation (RoPE), single layer, G_s = all-ones, n = 64. Its capacity numbers are not BDH's. What it does show faithfully is the write rule, the read rule, the fixed state size, and the two forgetting channels (interference and decay) that any additive fast-weight memory inherits, including the linear-attention special case named in the BDH-CQ report. Delta-rule variants (Gated DeltaNet, 2025) erase the old value for a key before writing the new one, which trades this interference for a different set of limits; the toy does not implement that.
In two or three sentences, explain to a colleague why a Hebbian matrix can hold many associations without growing, and what makes it fail. Nothing you type is stored or sent anywhere.
Check your explanation against these: