Memory
Weights are parameters × bytes per parameter. KV cache per token is
2 × layers × KV heads × head dim × bytes — the factor of two being K and V.
Total cache scales with concurrency and sequence length, which is why it, not the
weights, is usually what actually runs you out of memory at long context.