Nearly Optimal Attention Coresets
arXiv:2605.05602v1 Announce Type: cross
Abstract: We consider the problem of estimating the Attention mechanism in small space, and prove the existence of coresets for it of nearly optimal size. Specifically, we show that for any set of unit-norm keys…