Comment on State-space models can learn in-context by gradient descentComments−eli_gottlieb1yOur key insight is that the diagonal linear recurrent layer can act as a gradient accumulatorSo they're sort of reinventing the discrete-time differentiator from signal processing, but parameterized neurally?−radarsat11yConverging slowly on Kalman filters, calling it now.
Comments
So they're sort of reinventing the discrete-time differentiator from signal processing, but parameterized neurally?
Converging slowly on Kalman filters, calling it now.