Comment on Every Model Learned by Gradient Descent Is Approximately a Kernel MachineparentComments−Flashtoo5yThe claim also applies to GANs as you can simply use a masking function to indicate which inputs were used for each optimization timestep, like the author suggests for stochastic gradient descent in remark 5.
Comments
The claim also applies to GANs as you can simply use a masking function to indicate which inputs were used for each optimization timestep, like the author suggests for stochastic gradient descent in remark 5.