Comment on Every Model Learned by Gradient Descent Is Approximately a Kernel MachineparentComments−dkural5yWouldn't this one-layer network be a lot less "compressive" than the multi-layer net, and in some sense "duplicate" subnetworks in earlier layers?
Comments
Wouldn't this one-layer network be a lot less "compressive" than the multi-layer net, and in some sense "duplicate" subnetworks in earlier layers?