Comment on New attention mechanisms that outperform standard multi-head attentionparentComments−CGamesPlay2yThis is the paper, which appears to be cited in hundreds of others, some of which appear to be about efficiency gains. https://arxiv.org/abs/2105.03824−westurner2y"Fnet: Mixing tokens with fourier transforms" (2021) https://arxiv.org/abs/2105.03824https://scholar.google.com/scholar?cites=1423699627588508486...Fourier Transform and convolution.Shouldn't a deconvolvable NN be more explainable? #XAIDeconvolution: https://en.wikipedia.org/wiki/Deconvolution
Comments
This is the paper, which appears to be cited in hundreds of others, some of which appear to be about efficiency gains. https://arxiv.org/abs/2105.03824
"Fnet: Mixing tokens with fourier transforms" (2021) https://arxiv.org/abs/2105.03824
https://scholar.google.com/scholar?cites=1423699627588508486...
Fourier Transform and convolution.
Shouldn't a deconvolvable NN be more explainable? #XAI
Deconvolution: https://en.wikipedia.org/wiki/Deconvolution