Comment on New attention mechanisms that outperform standard multi-head attentionparentComments−daavidhauser2yxLSTM has a working memory and seems to outperform transformer architectures: https://arxiv.org/abs/2405.04517−jawon2yThanks for that. It looks like the kind of thing I'm looking for. I'll give it a read.
Comments
xLSTM has a working memory and seems to outperform transformer architectures: https://arxiv.org/abs/2405.04517
Thanks for that. It looks like the kind of thing I'm looking for. I'll give it a read.