I believe the attention mechanism we use now was introduced in 2014 by Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio in their paper titled "Neural Machine Translation by Jointly Learning to Align and Translate."
2014. It took almost a decade for the potential of this technique to be realized and come to the attention (heh) of most developers. I don't know what researchers are doing with Mamba and RWKV, but we should let them cook.
Comments
I believe the attention mechanism we use now was introduced in 2014 by Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio in their paper titled "Neural Machine Translation by Jointly Learning to Align and Translate."
2014. It took almost a decade for the potential of this technique to be realized and come to the attention (heh) of most developers. I don't know what researchers are doing with Mamba and RWKV, but we should let them cook.