Retentive Network: A Successor to Transformer for Large Language Modelsarxiv.org 11 pointsvagabund3 years ago3 commentsSaveHideCopy link On HNComments−fgfm3yThe repo the paper is pointing to indicates that the code will be released within ~1 week! If the sheer difference in VRAM requirements and latency holds up, this will seriously be a major breakthrough for LLM architectures!Can't wait to try it out−deimos0x023yhttps://github.com/Jamie-Stirling/RetNet−vagabundOP3yBrief twitter thread with performance metrics. Looks big if results hold up.https://twitter.com/arankomatsuzaki/status/16811139775001845...
Comments
The repo the paper is pointing to indicates that the code will be released within ~1 week! If the sheer difference in VRAM requirements and latency holds up, this will seriously be a major breakthrough for LLM architectures!
Can't wait to try it out
https://github.com/Jamie-Stirling/RetNet
Brief twitter thread with performance metrics. Looks big if results hold up.
https://twitter.com/arankomatsuzaki/status/16811139775001845...