Comment on Zamba2-7BparentComments−oatsandsugar1yOn the page it states:Our novel shared-attention architecture allows more parameters to be allocated to the Mamba2 backbone. In turn, the shared transformer block preserves the rich cross-sequence dependencies of the attention computation.so sounds like it is transformer based?
Comments
On the page it states:
Our novel shared-attention architecture allows more parameters to be allocated to the Mamba2 backbone. In turn, the shared transformer block preserves the rich cross-sequence dependencies of the attention computation.
so sounds like it is transformer based?