Tautology: a statement that is true by virtue of its logical form alone (Merriam Webster). Fits perfectly the “a neural network that can't process text is less good at processing text that one that can”
It's perfectly legit to discuss how a Transformer would perform if only the Self-Attention part was removed
It only shows that you don't understand the topic at all (but hey, you talked about closed-form solutions and quantum computing elsewhere in this discussion with others so why I am even surprised…)
Insofar as the actual other networks you've mentioned they fail to beat Transformers
They don't “fail to beat transformers”, they beat transformers that aren't the state of the art and are less good that the ones that are. And that's not really a surprise given that they are more recent and have much less manpower working on them. I don't expect them to replace transformers until they make some hypothetic breakthrough that'd makes them significantly better than transformers. That's what path dependence is. But they are still a good illustration to the point that you don't need to have attention heads to exhibit the capabilities of LLMs. (Remember you set the bar at GPT-2 level, and they are far beyond that)
because language comprehension simply cannot be done without sensitivity to word context
And these models actually have a way to represent context so this criticism completely miss the mark. That's really hilarious that you make this kind of claim in an HN thread about SSM. How come you have no idea at all about what a state-space model is and then feels confident enough to come and argue in the comment section…
No, that's what you're missing from the beginning: the breakthrough of transformers was scalability. Now we have other models that are equally scalable and as such roughly equally performant (and that's not a surprise).
But the ship has sailed and nobody is gonna switch to something else than transformers if it's not significantly better, and as such the other approaches are going to stay behind because every marginal improvement come to transformers first (because that's what practically everyone is working on) and alternative models are playing catch-up.
This is a remarkable example of path dependence.
Interpreting this as “transformers are fundamentally superior” is the mistake I'm trying to help you correct.
The breakthrough of transformers was scalability. The next breakthrough of equivalent importance will be entirely different or it won't be.
Comments
Tautology: a statement that is true by virtue of its logical form alone (Merriam Webster). Fits perfectly the “a neural network that can't process text is less good at processing text that one that can”
It only shows that you don't understand the topic at all (but hey, you talked about closed-form solutions and quantum computing elsewhere in this discussion with others so why I am even surprised…)
They don't “fail to beat transformers”, they beat transformers that aren't the state of the art and are less good that the ones that are. And that's not really a surprise given that they are more recent and have much less manpower working on them. I don't expect them to replace transformers until they make some hypothetic breakthrough that'd makes them significantly better than transformers. That's what path dependence is. But they are still a good illustration to the point that you don't need to have attention heads to exhibit the capabilities of LLMs. (Remember you set the bar at GPT-2 level, and they are far beyond that)
And these models actually have a way to represent context so this criticism completely miss the mark. That's really hilarious that you make this kind of claim in an HN thread about SSM. How come you have no idea at all about what a state-space model is and then feels confident enough to come and argue in the comment section…
Yes, a breakthrough that does what Self-Attention is doing, rather than just scaling up.
No, that's what you're missing from the beginning: the breakthrough of transformers was scalability. Now we have other models that are equally scalable and as such roughly equally performant (and that's not a surprise).
But the ship has sailed and nobody is gonna switch to something else than transformers if it's not significantly better, and as such the other approaches are going to stay behind because every marginal improvement come to transformers first (because that's what practically everyone is working on) and alternative models are playing catch-up.
This is a remarkable example of path dependence.
Interpreting this as “transformers are fundamentally superior” is the mistake I'm trying to help you correct.
The breakthrough of transformers was scalability. The next breakthrough of equivalent importance will be entirely different or it won't be.
Scalability wasn't an architectural breakthrough. It was merely a discovery..
How are these words even in contradiction to each other?
By intentionally lacking context.
You aren't even trying to pretend your sentence make sense, I see…
so Kindergarten bro.
We are now at the Markov chain level of sentence, it's not even grammatically resembling a sentence at that point. Congratz!
Like your 1 word replies.
No.
Your post attacking my grammar contains two grammatical errors of it's own. Can you spot them?
I'm not attacking your grammar, but now I question your ability to understand a sentence.
(English isn't my first language, BTW, so I'd be grateful if you could point the grammar errors)
The 4 words that inadvertently summarized your entire disposition.
It's not inadvertent.
You've spent all your credit for any form of consideration at that point.
And now you've even spent your credit for my attention at all, this conversation ends here.
By "inadvertently" I meant you didn't intend the 4 words as an encapsulation of your entire persona.