Yes, the paper says that the Transformer model is superior because it is more parallelizable and requires less time to train.
"Superior because it is X" is not the same as "superior while also being X". GPT-3 has managed to say both of those things within your short session, only the latter actually being correct.
You are right. I got so excited I missed it. I will play with it a little more to see if I can get it to make similar statements. It does not change that it is really impressive, but you absolutely have a point.
Comments
"Superior because it is X" is not the same as "superior while also being X". GPT-3 has managed to say both of those things within your short session, only the latter actually being correct.
You are right. I got so excited I missed it. I will play with it a little more to see if I can get it to make similar statements. It does not change that it is really impressive, but you absolutely have a point.