Comment on Highly efficient matrix transpose in MojoparentComments−77pt771yIt looks because it does.(2771.35/2775.49 - 1) * 100 = -.14916285052369131300Flagged.−timmydOP1yUpdated the title to the original. I did base the numbers on"This kernel archives 1437.55 GB/s compared to the 1251.76 GB/s we get in CUDA" (14.8%) which is still impressive
Comments
It looks because it does.
Flagged.
Updated the title to the original. I did base the numbers on
"This kernel archives 1437.55 GB/s compared to the 1251.76 GB/s we get in CUDA" (14.8%) which is still impressive