Skip to content

Comment on Highly efficient matrix transpose in Mojoparent

Comments

It looks because it does.

(2771.35/2775.49 - 1) * 100 = -.14916285052369131300

Flagged.

Updated the title to the original. I did base the numbers on

"This kernel archives 1437.55 GB/s compared to the 1251.76 GB/s we get in CUDA" (14.8%) which is still impressive

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.