Skip to content

Comment on Rust SIMD on the GPUparent

Comments

A VGPR is not the same thing as a vector register like in SSE4 or AVX. Each addressed register contains a single 32-bit value. A VGPR differs from an SGPR in that each thread in a thread group can have a different value in that register. An SGPR will have a uniform value shared with all threads in a group.

An add instruction on an AMD GPU adds two scalar values. If they're in a VGPR then each thread will add two values unique to that thread. A SIMD ISA as is common on a CPU is different because an add instruction explicitly adds a vector of values. xmm1 stores 128-bits of data. VGPR[1] stores 32-bits of data vectored over 32-64 threads in a thread group.

Without special instructions a thread can't access the VGPR values stored in other threads.

But those instructions exist, see Section 7.9 "Cross-Lane and Data Parallel Processing (DPP)"

AMD documentation describes RDNA as a vector ISA, so I don't understand why you say it is not.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.