Skip to content

Comment on The Alpha 21264 CPU: NT's Greatest RISC (1998)parent

Comments

In 1999, AMD Athlon became the first x86 CPU that was able to do both an addition and a multiplication during one clock cycle, for the 80-bit x87 floating-point numbers.

The previous Intel CPUs of the P6 family, from Pentium Pro to Pentium III, required 2 clock cycles for that, i.e. they reached at most half the throughput of Athlon at the same clock frequency. And Athlon had an even higher clock frequency.

So the launch of Athlon was one of the greatest jumps in floating-point performance per socket in the history of x86 CPUs.

It had a higher clock frequency than any Alpha. IBM POWER CPUs could do much more per clock cycle than Athlon, but their clock frequency was much lower, so Athlon was still faster.

One year and a half later Intel launched Pentium 4, which could match the throughput per clock cycle of Athlon, but only when executing new SSE2 programs, not when executing any legacy program.

This was such a huge transition for FEM on x86-64. We went from UltraSPARC III 1.2GHz 24 CPU system with 128GB of RAM to a smaller Opteron two chassis cluster linked with Infiniband, 500GB 32core/8NUMA nodes per server, and the speedup was almost 10x.

P4 was such a curious thing in it's own right; a lot of ambition that was perhaps too forceful.

Hell, if you -could- keep the pipeline from mispredicting and fed with data, one or two of it's internal ALUs actually ran at 2x the main CPU clock. Alas, that's an even bigger ask than adding SSE2 branching, and they decided to do RDRAM (Which, AFAIR was worse for overall latency than SDR or DDR)

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.