Yeah, even in like 1996/1997 for certain industries there were hints as to the way things were going, even if it took 4/5 years for the transition to fully take place.
For example, in 1996/1997, Digital Domain (VFX industry) used a 'render farm' cluster of Carrera Alpha workstations running NT to render the Titanic film, instead of SGIs running IRIX. (SGIs were still often used on the artists workstations though, but progressively that shifted).
By 2001, many of those machines were x86 which were then often as fast as the SGIs and Alphas, even with x86's stack-based floating point architecture which handicapped it a bit, and the significantly higher memory bandwidth and larger caches of the SGIs.
In 1999, AMD Athlon became the first x86 CPU that was able to do both an addition and a multiplication during one clock cycle, for the 80-bit x87 floating-point numbers.
The previous Intel CPUs of the P6 family, from Pentium Pro to Pentium III, required 2 clock cycles for that, i.e. they reached at most half the throughput of Athlon at the same clock frequency. And Athlon had an even higher clock frequency.
So the launch of Athlon was one of the greatest jumps in floating-point performance per socket in the history of x86 CPUs.
It had a higher clock frequency than any Alpha. IBM POWER CPUs could do much more per clock cycle than Athlon, but their clock frequency was much lower, so Athlon was still faster.
One year and a half later Intel launched Pentium 4, which could match the throughput per clock cycle of Athlon, but only when executing new SSE2 programs, not when executing any legacy program.
This was such a huge transition for FEM on x86-64. We went from UltraSPARC III 1.2GHz 24 CPU system with 128GB of RAM to a smaller Opteron two chassis cluster linked with Infiniband, 500GB 32core/8NUMA nodes per server, and the speedup was almost 10x.
P4 was such a curious thing in it's own right; a lot of ambition that was perhaps too forceful.
Hell, if you -could- keep the pipeline from mispredicting and fed with data, one or two of it's internal ALUs actually ran at 2x the main CPU clock. Alas, that's an even bigger ask than adding SSE2 branching, and they decided to do RDRAM (Which, AFAIR was worse for overall latency than SDR or DDR)
We rapidly concluded the DEC Alpha-based systems served our batch-processing needs very well. They provide extremely high floating-point performance in commodity packaging. We were able to identify certain floating-point-intensive applications as port targets. The Alpha systems could be configured with large amounts of memory and fast networking at extremely attractive price points. Overall, the DEC Alpha had the best price/performance match for our needs. [...]
At this point, the decision was made to purchase 160 433MHz DEC Alpha systems from Carrera Computers of Newport Beach, California. Of those 160 machines, 105 of the machines are running Linux, the other 55 are running NT. The machines are connected with 100Mbps Ethernet to each other and to the rest of our facility. [...]
The floating-point power of the DEC Alpha made jobs run about 3.5 times faster than on our old SGI systems.
Comments
Yeah, even in like 1996/1997 for certain industries there were hints as to the way things were going, even if it took 4/5 years for the transition to fully take place.
For example, in 1996/1997, Digital Domain (VFX industry) used a 'render farm' cluster of Carrera Alpha workstations running NT to render the Titanic film, instead of SGIs running IRIX. (SGIs were still often used on the artists workstations though, but progressively that shifted).
By 2001, many of those machines were x86 which were then often as fast as the SGIs and Alphas, even with x86's stack-based floating point architecture which handicapped it a bit, and the significantly higher memory bandwidth and larger caches of the SGIs.
In 1999, AMD Athlon became the first x86 CPU that was able to do both an addition and a multiplication during one clock cycle, for the 80-bit x87 floating-point numbers.
The previous Intel CPUs of the P6 family, from Pentium Pro to Pentium III, required 2 clock cycles for that, i.e. they reached at most half the throughput of Athlon at the same clock frequency. And Athlon had an even higher clock frequency.
So the launch of Athlon was one of the greatest jumps in floating-point performance per socket in the history of x86 CPUs.
It had a higher clock frequency than any Alpha. IBM POWER CPUs could do much more per clock cycle than Athlon, but their clock frequency was much lower, so Athlon was still faster.
One year and a half later Intel launched Pentium 4, which could match the throughput per clock cycle of Athlon, but only when executing new SSE2 programs, not when executing any legacy program.
This was such a huge transition for FEM on x86-64. We went from UltraSPARC III 1.2GHz 24 CPU system with 128GB of RAM to a smaller Opteron two chassis cluster linked with Infiniband, 500GB 32core/8NUMA nodes per server, and the speedup was almost 10x.
P4 was such a curious thing in it's own right; a lot of ambition that was perhaps too forceful.
Hell, if you -could- keep the pipeline from mispredicting and fed with data, one or two of it's internal ALUs actually ran at 2x the main CPU clock. Alas, that's an even bigger ask than adding SSE2 branching, and they decided to do RDRAM (Which, AFAIR was worse for overall latency than SDR or DDR)
Regarding the "Titanic" film, from the Linux Journal (https://www.linuxjournal.com/article/2494):
I believe the original Toy Story was 'created' on SGIs, but the render farm was SPARC.