A lot of the AI progress is stalled because these chips do not handle thermal contraction/expansion well at all, and end up permanently destroying these chips. HBM has been a failure at this scale with chips getting destroyed every 4 months or so during operations.
I'd like to see some actual science saying, here was the problem, here's how we solved it, here's the AFR data, here's this running after X cycles etc. Nobody has done this reliably yet. That entire industry is hiding the bodies.
IIRC it's thousands per day on these systems, mainly due to high power density and low mass, even small lapses in computation (100-500ms) rapidly change the temperatures of stacked die.
So even a GPU averaging 98% utilization may have thousands of cycles per day.
Compared to a regular server blade it may be dozens or barely any at all.
Comments
A lot of the AI progress is stalled because these chips do not handle thermal contraction/expansion well at all, and end up permanently destroying these chips. HBM has been a failure at this scale with chips getting destroyed every 4 months or so during operations.
I'd like to see some actual science saying, here was the problem, here's how we solved it, here's the AFR data, here's this running after X cycles etc. Nobody has done this reliably yet. That entire industry is hiding the bodies.
How many contraction/expansion cylces do you normally see in a DC setting, generally? Is there a measure for that to baseline against?
IIRC it's thousands per day on these systems, mainly due to high power density and low mass, even small lapses in computation (100-500ms) rapidly change the temperatures of stacked die.
So even a GPU averaging 98% utilization may have thousands of cycles per day.
Compared to a regular server blade it may be dozens or barely any at all.
Naive question, but could this not be fixed by a scheduler? If it's only idle 2% of the time, give it busy work for that 2%
the easier answer would be to remove the ability to idle