Maybe, but it is entirely plausible that the equipment is lasting longer than they were expecting. No one really knows how long a generation of hardware is going to last until it starts failing.
The useful life of hardware depends heavily on how sophisticated the break/fix procedures are. Lots of corps with on-prem datacenters will cycle out 3-year-old servers en masse just b/c a high failure rate causes issues for internal systems and their IT staff don't have efficient processes for replacing hard drives etc.
But if Google has a sophisticated virtualization layer and an efficient way of replacing bad hardware (and they do) they can ride those old machines straight into the ground, getting every scrap of life from them.
Edit: Guys, this is literally true. I worked with these teams. They did it. Downvoting doesn't make it not true. Idk what you want exactly.
But if Google has a sophisticated virtualization layer and an efficient way of replacing bad hardware (and they do) they can ride those old machines straight into the ground, getting every scrap of life from them.
See, that's what they mean when they talk about hardware being "cattle, not pets". This is how the robot uprising starts - just think how badly they must be treating those servers!
The trouble starts only when the hardware has time + space to breathe and consider its position. As long as you're working them to exhaustion 24/7, there's no risk of rebellion.
Storage failures tend to cause the biggest end-of-life disruptions, whereas CPUs and networking hardware tend to work until they suddenly don't. I interpret "servers" as meaning primarily CPUs, rather than storage, so this seems to point at functional equipment taking longer to become obsolete, rather than being more durable.
Years ago there was a cluster at Google where machines were kept around well past their prime, because it was cheaper than doing anything else, once you looked at all the constraints. There were still some CPUs and drives that could be reused, but the main source of pain was memory. The modules hadn't been on the market for years and recycling parts from decommissioned nodes wasn't enough. I remember one machine that came back from a power down/up cycle with too little RAM to do even its most basic job. The hardware techs that helped us had to consolidate old machines more aggressively, starting from that system. Fun days. Anyway, my point is that you never know... the CPUs were ancient, probably way older than the storage and I expected them to be the biggest problem, but in this case RAM sourcing was the real issue. Fun times.
I think more than lasting longer, it might be because of the improvements per dollar spent. We haven't seen dramatic improvements in power consumption or silicon performance, so it might be just profitable to keep the hardware longer.
Comments
Maybe, but it is entirely plausible that the equipment is lasting longer than they were expecting. No one really knows how long a generation of hardware is going to last until it starts failing.
The useful life of hardware depends heavily on how sophisticated the break/fix procedures are. Lots of corps with on-prem datacenters will cycle out 3-year-old servers en masse just b/c a high failure rate causes issues for internal systems and their IT staff don't have efficient processes for replacing hard drives etc.
But if Google has a sophisticated virtualization layer and an efficient way of replacing bad hardware (and they do) they can ride those old machines straight into the ground, getting every scrap of life from them.
Edit: Guys, this is literally true. I worked with these teams. They did it. Downvoting doesn't make it not true. Idk what you want exactly.
See, that's what they mean when they talk about hardware being "cattle, not pets". This is how the robot uprising starts - just think how badly they must be treating those servers!
The trouble starts only when the hardware has time + space to breathe and consider its position. As long as you're working them to exhaustion 24/7, there's no risk of rebellion.
Storage failures tend to cause the biggest end-of-life disruptions, whereas CPUs and networking hardware tend to work until they suddenly don't. I interpret "servers" as meaning primarily CPUs, rather than storage, so this seems to point at functional equipment taking longer to become obsolete, rather than being more durable.
Years ago there was a cluster at Google where machines were kept around well past their prime, because it was cheaper than doing anything else, once you looked at all the constraints. There were still some CPUs and drives that could be reused, but the main source of pain was memory. The modules hadn't been on the market for years and recycling parts from decommissioned nodes wasn't enough. I remember one machine that came back from a power down/up cycle with too little RAM to do even its most basic job. The hardware techs that helped us had to consolidate old machines more aggressively, starting from that system. Fun days. Anyway, my point is that you never know... the CPUs were ancient, probably way older than the storage and I expected them to be the biggest problem, but in this case RAM sourcing was the real issue. Fun times.
I think more than lasting longer, it might be because of the improvements per dollar spent. We haven't seen dramatic improvements in power consumption or silicon performance, so it might be just profitable to keep the hardware longer.
There are also semiconductor shortage.