3 ms·
Refresh cycles depend on 3 values -- end of warranty, end of life, and end of support. If you're end of support, you have no bug fixes, no technical support, a
by caw 10y ago
Refresh cycles depend on 3 values -- end of warranty, end of life, and end of support.
If you're end of support, you have no bug fixes, no technical support, and very limited part availability. You probably won't be running EOS hardware anywhere in your datacenter without a very defined plan to migrate off of it quickly or it's in a lab somewhere no one cares if it dies.
End of Life is when you can't buy the hardware anymore, so you start (or have already started) with the next generation of hardware. Depending on the policy, this could be the oldest hardware you have. You'll get the new stuff in and migrate over, and get rid of these machines most likely once they're out of warranty. Maybe you'll keep this hardware until EOS for low-priority work.
End of Warranty is the last major date, and it depends on date of purchase of a particular piece of hardware, rather than a product line. If you're on a frequent refresh cycle, once the product costs you money to repair then you'll get rid of it. You'll buy a warranty that most likely matches your depreciation cycle - 4 or 5 years. If you choose to keep machines that are out of warranty they start getting downgraded to less critical roles, or are scrapped as soon as they break. Repairs are variable costs in terms of parts and labor, and that doesn't play nicely with budgeting to manage a machine.
In general the harder it is to replace a single piece of equipment or the more expensive it is, the more likely you'll be running until EOL or EOS. Fileservers and networking (especially core networking) will be kept longer than cheap web servers.
EOW, EOS, and EOL all vary depending on vendors and product lines. It's the reason Dell and HP has a separate business line of desktops and laptops, because the EOL and EOS dates are known in advance, and you can standardize your equipment even when purchasing in multiple batches over 2 years. This reduces how much compatibility testing you have to perform, which is really beneficial for huge deployments.
- webmaven 10y agoInteresting answer, but how do those values affect the sort of massive DCs that AMZN/GOOG run? They don't even replace individual machines, they replace whole racks once a designated % of machines in it are failing. Presumably the HW within a rack is homogenous, though, so perhaps it is treated as a single large computer from this perspective?
- caw 10y agoYes, when you're buying enough hardware you treat the rack or row as a purchase unit. A rack is typically hardware plus switching for that rack, so just a few network and power cables are needed to connect a rack to the rest of the infrastructure. If these companies replace a whole rack when a percentage of machines fail that's a business decision based on cost of hardware versus repair costs. Where I worked we had warranties, so the hardware would get fixed (or die in place for things past warranty) and the entire rack would get rolled out at one time, since we bought in rack increments of capacity and thus they all had the same warranty dates. Letting things die in place obviously depends on how much spare capacity you have in your DC, which could be site-specific within large companies.
- webmaven 10y ago> Letting things die in place obviously depends on how much spare capacity you have in your DC, by "capacity" do you mean room? I associate capacity in this context more with power & cooling (neither of which are consumed by disabled machines). Or do you mean that if the DC's capacity is constrained, letting systems die in place and only replacing "mostly dead" racks on a rolling basis keeps capacity utilization below a critical threshold?
- caw 10y agoGenerally dead systems don't hurt the datacenter in terms of infrastructure utilization, as they don't use power and they don't use cooling, but they could prevent you from bringing in new compute capacity. So you have to figure out whether you can run both the old and slightly broken systems at the same time as the new and faster systems. Capacity could mean physical space, since if floor tiles are occupied by a rack you'll have to put it somewhere else. But even if you have physical space to put another rack somewhere, that needs connected to power and cooling. Depending on your cooling layout (hot boxing, cold boxing, vented floors, etc) you may not have appropriate cooling for a new rack, especially a newer, higher density rack. Or you have sufficient cooling overall, but not for that density in that location because all your high density stuff was planned to go in a designated area. Same thing with power -- you may not have enough power cables for the power distribution units within a rack (you'd typically have 2 so the redundant power supplies don't share a PDU as a point of failure), the cabling may not reach, the electrical panel could be maxed out on amperage or breaker slots, or you'd throw the load on each of the electrical phases too far out of balance (A previous manager of mine was an electrical engineer, I'm not too sure about the real-world technicalities of this other than he tried to keep everything balanced). Networking could also be a limiting factor. You could run out of switch ports or SFPs because you only planned for N connections per racks on however many racks, and now you want to keep hardware around longer.