Agreed that technical debt is a pain to deal with. But it's definitely possible to migrate and upgrade live systems without significant technical risk. The company just needs to make the decision to invest the proper resources and decide to have executive-level focus on Operations. From your comments it sounds like this was a failure of management to properly plan out these changes. Lack of resources shouldn't be an issue here, since clearly github's $100 Million+ in funding is enough to build a reliable HA system with adequate redundancies and hire the people who are qualified to run it. There just isn't anyone to point fingers at, including vendor screw-ups. All of those items would be addressed in a proper risk assessment plan.