The structured process of identifying, responding to, resolving, and learning from service disruptions or failures, minimizing downtime and impact on users, and improving system reliability and resilience.