Job Summary
Major Incident Manager Role Summary The Major Incident Manager is responsible for governing and coordinating the end-to end Major Incident Management (MIM) process for high-severity incidents that significantly impact business operations. The role ensures rapid service restoration through structured stakeholder engagement, real-time incident governance, decision oversight, and adherence to established operational processes. The Major Incident Manager provides 24x7 lifecycle control, coordinates cross functional teams, drives effective communication, and leads post-incident reviews to improve service resilience and prevent recurrence.
Key Responsibilities
Key Responsibilities 1. Major Incident Governance & Lifecycle Management • Drive the end-to-end Major Incident lifecycle in accordance with defined processes and operating frameworks. • Validate incident severity, progression, escalation paths, and closure readiness. • Ensure structured execution across all phases of the incident lifecycle. • Govern incident timelines, SLA compliance, and restoration objectives. • Maintain process adherence and operational governance throughout incident resolution. 2. 24x7 Incident Coordination & Service Restoration • Monitor, triage, and manage high-severity incidents on a 24x7 basis. • Mobilize and coordinate technical, operational, and business teams during critical incidents. • Conduct and manage incident bridge calls to facilitate rapid resolution. • Drive structured escalation management and controlled decision-making. • Ensure timely restoration of services while minimizing business impact. 3. Stakeholder Communication & Governance • Coordinate engagement with all relevant stakeholders throughout the incident lifecycle. • Deliver accurate, timely, and consistent status communications. Classification: Internal • Manage executive, customer, and operational communications during major incidents. • Publish incident reports, updates, and resolution summaries. • Ensure communication standards and governance requirements are consistently followed. 4. Resolution Oversight & Decision Management • Review and validate workaround and resolution strategies. • Ensure resolution actions align with operational, business, and risk management requirements. • Facilitate collaboration between cross-functional teams to remove blockers and accelerate recovery. • Provide governance and oversight for critical decisions during incident management activities. • Track and govern action plans until service restoration is achieved. 5. Shift Handover & Operational Continuity • Conduct structured shift handovers to maintain 24x7 incident management continuity. • Ensure comprehensive briefing on: o Incident status and current severity o Business impact assessment o Actions completed and ongoing activities o Workaround and resolution progress o Pending tasks and next steps o Risks, dependencies, and escalation points • Ensure explicit ownership transfer and accountability for all open actions. • Validate handover completeness through standardized templates, incident logs, and bridge notes. • Perform handover reviews to confirm understanding of priorities, action plans, and communication cadence. 6. Post-Incident Review & Continuous Improvement Classification: Internal • Lead and govern Major Incident Reviews (MIR) and Post-Incident Reviews (PIR). • Analyze incident trends, Root Cause Analysis (RCA) findings, and SLA performance. • Identify corrective and preventive improvement opportunities. • Track implementation of improvement actions through closure. • Drive continuous improvement initiatives to enhance service stability and operational resilience. Key Deliverables • Major Incident governance and lifecycle management. • High-severity incident coordination and service restoration. • Executive and stakeholder communication management. • Incident bridge management and escalation coordination. • Shift handover governance and operational continuity. • MIR/PIR facilitation and follow-up actions. • SLA compliance reporting and service improvement recommendations.
Skill Requirements
Required Skills & Competencies • Strong Major Incident Management and IT Service Management (ITSM) expertise. • Knowledge of ITIL Incident and Major Incident Management processes. • Excellent stakeholder management and communication skills. • Strong crisis management and decision-making abilities. • Experience coordinating cross-functional technical teams. • Ability to manage high-pressure situations and critical business-impacting incidents. • Strong analytical, governance, and reporting capabilities. • Experience conducting post-incident reviews and driving continuous improvement initiatives. Operating Model Classification: Internal • 24x7 operational support model. • Structured shift handovers with formal governance checkpoints. • Continuous stakeholder engagement throughout incident resolution. • End-to-end oversight from incident initiation through closure and post-incident review.