Job Summary
Critical Incident Manager is responsible for leading the end-to-end management of high-impact and high-urgency IT incidents that affect business operations, critical services, users, customers, or revenue. The role acts as the central point of coordination during major incidents, driving rapid service restoration, structured communication, stakeholder alignment, escalation governance, and post-incident improvement. The ideal candidate should have strong ITIL-based incident management experience, proven crisis leadership capability, and the ability to coordinate technical teams, vendors, service owners, and leadership stakeholders under pressure.
Key Responsibilities
Lead, coordinate, and govern all critical and major incidents from identification through service restoration and closure. • Initiate and facilitate incident bridges, war rooms, technical triage calls, and executive communication channels as required. • Assess incident impact, urgency, priority, affected services, business risk, and customer impact to ensure correct classification and escalation. • Coordinate cross-functional resolver groups, service owners, vendors, infrastructure teams, application teams, and leadership stakeholders to drive timely resolution. • Maintain clear command and control during incidents by tracking actions, owners, timelines, dependencies, decisions, risks, and next steps. • Provide timely, accurate, and business-friendly updates to stakeholders, including senior leadership, customer teams, and service delivery management. • Ensure adherence to ITIL incident management, major incident management, escalation, communication, and service restoration processes. • Drive workaround identification, service recovery actions, permanent fix coordination, and handover to problem management where required. • Prepare and publish incident summaries, major incident reports, executive updates, post-incident review notes, and lessons-learned documentation. • Collaborate with problem management teams to support root cause analysis, corrective actions, preventive actions, and recurrence reduction. • Track SLA compliance, response timelines, resolution timelines, communication quality, and incident aging to ensure governance discipline. • Identify process improvement opportunities and contribute to continual service improvement across incident response, communication, reporting, and escalation management.
Skill Requirements
Minimum 5 years of experience in Incident Management, Major Incident Management, IT Service Management, or IT Operations in a complex enterprise environment. • Hands-on experience managing P1/P2, Sev1/Sev2, business-critical, or customer-impacting incidents. • Strong working knowledge of ITIL practices, including incident management, major incident management, problem management, change management, service request management, and continual improvement. • Experience coordinating multiple technical teams such as infrastructure, cloud, network, database, application, security, workplace, service desk, and vendor support teams. • Ability to manage high-pressure situations with calmness, discipline, urgency, and structured decision-making. • Excellent verbal and written communication skills, with the ability to simplify technical updates for business and leadership audiences. • Experience preparing major incident reports, post-incident reviews, RCA inputs, trend analysis, and service improvement recommendations. • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline is preferred. • ITIL Foundation certification is required; ITIL Intermediate, ITIL 4 Managing Professional, PMP, Agile, or relevant service management certifications are preferred. • Experience with ServiceNow, BMC Remedy, Jira Service Management, PagerDuty, Splunk, monitoring tools, or similar ITSM and observability platforms is preferred.
Other Requirements
Major Incident Management and Crisis Coordination • ITIL Incident, Problem, Change, and Continual Improvement Practices • War Room Facilitation and Technical Bridge Management • Executive Communication and Stakeholder Management • Escalation Management and Decision Governance • Service Restoration Planning and Business Impact Assessment • RCA Coordination, PIR Support, and Corrective Action Tracking • SLA, KPI, Dashboard, and Incident Reporting Governance • Cross-functional Collaboration with Technical, Vendor, and Business Teams • Strong Analytical Thinking, Prioritization, and Risk Management