Job Summary
Senior Site Reliability Engineer with extensive experience driving operational excellence, reliability engineering, service maturity initiatives, observability strategies, automation programs, and large-scale production support transformations. Proven ability to lead major incidents, influence engineering and business stakeholders, improve service reliability, and implement enterprise-wide monitoring, automation, and operational governance frameworks.
Key Responsibilities
Operational Leadership & Reliability Strategy
- Service Reliability Leadership
- Operational Excellence Programs
- Reliability Engineering Strategy
- Service Maturity Improvement
- Operational Governance
- Resiliency and Stability Planning
- Service Performance Management
- Reliability Roadmap Development
- Strategic Operational Initiatives
Incident & Problem Management Leadership
- Major Incident Management
- Executive Incident Communications
- Critical Production Event Leadership
- Root Cause Elimination
- Problem Management Governance
- Preventive Control Implementation
- Incident Trend Analysis
- Risk Reduction Initiatives
Observability, Monitoring & Reliability Engineering
- Enterprise Monitoring Strategy
- Splunk Architecture and Observability
- Dashboard Governance and Standardization
- Alert Optimization Programs
- Event Correlation Strategy
- KPI and SLA Reporting
- Service Health Metrics
- Operational Analytics
- Error Budget and Reliability Metrics
- Service Performance Reporting
Automation & Continuous Improvement
- Toil Reduction Programs
- Operational Automation Strategy
- REXX Automation Solutions
- Continuous Service Improvement (CSI)
- Process Optimization
- Operational Risk Reduction
- Platform Efficiency Improvements
- Reliability Improvement Programs
Mainframe & Enterprise Platform Support
- IBM Mainframe Operations Leadership
- z/OS Platform Support
- CICS Transaction Processing
- JCL and REXX Expertise
- Batch Processing Governance
- OPC Scheduling Operations
- Endevor Release Governance
- Enterprise Production Support
Release & Operational Readiness Governance
- Operational Readiness Reviews
- Production Acceptance Criteria
- Release Governance
- Change Risk Assessment
- Customer Onboarding Support
- Migration and Modernization Readiness
- Production Support Models
- Recovery and Supportability Standards
Stakeholder & Cross Functional Leadership
- Engineering Partnership
- Product Owner Collaboration
- Infrastructure Coordination
- Business Stakeholder Engagement
- Executive Reporting
- Technical Risk Communication
- Service Strategy Recommendations
- Operational Governance Reviews
Knowledge Management & Mentorship
- Team Mentoring
- SME Enablement
- Knowledge Management Programs
- Runbooks and Documentation Standards
- Operational Best Practices
- Support Organization Enablement
- Operational Culture Transformation
Skill Requirements
Splunk Netcool Omnibus xMatters Domo Remedy Jira Bitbucket XL Release (XLR) Endevor SDSF File-AID Abend-Aid IDCAMS DFSORT / SyncSort Key Strengths Reliability Engineering Operational Leadership Incident Command and Crisis Management Service Availability Improvement Automation and Toil Reduction Service Maturity Advancement Operational Risk Mitigation Execut