Job Summary
| We are looking for experienced SRE leader to drive enterprise-wide reliability transformation across cloud, digital, SaaS, industrial, AI-enabled, and mission-critical platforms. This role combines technical expertise with strategic transformation leadership to modernize operations through SRE, AI-driven observability, autonomous operations, platform engineering, and intelligent automation practices. As a trusted advisor and hands-on transformation leader, you will partner with engineering, operations, cloud, platform, architecture, security and AI-engineering teams to institutionalize modern Site Reliability Engineering practices at scale. |
Key Responsibilities
"
Key Responsibilities:
Lead and drive enterprise-wide SRE transformation initiatives across business applications and platforms.
Assess current operational maturity and define SRE adoption roadmaps, operating models, governance frameworks, reliability standards, and implementation strategies.
Coach and mentor engineering, operations, cloud, platform, and application teams on SRE principles, SLI/SLO frameworks, Error Budgets, reliability engineering, and operational excellence.
Establish and institutionalize reliability metrics, service health management practices, reliability governance models, and continuous service improvement frameworks.
Define and govern enterprise observability strategies, telemetry standards, monitoring frameworks, and best practices leveraging Splunk Observability Cloud, Splunk ITSI, and Splunk Enterprise.
Lead and guide the implementation of enterprise observability capabilities, including monitoring standards, OpenTelemetry adoption, Business Observability, service health monitoring, KPI frameworks, alert quality management, and operational analytics.
Guide teams in implementing effective Incident Management, Problem Management, Root Cause Analysis (RCA), and Blameless Postmortem practices.
Drive continuous improvement initiatives focused on reducing MTTD and MTTR, improving MTBF, and enhancing overall service reliability and operational efficiency.
Promote automation-first operations through toil reduction, intelligent automation, self-healing, and operational excellence initiatives.
Advise and guide teams on implementing AIOps capabilities, including event correlation, anomaly detection, intelligent alerting, operational intelligence, and predictive operations.
Collaborate with business, engineering, architecture, cloud, platform, security, and operations stakeholders to embed reliability into technology delivery and operational processes.
Drive adoption of Agile and Scrum practices through reliability reviews, retrospectives, and data-driven continuous improvement initiatives.
Measure and report SRE adoption, reliability KPIs, operational maturity, and business outcomes to leadership teams."
Skill Requirements
"Must Have Skills :
- Deep expertise in Site Reliability Engineering (SRE) including SLI, SLO, Error Budgets, reliability governance, and service health management with 13+ years of experience
- Proven experience leading enterprise SRE transformation programs and coaching engineering and operations teams
- Strong hands-on expertise with Splunk Observability Cloud (OpenTelemetry, APM, RUM, Synthetic Monitoring), Splunk ITSI (Event Analytics, Service Modeling & KPI Mgmt., Glass Tables, Business Observability), and Splunk Enterprise (Logging, SPL, Log Analytics, Dashboarding))
- Strong understanding of enterprise observability architecture, telemetry strategy, monitoring standards, and operational analytics
- Experience driving Incident Management, Root Cause Analysis (RCA), and Blameless Postmortems
- Strong track record in Automation and Toil Reduction initiatives (Ansible, Python, RPA and other automation platforms)
- Experience implementing AIOps and Intelligent Operations practices to improve service reliability and operational efficiency
- Hands-on experience with cloud platforms such as Azure, AWS, and GCP
- Experience institutionalizing SRE practices for business applications including custom applications (Java, .NET, and other technologies running on VM and container platforms), SAP, Salesforce, SaaS/COTS, and - business-critical enterprise applications
- Drive continuous improvement through Agile and Scrum practices, using regular retrospectives, reliability reviews, and data-driven actions to improve service reliability, team effectiveness, MTTD, MTTR, and MTBF.
- Strong stakeholder management, leadership, communication, and mentoring skills"
Other Requirements
Good to Have Skills : - Knowledge of DevOps practices, CI/CD pipelines, GitOps, and release automation - Experience with Reliability Engineering practices including Resilience Testing and Chaos Engineering - Experience with Agentic AI, AI Agents, and AI-driven solutions for SRE, DevOps, Automation, and AMS operations - Experience with Platform Engineering, Internal Developer Platforms (IDP), and self-service engineering models - SRE, Splunk, Cloud, Observability, or related industry certifications. |