Job Summary
The SRE Engineer is responsible for the day-to-day reliability, availability, observability, and operational support of business applications and platforms. The role focuses on maintaining and operating production services, improving reliability through SLI/SLO practices, leveraging Splunk-based observability solutions, driving automation, and continuously improving operational performance and service quality.
Key Responsibilities
"Key Responsibilities:
- Maintain and operate production applications and platforms to ensure high availability, reliability, performance, and operational stability.
- Configure, maintain, and optimize observability, monitoring, logging, alerting, and service health solutions using Splunk Enterprise, Splunk ITSI, and Splunk Observability Cloud.
- Develop and maintain dashboards, alerts, reports, service health views, and KPI monitoring solutions.
- Perform Incident Management, troubleshooting, Root Cause Analysis (RCA), blameless postmortems, and service restoration activities.
- Implement and support SLI, SLO, Error Budget tracking, service health monitoring, reliability reporting, and continuous improvement initiatives.
- Improve alert quality by reducing noise, eliminating false positives, and optimizing monitoring coverage.
- Support AIOps capabilities such as event correlation, anomaly detection, alert enrichment, and operational analytics.
- Automate operational tasks, monitoring, reporting, and remediation activities using scripting and automation tools.
- Collaborate with application, cloud, infrastructure, and operations teams to improve reliability, performance, and operational efficiency.
- Maintain runbooks, SOPs, knowledge articles, and operational documentation.
- Participate in reliability reviews, blameless postmortems, Agile ceremonies, and continuous service improvement initiatives.
- Support production releases, platform upgrades, maintenance activities, and operational readiness reviews."
Skill Requirements
"Must Have Skills:
- 6+ years of experience in Production Operations, SRE, Observability, Application Support, or Operations Engineering roles.
- Experience with observability, logging, alert management, alert correlation, and alert quality improvement practices using Splunk ITSI, Splunk Observability Cloud & Splunk Enterprise logging.
- Experience implementing and supporting SLI, SLO, Error Budget tracking, service health monitoring, reliability reporting, and continuous improvement initiatives.
- Experience with Incident Management, Problem Management, Troubleshooting, Change Management, Root Cause Analysis (RCA), and Blameless Postmortem practices.
- Experience supporting business applications running on VM and container platforms.
- Experience with cloud platforms such as Azure, AWS, and GCP.
- Experience with automation technologies to reduce the operational toil (Ansible, Python, RPA and other automation platforms)
- Experience with ITSM processes and tools such as ServiceNow.
- Strong analytical, troubleshooting, communication, and collaboration skills."
Other Requirements
"Good to Have Skills:
- Knowledge of DevOps practices, CI/CD pipelines, GitOps, and release automation.
- Experience supporting Java, .NET, SAP, Salesforce, SaaS/COTS, or other enterprise applications.
- Experience with Reliability Engineering practices including Resilience Testing and Chaos Engineering.
- Exposure to OpenTelemetry and modern observability frameworks.
- Exposure to Agentic AI, AI-driven Operations, and AI-assisted observability solutions
- Splunk, SRE, Cloud, or Observability-related certifications."