Job Summary
We are seeking a technically strong L2 SRE / Production Support Analyst with expertise in production support, site reliability engineering, and platform operations. The role involves managing critical incidents, troubleshooting complex application and platform issues, performing RCA, supporting Kubernetes environments, improving observability, and driving automation initiatives to enhance service reliability.
Key Responsibilities
- Handle L1 escalations and resolve complex application, infrastructure, and Kubernetes/OpenShift issues.
- Perform RCA and drive permanent fixes for recurring incidents.
- Own P1/P2 incidents and coordinate with Development, DevOps, Infrastructure, and QA teams.
- Support deployments, release validations, post-implementation checks, and environment stability.
- Monitor and optimize application health using Splunk, AppDynamics, Grafana, BigPanda, Kibana, and Prometheus.
- Troubleshoot Linux, API, database, and batch processing issues.
- Automate operational tasks using Python and Shell scripting.
- Maintain SOPs, runbooks, and knowledge articles while supporting continuous service improvements.
Skill Requirements
- Strong Linux/Unix administration and troubleshooting.
- Hands-on Kubernetes/OpenShift support and troubleshooting.
- Python and Shell scripting.
- Monitoring tools: Splunk, Grafana, AppDynamics, BigPanda, Kibana, Prometheus/Dynatrace.
- Autosys (mandatory); Control-M (good to have).
- SQL knowledge with Oracle, MSSQL, or MongoDB.
- Exposure to AWS, Docker, CI/CD, and DevOps practices.
- Application log analysis, API troubleshooting, and production support.
Other Requirements
- Strong debugging, analytical, and problem-solving skills.
- Knowledge of ITIL Incident, Problem, and Change Management processes.
- Ability to manage critical incidents and lead technical bridge calls.
- Strong communication and cross-functional collaboration skills.
- Proven ability to automate recurring operational tasks and improve platform reliability.