Job Summary
Cloud Application Operations Engineer / Specialist SRE
We are looking for a skilled Cloud Operations Support professional with 5 to 10 years of experience to support and operate enterprise applications and infrastructure hosted on cloud platforms. The role requires hands-on experience in cloud operations, production support, monitoring, incident/problem management, and coordination across infrastructure, DevOps, and application teams to ensure stable, secure, and high-performing services.
Key Responsibilities
• optimize CI/CD pipelines to improve deployment efficiency and reliability.
• Lead design and implementation of scalable, secure, and highly available cloud architectures across enterprise environments.
• Implement SRE best practices including SLI/SLO definition, error budgets, and proactive reliability engineering.
• Drive automation-first culture by implementing Infrastructure-as-Code, orchestration frameworks, and self-healing systems.
• Establish security and compliance frameworks aligned with enterprise and regulatory requirements.
• Mentor engineering teams, drive capability development, and support hiring initiatives.
• Engage with stakeholders and leadership to provide insights on system health, risks, and improvement initiatives.
• Engage with stakeholders and leadership to provide insights on system health, risks, and improvement initiatives.
Skill Requirements
• Hands-on experience in Azure, AWS.
• Strong understanding of cloud infrastructure: VMs, storage, networking, IAM.
• Experience with Windows and/or Linux server administration.
• Knowledge of monitoring tools like Splunk, CloudWatch, Azure Monitor, Dynatrace, etc.
• Experience in incident management, RCA, and troubleshooting production issues.
• Exposure to Docker, Kubernetes, AKS or OpenShift is an advantage.
• Basic scripting skills in PowerShell, Shell, or Python.
• Understanding of CI/CD and Infrastructure as Code (Terraform preferred).
Preferred Qualifications
• Relevant cloud certifications are preferred.
• Experience in enterprise cloud operations and support models.
• Experience in DevSecOps / CI-CD pipelines
• Knowledge of scripting/automation (Shell, Python)
• Exposure to Site Reliability Engineering (SRE) or Platform Reliability Engineering (PRE)
• Experience in hybrid cloud / multi-cloud environments
Other Requirements
• Experience in DevSecOps / CI-CD tools (Azure DevOps, Jenkins, GitHub Actions)
• Knowledge of scripting/automation (Shell, Python); Exposure to automation tools (like Ansible)
• Exposure to Site Reliability Engineering (SRE) or Platform Reliability Engineering (PRE)
• Experience in hybrid cloud / multi-cloud environments