Job Summary
Cloud Application Operations Engineer / Specialist - SRE
Job Summary
We are looking for a skilled Cloud Operations Support professional with 5 to 10 years of experience to support and operate enterprise applications and infrastructure hosted on cloud platforms. The role requires hands-on experience in cloud operations, production support, monitoring, incident/problem management, and coordination across infrastructure, DevOps, and application teams to ensure stable, secure, and high-performing services.
Key Responsibilities
• Monitor, support, and maintain cloud-hosted applications and infrastructure to ensure service availability and performance.
• Provide L2/L3 production support including alert monitoring, log analysis, and troubleshooting.
• Handle incidents, service requests, and problem management including root cause analysis (RCA).
• Coordinate with infrastructure, database, network, and application teams for issue resolution.
• Support change management, patching, and maintenance activities with proper validation.
• Work with monitoring and observability tools to detect performance issues and improve reliability.
• Drive automation initiatives for operations and deployments.
• Ensure security, governance, and compliance across cloud environments.
• Maintain documentation, SOPs, and operational runbooks.
• Participate in 24x7 support or rotational shifts as required.
Skill Requirements
Skill Requirement :
• Hands-on experience in Azure, AWS.
• Strong understanding of cloud infrastructure: VMs, storage, networking, IAM.
• Experience with Windows and/or Linux server administration.
• Knowledge of monitoring tools like Splunk, CloudWatch, Azure Monitor, Dynatrace, etc.
• Experience in incident management, RCA, and troubleshooting production issues.
• Exposure to Docker, Kubernetes, AKS or OpenShift is an advantage.
• Basic scripting skills in PowerShell, Shell, or Python.
• Understanding of CI/CD and Infrastructure as Code (Terraform preferred).
Preferred Skills
• Experience in DevSecOps / CI-CD tools (Azure DevOps, Jenkins, GitHub Actions)
• Knowledge of scripting/automation (Shell, Python); Exposure to automation tools (like Ansible)
• Exposure to Site Reliability Engineering (SRE) or Platform Reliability Engineering (PRE)
• Experience in hybrid cloud / multi-cloud environments
Other Requirements
Preferred Qualifications
• Relevant cloud certifications are preferred.
• Experience in enterprise cloud operations and support models.
• Experience in DevSecOps / CI-CD pipelines
• Knowledge of scripting/automation (Shell, Python)
• Exposure to Site Reliability Engineering (SRE) or Platform Reliability Engineering (PRE)
• Experience in hybrid cloud / multi-cloud environments
Other Requirement :
• Knowledge of scripting/automation (Shell, Python); Exposure to automation tools (like Ansible)