Job Summary
Seeking an experienced Support Engineer having experience of 7-10 years working as Site Reliability Engineer (SRE) with strong expertise in Kubernetes, Linux, Cloud Platforms (AWS/Azure/GCP), Observability, Automation, and Production Support. Responsible for managing and supporting production Kubernetes environments, ensuring platform reliability, availability, security, scalability, and operational excellence.
Key Responsibilities:
- Manage and maintain Kubernetes clusters, including deployments, upgrades, capacity planning, RBAC, networking, storage, Ingress, ConfigMaps, Secrets, Services, Persistent Volumes, StatefulSets, DaemonSets, Jobs/CronJobs, Helm, and Autoscaling (HPA/VPA).
- Provide 24x7 production support, incident management, RCA, postmortems, service restoration, and SLA/SLO compliance.
- Monitor and improve platform reliability using Prometheus, Grafana, Loki, Elastic Stack, OpenTelemetry, and AlertManager.
- Troubleshoot Kubernetes, Linux, container (Docker/OCI), networking (DNS, Load Balancers, TLS, Ingress), cloud, and infrastructure issues.
- Automate operational tasks through Bash, Python, Terraform, and Ansible, and support Infrastructure as Code practices.
- Support CI/CD and release management using GitHub Actions, GitLab CI, Jenkins, and ArgoCD (preferred).
- Perform patching, cluster maintenance, security updates, backups, disaster recovery validation, and platform upgrades.
- Create runbooks, operational documentation, dashboards, alerts, and capacity planning reports.
- Collaborate with Development, Platform Engineering, Security, Networking, Cloud Operations, and DevOps teams to improve system resilience and operational efficiency.
Required Skills: Kubernetes Administration, Linux, Docker/OCI, AWS/Azure/GCP, Networking, CI/CD, GitOps, Observability, Incident Management, RCA, Automation, Terraform, Ansible, Bash, Python.
Preferred: CKA/CKS certification, Cloud certifications, Multi-cluster/Multi-region Kubernetes, Service Mesh (Istio/Linkerd), High Availability, Disaster Recovery, Security Hardening, Capacity Planning, Performance Tuning, Cost Optimization, Chaos Engineering, AI-assisted Observability.
Key Competencies: Strong troubleshooting, ownership, production support, customer focus, communication, collaboration, continuous improvement, and ability to perform under pressure.
Success Metrics: High platform availability, improved MTTR, reduced incidents and alert noise, SLA/SLO compliance, increased automation coverage, successful upgrades/maintenance, and customer satisfaction.
Key Responsibilities
Skill Requirements
Other Requirements
Preferred Qualifications
- Certified Kubernetes Administrator (CKA)
- Certified Kubernetes Security Specialist (CKS)
- Cloud certifications (AWS/Azure/GCP)
- Experience supporting multi-cluster Kubernetes environments.
- Experience with service mesh technologies (Istio/Linkerd).
- Experience with GitOps workflows.