Job Summary
We are seeking an experienced Site Reliability Engineer (SRE) with strong hands-on expertise in cloud infrastructure, platform reliability, automation, observability, and production support. This role focuses on improving service reliability through SLIs/SLOs and error budgets, reducing operational toil through automation, and partnering with global engineering teams to ensure resilient, scalable, and secure platforms.
Key Responsibilities
Reliability Engineering
• Define, measure, and report SLIs, SLOs, and error budgets for critical services.
• Drive service reliability improvements and systematically reduce operational toil through automation.
• Own capacity planning, performance tuning, and scalability initiatives.
• Lead blameless postmortems, root cause analyses, and corrective action tracking.
Platform Reliability & Automation
• Design and maintain CI/CD pipelines using Azure DevOps and Jenkins.
• Operate and manage Azure cloud infrastructure and Kubernetes platforms.
• Deploy and support containerized applications using Docker, Kubernetes, and Helm.
• Automate infrastructure provisioning and configuration using Terraform and Ansible.
• Manage artifacts and repositories using JFrog Artifactory.
Observability & Production Support
• Implement monitoring, logging, and observability using Prometheus, Grafana, Loki, and OpenTelemetry.
• Provide L2/L3 production support, incident management, troubleshooting, and RCA.
• Participate in on-call rotations supporting critical production services.
• Support PostgreSQL, Redis, and RabbitMQ environments, including high availability, backups, replication, and performance tuning.
Collaboration
• Partner with Development, QA, Product, Operations, and global engineering teams.
• Ensure platform availability, scalability, security, and performance.
Skill Requirements
• 8-12 years of experience in Site Reliability Engineering
• Strong hands-on experience with:
o Azure
o Kubernetes
o Docker
o Jenkins
o Terraform
o Ansible
• Proficiency in Python, Go, or Bash.
• Strong understanding of Networking & Security.
• Experience in L2/L3 production support, incident management, and root cause analysis.
• Experience working with global teams across multiple time zones.
• Strong communication, stakeholder management, and ownership mindset.
• Willingness to work from the Bangalore ITPL office, 5 days a week.
• Proficient in SECS/GEM protocols (E4, E5, E30, E37, E39, E40, E87, E90, E94, E116) for Equipment
• Experience in SEMI manufacturing process
Good to Have
• Experience with AI/GenAI concepts and AIOps practices.
• Exposure to chaos engineering and resilience testing.
• Azure (AZ-104/AZ-400), CKA, or CKAD certifications.
• Bachelor’s degree in computer science or a related field.