Sr Subject Matter Expert (Support&Ops)
India
Job Description
Sr Subject Matter Expert (Support&Ops)
Lucknow, Uttar Pradesh

Job Summary

Job Description – Site Reliability Engineer (SRE)Position: Site Reliability Engineer (SRE)
Experience: 12-15 Years
Location: Flexible / HybridRole SummaryWe are seeking a highly motivated Site Reliability Engineer (SRE) responsible for ensuring the reliability, availability, scalability, and performance of enterprise applications and infrastructure. The ideal candidate will combine software engineering, automation, cloud operations, and observability expertise to reduce operational toil and improve platform resilience through automation and proactive engineering practices. 

Key Responsibilities

Reliability & OperationsEnsure high availability, performance, and scalability of critical business services.Define, monitor, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).Lead major incident response, root cause analysis (RCA), and post-incident reviews.Implement proactive monitoring, alerting, and self-healing capabilities.Conduct capacity planning, performance tuning, and reliability assessments.Automation & EngineeringDesign and implement Infrastructure as Code (IaC) solutions using Terraform, Ansible, or similar tools.Develop automation scripts using Python, PowerShell, Bash, or equivalent technologies.Drive CI/CD pipeline enhancements and deployment automation.Reduce manual effort (toil) through engineering-led automation initiatives.Cloud & Platform ManagementManage and optimize Azure, AWS, GCP, and hybrid-cloud environments.Support containerized workloads using Kubernetes and Docker.Implement high-availability, disaster recovery, and resilience architectures.Collaborate with infrastructure, application, security, and DevOps teams to improve platform stability.Observability & MonitoringImplement and maintain observability platforms for logs, metrics, traces, and events.Establish monitoring dashboards, reliability KPIs, and performance baselines.Drive noise reduction, event correlation, and predictive analytics initiatives.Utilize tools such as Azure Monitor, Dynatrace, AppDynamics, ELK, Grafana, Prometheus, Splunk, or equivalentRequired Skills 

Skill Requirements

Technical SkillsStrong experience in Cloud Platforms (Azure/AWS/GCP).Expertise in Linux and/or Windows administration.Hands-on experience with Terraform, Ansible, Puppet, or similar automation tools.Strong scripting/programming skills in Python, PowerShell, Bash, or Go.Experience with CI/CD tools such as Azure DevOps, Jenkins, GitHub Actions, or GitLab.Good understanding of networking, storage, databases, and middleware technologies.Experience with container orchestration platforms such as Kubernetes 

Other Requirements

null
Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.