Track Lead (Support & Operations)
India
Job Description
Track Lead (Support & Operations)
Hyderabad, Telangana

Job Summary

Site Reliability Engineer (SRE) Lot 3 - HCBU | L2 - Senior SRE Engineer Experience Operating Model Focus Reporting 5-8 Years 24x7 Reliability Engineering Operations and Transformation SRE Lead / Reliability Architect ROLE PURPOSE Drive service reliability, resilience, observability, and automation across cloud, infrastructure, platform, and application services. Apply SRE practices to improve availability, reduce operational toil, accelerate recovery, and support the transition toward proactive and autonomous operations.

Key Responsibilities

KEY RESPONSIBILITIES Reliability and SLO Management Define and maintain SLIs, SLOs, error budgets, golden signals, and service-health measures for critical services. Monitor SLO performance and error-budget consumption, identify reliability risks, and coordinate corrective actions with service owners. Contribute to service classification, availability targets, operational readiness, and reliability improvement roadmaps. Observability and AIOps Implement and enhance full-stack observability across applications, APIs, cloud, infrastructure, Kubernetes, databases, and network services. Build dashboards, service maps, alerts, synthetic checks, log analytics, tracing, and dependency views using enterprise monitoring platforms. Improve signal quality through event correlation, alert tuning, noise reduction, anomaly detection, and actionable routing. Automation and Toil Reduction Identify repetitive operational activities and develop automated runbooks, health checks, self-service workflows, and self-healing solutions. Create reliable automation using Python, PowerShell, Shell, Terraform, Ansible, CI/CD, and platform-native capabilities. Ensure automation includes testing, approvals, audit evidence, exception handling, rollback, and human oversight where required. Incident, Problem, and Resilience Engineering Support P1/P2 incident response, technical triage, service restoration, and cross-tower coordination. Lead or contribute to RCA, post-incident reviews, problem investigations, corrective actions, and recurrence-prevention measures. Conduct capacity reviews, failover validation, DR exercises, resilience testing, performance analysis, and single-point-of-failure assessments. Governance and Continuous Improvement Prepare reliability reports covering SLO compliance, MTTR, incidents, recurring failures, telemetry gaps, automation benefits, and technical debt. Participate in service reliability reviews and maintain evidence for agreed reliability and operational controls. Collaborate with Infrastructure, Cloud, DevOps, Database, Network, Security, Application, and ITSM teams to embed reliability by design.

Skill Requirements

REQUIRED SKILLS AND EXPERIENCE Hands-on experience in SRE, production engineering, platform operations, cloud operations, or infrastructure reliability. Practical knowledge of SLI/SLO definition, error budgets, availability, latency, throughput, saturation, and service-health measurement. Experience with Datadog, Splunk, Grafana, Prometheus, AppDynamics, Dynatrace, Azure Monitor, GCP Operations Suite, or similar tools. Operational experience with Azure and/or GCP, Kubernetes/container platforms, Linux, and Windows environments. Automation and scripting skills in Python, PowerShell, Bash, or Shell, with exposure to APIs and Git-based version control. Knowledge of incident, problem, change, availability, capacity, and service-continuity processes using ServiceNow and ITIL practices. Strong troubleshooting, analytical, documentation, stakeholder communication, and cross-functional collaboration skills. PREFERRED SKILLS Terraform and Ansible AIOps, event correlation, and automated remediation Chaos engineering and resilience testing Azure DevOps, GitHub Actions, GitLab CI/CD, or Jenkins Distributed tracing, real-user monitoring, and synthetic monitoring Cloud-native architecture, FinOps awareness, and security-by-design practices KEY DELIVERABLES AND SUCCESS MEASURES Approved SLI/SLO definitions, service-health dashboards, and error-budget reporting for assigned services. Reduced alert noise, detection time, restoration time, repeat incidents, and manual operational effort. Increased adoption of automated runbooks, self-healing, proactive detection, and preventive remediation. Improved service availability, resilience, capacity readiness, observability coverage, and SLO compliance. Complete RCA, post-incident actions, operational documentation, and auditable evidence for reliability controls.

Other Requirements

PREFERRED CERTIFICATIONS SRE Foundation | Certified Kubernetes Administrator (CKA) | Microsoft Azure Administrator / Architect | Google Professional Cloud DevOps Engineer | Datadog Certification | ITIL v4

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.