SeniorAdministrator - AWS IAC, Terraform,Python
India
Job Description
SeniorAdministrator - AWS IAC, Terraform,Python
Goutam Budd Nagar, Uttar Pradesh

Job Summary

  Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js) 

Key Responsibilities

  Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js) 

Skill Requirements

 Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Other Requirements

 Mandatory Preference -Candidate with strong development skills with coding debugging AWS serverless with Node.js Resiliency & Operational Excellence — AWS Serverless Reliability, resiliency, and operational excellence for mission‑critical AWS serverless platforms, ensuring high availability, low MTTR, and strong production governance using Dynatrace‑driven observability. Resiliency strategy for serverless architectures (Lambda, API Gateway, async/event‑driven systems) SLOs / SLIs / Error Budgets for critical API’s Incident analysis and post‑incident reviews Dynatrace observability: dashboards, alert tuning, dependency mapping, RCA acceleration Operational excellence improvements: incident reduction, MTTR improvement, toil automation Reliability guardrails embedded into CI/CD and production readiness reviews Core Responsibilities Design & enforce resiliency patterns: timeouts, retries, circuit breakers, throttling, graceful degradation Lead major incidents and drive actionable RCAs with sustained fixes Build signal‑driven alerts aligned to SLOs (noise reduction focus) Enable automation & self‑healing where feasible Required Experience 5-6+ years in SRE/DevOps/Production Engineering Deep hands‑on with AWS serverless (Lambda, API Gateway, SQS/SNS, DynamoDB/RDS) Strong expertise in Dynatrace for serverless monitoring & triage Proven success improving availability, MTTR, and incident trends Solid coding/scripting (Python / Java / Node.js)

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.