Tower Lead (Support & Operations)
India
Job Description
Tower Lead (Support & Operations)
Hyderabad, Telangana

Job Summary

Role Summary

The Observability & Datadog Engineer is responsible for designing, implementing, and managing enterprise monitoring and observability solutions using Datadog. The role focuses on proactive monitoring, performance management, incident reduction, root cause analysis, and improving service reliability across infrastructure, applications, cloud platforms, and business services. Internal observability frameworks emphasise Metrics, Events, Logs, and Traces (MELT), alert correlation, anomaly detection, dashboards, and SRE-driven operations.

Key Responsibilities

Key Responsibilities

  • Design and implement end-to-end observability solutions using Datadog.
  • Configure and maintain monitoring for:
    • Infrastructure (Servers, VMs, Network, Storage)
    • Cloud platforms (Azure, AWS, GCP)
    • Containers and Kubernetes
    • Applications and Microservices
    • Databases and Middleware
  • Build and manage:
    • Dashboards
    • Alerts and Monitors
    • Service Level Indicators (SLIs)
    • Service Level Objectives (SLOs)
  • Implement Application Performance Monitoring (APM), Log Management, Distributed Tracing, and Synthetic Monitoring.
  • Perform root cause analysis using metrics, logs, traces, and events.
  • Reduce alert noise through correlation, automation, and threshold optimisation.
  • Partner with Development, DevOps, Cloud, and SRE teams to improve application reliability.
  • Automate monitoring deployment using Terraform, APIs, Python, or Shell scripting.
  • Support incident, problem, and change management processes.
  • Define observability standards, governance, and best practices.

Skill Requirements

Required Skills

Datadog

  • Infrastructure Monitoring
  • APM (Application Performance Monitoring)
  • Log Management
  • Distributed Tracing
  • RUM (Real User Monitoring)
  • Synthetic Monitoring
  • Network Performance Monitoring
  • Dashboard & Alert Configuration
  • SLO/SLI Management
  • Datadog API & Integrations

Technical Skills

  • Linux & Windows Administration
  • Kubernetes / OpenShift
  • Docker Containers
  • Azure / AWS / GCP
  • CI/CD Tools
  • REST APIs
  • Terraform / Infrastructure as Code
  • Python, PowerShell, Bash

Observability Skills

  • Metrics, Logs, Events & Traces (MELT)
  • Monitoring Strategy
  • Alert Correlation
  • Anomaly Detection
  • Capacity & Performance Management
  • Incident Management
  • Root Cause Analysis

Other Requirements

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.