Senior Technical Architect
India
Job Description
Senior Technical Architect
Bangalore, Karnataka

Job Summary

AI Observability Principal Architect(18+ Years)

Description

The AI Observability Principal Architect is responsible for defining and delivering end-to-end observability and AIOps capabilities across enterprise applications and platforms. The role focuses on enabling proactive detection, intelligent event correlation, and automated incident response, improving system reliability and operational efficiency.

This role will lead observability strategy, standardization, and implementation across teams, ensuring clear visibility into system health, business workflows, and performance outcomes. The lead will work closely with SRE, application, infrastructure, and service management teams to embed outcome-driven observability (CUJs, SLIs/SLOs) and drive continuous improvement in incident detection and resolution.

Key Responsibilities

Assignment Deliverables

  • Define and implement observability strategy, standards, and roadmap
  • Establish telemetry framework (logs, metrics, traces) and instrumentation standards
  • Define Critical User Journeys (CUJs) and map Service Level Indicators (SLIs) / SLOs
  • Enable actionable alerting aligned to user/business impact
  • Implement alert-to-incident automation with correct routing and ownership
  • Drive AIOps capabilities:
    • Event correlation and alert noise reduction
    • Root-cause-based incident generation
    • Predictive detection and anomaly identification
  • Build and optimize observability dashboards for operations and leadership visibility
  • Enable automation and self-healing playbooks for recurring incidents
  • Lead major incident support and post-incident improvements

Ensure continuous improvement through observability lifecycle management

Skill Requirements

Required Skills

  • 18+ years experience in IT Operations / SRE / Observability / Platform Engineering
  • Strong expertise in:
    • Observability (logs, metrics, traces, distributed tracing)
    • SRE practices (SLIs, SLOs, error budgets)
    • Incident management and automation
  • Experience in AIOps / Event Management:
    • Event correlation, alert deduplication, noise reduction
  • Hands-on experience with observability platforms and ITSM integrations
  • Strong understanding of:
    • Distributed systems and cloud environments
    • Telemetry pipelines and data integration
  • Experience in designing automation workflows, runbooks, and self-healing mechanisms
  • Strong stakeholder management and ability to work across application, infra, and operations teams

Other Requirements

Nice to Have

  • Experience with Azure observability stack / OpenTelemetry
  • Experience with ServiceNow ITOM / AIOps
  • Exposure to AI-driven observability (anomaly detection, predictive analytics)
  • Experience in building executive dashboards and reporting frameworks
Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.