Tower Lead - Red Hat Enterprise Linux, Red Hat Satellite, Red Hat Cluster
India
Job Description
Tower Lead - Red Hat Enterprise Linux, Red Hat Satellite, Red Hat Cluster
Bengaluru, Karnataka

Job Summary

The RoleThe AI Infrastructure Engineer (L4) provides enterprise‑level architectural leadership and strategic ownership for high‑performance AI and ML infrastructure platforms. This role is responsible for designing, governing, and scaling next‑generation GPU/accelerator platforms, ensuring high availability, performance, and cost optimization for large‑scale training and inference workloads.The L4 engineer operates as a domain leader and platform owner, driving roadmap, standards, and cross‑functional alignment across engineering, operations, and business stakeholders.Competency Focus:
AI Platform Architecture, HPC at scale, distributed systems leadership, GPU platform strategy, cloud & hybrid optimization, SRE practicesKeywords:
AI Platform Architect, GPU Platform Owner, HPC Leader, Kubernetes Enterprise Architect, Infrastructure Strategy, AI Ops, SRE Governance

Linux, Kubernetes ,Pacemaker, DevOps Networking and Nvidia GPU Clusters

Key Responsibilities

Responsibilities (L4 Enhancements Applied)1. Architecture & Platform OwnershipDefine and own end‑to‑end architecture for enterprise AI infrastructure platforms (GPU, storage, network, orchestration)Establish reference architectures, design standards, and best practices for AI/ML platforms across hybrid/cloud environmentsLead platform modernization (AI‑native infrastructure, automation-first, GPUaaS models)Drive platform scalability strategy for multi-cluster, multi-region deploymentsOptimize Linux systems (Ubuntu, RHEL, Rocky) for AI/HPC workloads through NUMA, kernel, and clock tuning.Knowledge on Pacemaker and devopsLinux, Kubernetes ,Pacemaker, DevOps Networking and Nvidia GPU Clusters2. Strategic Leadership & RoadmapDefine long-term roadmap for AI infrastructure aligned with business and AI/GenAI adoption strategyEvaluate and onboard emerging GPU/accelerator technologies (NVIDIA, AMD, TPU, specialized AI hardware)Lead vendor strategy and partnerships (OEMs, cloud providers, NVIDIA ecosystem, etc.)Provide strategic advisory to leadership on AI infrastructure investments and scaling decisions3. Advanced Engineering & Performance OptimizationLead optimization of large-scale distributed training and inference architecturesDrive innovations in:GPU utilization efficiencyDistributed compute frameworks (Ray, Slurm, Kubernetes)High-speed interconnect optimization (InfiniBand, RDMA, NVLink)Establish best practices for LLM training/inference platforms (vLLM, Triton, TensorRT-LLM, DeepSpeed)Should have very good understanding on Hardware, Linux operating system, performance management, network, Good Knowledge on Kubernetes, Virtualization, Ansible, Red hat Satellite,, Knowledge on Hyperscale’s like AWS/Azure/GCP

Skill Requirements

Qualifications & Experience (Elevated for L4)Bachelor’s/master’s degree in computer science, Engineering, or related field12–18 years of overall infrastructure/platform engineering experience6–10 years in AI/ML infrastructure and distributed systems at scaleProven experience in architecting enterprise AI platforms (on-prem + cloud + hybrid)Deep expertise in:Kubernetes at scaleGPU infrastructure (NVIDIA ecosystem)HPC and distributed training frameworksStrong exposure to GenAI / LLM workloads and phantomizationDemonstrated experience in:Leading large programsDefining architecture & strategyCustomer-facing solutioningCertifications (L4 – Recommended / Advanced)Mandatory / Core:NVIDIA Certified AI Infrastructure (Associate/Professional)CKA / CKS (Kubernetes), DevOpsAWS / Azure Architect (Professional level preferred)Linux, PacemakerShould have very good understanding on Virtualization, Ansible, Red hat Satellite,, Knowledge on Hyperscalers like AWS/Azure/GCP

Other Requirements

1. Relevant Red Hat Certifications (E.G., Red Hat Certified Engineer Rhce, Red Hat Certified Specialist In Ansible Automation) Are Optional But Valuable
Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.