Job Summary
The RoleThe AI Infrastructure Engineer (L4) provides enterprise‑level architectural leadership and strategic ownership for high‑performance AI and ML infrastructure platforms. This role is responsible for designing, governing, and scaling next‑generation GPU/accelerator platforms, ensuring high availability, performance, and cost optimization for large‑scale training and inference workloads.The L4 engineer operates as a domain leader and platform owner, driving roadmap, standards, and cross‑functional alignment across engineering, operations, and business stakeholders.Competency Focus:
AI Platform Architecture, HPC at scale, distributed systems leadership, GPU platform strategy, cloud & hybrid optimization, SRE practicesKeywords:
AI Platform Architect, GPU Platform Owner, HPC Leader, Kubernetes Enterprise Architect, Infrastructure Strategy, AI Ops, SRE Governance
Linux, Kubernetes ,Pacemaker, DevOps Networking and Nvidia GPU Clusters
Key Responsibilities
Responsibilities (L4 Enhancements Applied)1. Architecture & Platform OwnershipDefine and own end‑to‑end architecture for enterprise AI infrastructure platforms (GPU, storage, network, orchestration)Establish reference architectures, design standards, and best practices for AI/ML platforms across hybrid/cloud environmentsLead platform modernization (AI‑native infrastructure, automation-first, GPUaaS models)Drive platform scalability strategy for multi-cluster, multi-region deploymentsOptimize Linux systems (Ubuntu, RHEL, Rocky) for AI/HPC workloads through NUMA, kernel, and clock tuning.Knowledge on Pacemaker and devopsLinux, Kubernetes ,Pacemaker, DevOps Networking and Nvidia GPU Clusters2. Strategic Leadership & RoadmapDefine long-term roadmap for AI infrastructure aligned with business and AI/GenAI adoption strategyEvaluate and onboard emerging GPU/accelerator technologies (NVIDIA, AMD, TPU, specialized AI hardware)Lead vendor strategy and partnerships (OEMs, cloud providers, NVIDIA ecosystem, etc.)Provide strategic advisory to leadership on AI infrastructure investments and scaling decisions3. Advanced Engineering & Performance OptimizationLead optimization of large-scale distributed training and inference architecturesDrive innovations in:GPU utilization efficiencyDistributed compute frameworks (Ray, Slurm, Kubernetes)High-speed interconnect optimization (InfiniBand, RDMA, NVLink)Establish best practices for LLM training/inference platforms (vLLM, Triton, TensorRT-LLM, DeepSpeed)Should have very good understanding on Hardware, Linux operating system, performance management, network, Good Knowledge on Kubernetes, Virtualization, Ansible, Red hat Satellite,, Knowledge on Hyperscale’s like AWS/Azure/GCP
Skill Requirements
Qualifications & Experience (Elevated for L4)Bachelor’s/master’s degree in computer science, Engineering, or related field12–18 years of overall infrastructure/platform engineering experience6–10 years in AI/ML infrastructure and distributed systems at scaleProven experience in architecting enterprise AI platforms (on-prem + cloud + hybrid)Deep expertise in:Kubernetes at scaleGPU infrastructure (NVIDIA ecosystem)HPC and distributed training frameworksStrong exposure to GenAI / LLM workloads and phantomizationDemonstrated experience in:Leading large programsDefining architecture & strategyCustomer-facing solutioningCertifications (L4 – Recommended / Advanced)Mandatory / Core:NVIDIA Certified AI Infrastructure (Associate/Professional)CKA / CKS (Kubernetes), DevOpsAWS / Azure Architect (Professional level preferred)Linux, PacemakerShould have very good understanding on Virtualization, Ansible, Red hat Satellite,, Knowledge on Hyperscalers like AWS/Azure/GCP