Job Summary
Skill: Network + LinuxL4 – Lead Engineer / Architect / Platform OwnerSkill RequirementDeep expertise in enterprise networking and Linux platform architectureResponsible for design, governance, scalability, and performance engineering of network and OS layers supporting AI/HPC/Kubernetes platforms at scaleOwns platform standards, reference architectures, and reliability frameworks across multi-cluster and multi-region environmentsCCNA, CCNP preferred with working knowledge on InfiniBand and Mellanix SwitchesNetworking (L4)Certifications (Preferred):CCNP (Mandatory baseline)CCIE (Strongly preferred)Specialized certifications in:Data Center / Service Provider NetworkingInfiniband, MellonixCloud Networking (AWS/Azure/GCP advanced tracks)Experience:10–15+ years in enterprise / hyperscale network engineering and architectureProven experience in:Designing AI/HPC network architectures with Infiniband and Spectrum-X (Ethernet with RoCEv2)Multi-cluster Kubernetes networking (CNI, service mesh, ingress architectures)Hybrid and multi-cloud connectivity (on-prem ↔ cloud GPU clusters)Deep exposure to:High-performance GPU networking environmentsLarge-scale, multi-tenant infrastructure platformsSkill Depth:Architect and govern:Low-latency, high-throughput AI networks (InfiniBand, RDMA, RoCE, NVLink awareness, UFM and NetQ)Distributed training network topologiesDefine enterprise-level:Network architecture standards, segmentation, and security modelsScalability and resilience patterns for AI platformsLead optimization of:Network performance for GPU-intensive and distributed workloadsEstablish and drive:Production readiness frameworks across platform, data, model, and security layersSLO/SLI models for latency, availability, throughput, and reliability with error budgetsDefine and publish:Reference architectures for AI workloads (LLM, RAG, distributed inference, batch/stream pipelines)Standardize: Observability patterns (network telemetry, latency tracking, cost-performance correlation)Own: Capacity engineering (network throughput, concurrency, GPU fabric utilization)Define resilience engineering:Failover models, network redundancy, fault isolation, and recovery strategiesEstablish AI security and compliance frameworks:Zero Trust networking, segmentation, secure egress, data protection policies
Key Responsibilities
Skill Depth:Architect and govern:Low-latency, high-throughput AI networks (InfiniBand, RDMA, RoCE, NVLink awareness, UFM and NetQ)Distributed training network topologiesDefine enterprise-level:Network architecture standards, segmentation, and security modelsScalability and resilience patterns for AI platformsLead optimization of:Network performance for GPU-intensive and distributed workloadsEstablish and drive:Production readiness frameworks across platform, data, model, and security layersSLO/SLI models for latency, availability, throughput, and reliability with error budgetsDefine and publish:Reference architectures for AI workloads (LLM, RAG, distributed inference, batch/stream pipelines)Standardize: Observability patterns (network telemetry, latency tracking, cost-performance correlation)Own: Capacity engineering (network throughput, concurrency, GPU fabric utilization)Define resilience engineering:Failover models, network redundancy, fault isolation, and recovery strategiesEstablish AI security and compliance frameworks:Zero Trust networking, segmentation, secure egress, data protection policies
Skill Requirements
Skill RequirementDeep expertise in enterprise networking and Linux platform architectureResponsible for design, governance, scalability, and performance engineering of network and OS layers supporting AI/HPC/Kubernetes platforms at scaleOwns platform standards, reference architectures, and reliability frameworks across multi-cluster and multi-region environmentsCCNA, CCNP preferred with working knowledge on InfiniBand and Mellanix Switches