Job Summary
Skill: Network (Enterprise / AI Infrastructure Context)L3 – Senior Engineer / SMESkill RequirementStrong hands-on expertise in enterprise networking and Linux systemsResponsible for troubleshooting, performance tuning, and operational stability of network and OS layers supporting AI/HPC/Kubernetes workloadsNetworking (L3)Certifications (Preferred):CCNA (Mandatory baseline)CCNP (Strongly preferred)Experience:5–10 years in data center / cloud networking operationsHands-on experience with:Routing & Switching (BGP, OSPF, VLANs)Load balancing, firewall basicsInfiniBand card and switchesNvidia UFMCCNA, CCNP preferred with working knowledge on InfiniBand and Mellanix SwitchesSkill Depth:Strong in incident troubleshooting and RCAAbility to diagnose:Latency, packet loss, connectivity failuresNetwork issues affecting GPU / distributed workloadsWorking knowledge of cloud networking (AWS/Azure/GCP)Working Knowledge of AI/HPC networking for distributed training for GPU-GPU communication.Define and maintain production readiness standards across platform, data, model, application, and security layers.Establish SLO/SLI frameworks for latency, availability, quality, safety, and drift implement error budget policies.Maintain compliance mappings (e.g., ISO 27001, SOC 2, GDPR/DPDP, HIPAA where applicable).Author PRR checklists, runbooks/playbooks, and DR/BCP blueprints (RTO/RPO, multi‑region/site failover). Drive enablement (trainings, brown-bags) and maintain knowledge repositories and decision records.Implement observability (tracing, metrics, logs), dashboards, and SLO burn and cost anomaly alerting.Execute safe releases (canary/shadow/blue green), prompt/model versioning, feature flags, and rollback plans.
Key Responsibilities
Skill Depth:Strong in incident troubleshooting and RCAAbility to diagnose:Latency, packet loss, connectivity failuresNetwork issues affecting GPU / distributed workloadsWorking knowledge of cloud networking (AWS/Azure/GCP)Working Knowledge of AI/HPC networking for distributed training for GPU-GPU communication.Define and maintain production readiness standards across platform, data, model, application, and security layers.Establish SLO/SLI frameworks for latency, availability, quality, safety, and drift implement error budget policies.Maintain compliance mappings (e.g., ISO 27001, SOC 2, GDPR/DPDP, HIPAA where applicable).Author PRR checklists, runbooks/playbooks, and DR/BCP blueprints (RTO/RPO, multi‑region/site failover). Drive enablement (trainings, brown-bags) and maintain knowledge repositories and decision records.Implement observability (tracing, metrics, logs), dashboards, and SLO burn and cost anomaly alerting.Execute safe releases (canary/shadow/blue green), prompt/model versioning, feature flags, and rollback plans.
Skill Requirements
Skill RequirementStrong hands-on expertise in enterprise networking and Linux systemsResponsible for troubleshooting, performance tuning, and operational stability of network and OS layers supporting AI/HPC/Kubernetes workloads
CCNA, CCNP preferred with working knowledge on InfiniBand and Mellanix Switches
Other Requirements
2. Excellent communication and presentation skills. Must be able to clearly communicate with the customer and be able to present solutions to customers at CIO, CXO level
3. Must be open for 24x7 environment