Job Summary
HCLTech is seeking a seasoned Network Reliability Engineer (NRE) with deep expertise in designing, operating, and optimizing hybrid & cloud network environments. You will focus on ensuring high availability, scalability, performance, and resilience of enterprise networks spanning on-premises data centers, SD-WAN, ACI, and major cloud platforms (AWS, Azure). This role combines network engineering with Reliability Engineering principle, emphasizing automation, observability, proactive incident prevention, and rapid recovery to meet stringent SLAs for mission-critical applications.
Key Responsibilities
Reliability & Availability: Design and implement highly reliable network infra. Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets. • Incident Management: Lead major incident response, root cause analysis (RCA), and post-mortem reviews. Implement blameless post-mortems and drive corrective actions to prevent recurrence. • Automation & Self-Healing: Build automation scripts, tools, and self-healing mechanisms for network provisioning, configuration management, monitoring, and failover. • Monitoring & Observability: Develop and maintain comprehensive dashboards, alerts, and logging using tools like Prometheus, Django, Grafana, Datadog, Splunk etc. • Capacity Planning & Performance: Conduct network capacity planning, performance tuning, and chaos engineering to validate resilience. • Security & Compliance: Collaborate on network security posture, zero-trust models, firewall policies, and compliance requirements. • Cross-functional Collaboration: Work with DevOps, SRE, Cloud, Security, and Application teams to embed reliability into the development lifecycle. • Mentorship: Guide junior engineers and contribute to knowledge sharing within the global team.
Skill Requirements
Education: o Bachelor’s or master’s in computer science, Engineering, or related field (or equivalent experience). o Preferred CCIE and/or CCNP certification • Core Networking o Deep expertise in Routing (BGP, OSPF, EIGRP), Switching (STP, VLANs, VXLAN), Firewalls, Load Balancers, VPNs, and SD-WAN. o Strong troubleshooting of complex Layer 2/3/4 issues. • SRE/NRE Practices: o Experience applying SRE principles (error budgets, toil reduction, automation). o Proficiency in scripting (Python) and Infrastructure as Code (Terraform, Ansible, etc.). • Tools & Technologies: o Monitoring: Prometheus, Grafana, Django, Datadog, etc. o CI/CD & Automation: Jenkins, GitOps, Ansible. o Packet analysis: Wireshark, tcpdump. • Soft Skills: o Excellent problem-solving, communication, and stakeholder management. Ability to work in a global, 24x7 on-call rotation.