Job Summary
HCLTech is seeking a seasoned Network Reliability Engineer (NRE) with deep expertise in designing, operating, and optimizing hybrid & cloud network environments. You will focus on ensuring high availability, scalability, performance, and resilience of enterprise networks spanning on-premises data centers, SD-WAN, and major cloud platforms (AWS, Azure).
This role combines network engineering with Reliability Engineering principles—emphasizing automation, observability, proactive incident prevention, and rapid recovery to meet stringent SLAs for mission-critical applications.
Key Responsibilities
- Reliability & Availability: Design and implement highly reliable network architectures. Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
- Incident Management: Lead major incident response, root cause analysis (RCA), and post-mortem reviews. Implement blameless post-mortems and drive corrective actions to prevent recurrence.
- Automation & Self-Healing: Build automation scripts, tools, and self-healing mechanisms for network provisioning, configuration management, monitoring, and failover.
- Monitoring & Observability: Develop and maintain comprehensive dashboards, alerts, and logging using tools like Prometheus, Django, Grafana, Datadog, Splunk etc.
- Capacity Planning & Performance: Conduct network capacity planning, performance tuning, and chaos engineering to validate resilience.
- Security & Compliance: Collaborate on network security posture, zero-trust models, firewall policies, and compliance requirements.
- Cross-functional Collaboration: Work with DevOps, SRE, Cloud, Security, and Application teams to embed reliability into the development lifecycle.
- Mentorship: Guide junior engineers and contribute to knowledge sharing within the global team.
Skill Requirements
Required Skills & Experience (15 Years Profile)
- Core Networking (10+ years hands-on):
- Deep expertise in Routing (BGP, OSPF, EIGRP), Switching (STP, VLANs, VXLAN), Firewalls, Load Balancers, VPNs, and SD-WAN.
- Strong troubleshooting of complex Layer 2/3/4 issues.
- SRE/NRE Practices:
- Experience applying SRE principles (error budgets, toil reduction, automation).
- Proficiency in scripting (Python) and Infrastructure as Code (Terraform, Ansible, etc.).
- Tools & Technologies:
- Monitoring: Prometheus, Grafana, Datadog, etc.
- CI/CD & Automation: Jenkins, GitOps, Ansible.
- Packet analysis: Wireshark, tcpdump.
- Soft Skills: Excellent problem-solving, communication, and stakeholder management. Ability to work in a global, 24x7 on-call rotation.
Education: Bachelor’s or Master’s in Computer Science, Engineering, or related field (or equivalent experience).
Preferred Certifications:
- CCIE (Enterprise, Data Center, or Security)/CCNP (R&S)