Senior Technical Lead - Post Silicon Validation
United States
Job Description
Senior Technical Lead - Post Silicon Validation
Santa Clara, California

Job Summary

You will be responsible for developing, executing, and maintaining tests for collective communication operations (AllReduce, AllGather, ReduceScatter, AllToAll) on custom networking silicon — validating that the ASIC correctly enables large-scale distributed AI training workloads.

 

Programming Languages

Language

Proficiency

Usage

C/C++

Expert (Must-have)

Writing test frameworks, RDMA verbs test suites, driver-level test development, loopback and traffic tests

Python

Strong (Must-have)

Test automation, CI/CD integration, orchestration of multi-node test scenarios, emulation test infrastructure

Bash/Shell Scripting

Proficient (Good-to-have)

Test execution scripts, environment setup, multi-host coordination

 

Key Responsibilities

  1. Develop test suites for collective operations (AllReduce, AllGather, ReduceScatter, AllToAll) targeting Trantor ASIC across emulation, FPGA, and silicon platforms
  2. Write RDMA verbs-level tests using the RoCE Verbs Testing Framework (rdma-core Verbs API) — covering positive, negative, and error-injection scenarios
  3. Validate multi-node, multi-NIC collective communication patterns, ensuring correct behavior under various topologies (rail-aligned, cross-rail, multiplanar)
  4. Develop traffic generation and validation tools for RDMA collectives at scale — covering data integrity, performance, and error handling
  5. Integrate tests into CI/CD pipelines for regression prevention on every code change and nightly builds
  6. Collaborate with driver, firmware, architecture, and modeling teams to define test plans and ensure complete coverage of networking features
  7. Run and analyze performance benchmarks (NCCL-tests, perftest, rdma_gen) to identify regressions and validate throughput/latency targets
  8. Debug and root-cause failures across the full stack — ASIC RTL, firmware, driver, rdma-core provider, and user-space collectives

 

Skill Requirements

Must-Have

  • RDMA (Remote Direct Memory Access) — Deep understanding of RDMA operations: READ, WRITE, SEND, RECEIVE; Queue Pairs (QPs), Completion Queues (CQs), Memory Regions (MRs), Protection Domains (PDs)
  • RoCE v2 — Understanding of RDMA over Converged Ethernet protocols, transport-level behavior, and conformance requirements
  • Collective Communication Operations — AllReduce, AllGather, ReduceScatter, AllToAll; ring/tree algorithms; understanding of how collectives map to network traffic patterns
  • Ethernet / L2 Networking — Layer 2 fundamentals, MTU, multiport networking, VLANs
  • PCIe Architecture — PCIe endpoint/switch topology, Gen5/Gen6, BAR regions, MSI-X interrupts, SR-IOV, multi-function devices
  • NCCL / Communication Libraries — Familiarity with NVIDIA Collective Communications Library or equivalent; understanding of how training jobs use collectives over RDMA NICs

Good-to-Have

  • InfiniBand / IB Verbs API — Experience with libibverbs, rdma-core, ibv_* APIs
  • Network Topologies for AI Training — Rail-optimized fabrics, fat-tree, multi-planar designs, PXN
  • Traffic Congestion & Flow Control — PFC (Priority Flow Control), ECN, congestion management for lossless fabrics
  • DMA & Memory Subsystems — GDR (GPUDirect RDMA), host memory registration, IOMMU
  • Protocol Conformance Testing — ANVL or similar automated conformance testing methodologies

 

Other Requirements

Preferred Qualifications

  • Background in silicon validation for networking chips (switches, NICs, DPUs)
  • Experience with pre-silicon validation environments (emulation, FPGA prototyping, software models/QEMU)
  • Familiarity with infrastructure or large-scale hyperscaler networking
  • Contributions to open-source RDMA/networking projects (rdma-core, Linux kernel networking, NCCL)
  • Understanding of AI/ML training workloads and how network performance impacts training efficiency
  • Experience with build systems and test infrastructure at scale

 

Education

  • B.S./M.S. in Computer Science, Electrical Engineering, Computer Engineering, or related field
  • Advanced degree preferred but not required with equivalent industry experience
Maximum Salary (US):  148000
Minimum Salary (US):  78000
Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion. 

 

Compensation and Benefits

A candidate’s pay within the range will depend on their skills, experience, education, and other factors permitted by law. This role may also be eligible for performance-based bonuses subject to company policies. In addition, this role is eligible for the following benefits subject to company policies: medical, dental, vision, pharmacy, life, accidental death & dismemberment, and disability insurance; employee assistance program; 401(k) retirement plan; 10 days of paid time off per year (some positions are eligible for need-based leave with no designated number of leave days per year); and 10 paid holidays per year.