Track Manager - High Performance Computing, Red Hat Enterprise Linux
India
Job Description
Track Manager - High Performance Computing, Red Hat Enterprise Linux
Bengaluru, Karnataka

Job Summary

Key Responsibilities • Own Linux infrastructure lifecycle: provisioning, configuration, patching, upgrades, and capacity planning. • Design & implement: Server provisioning (Kickstart, PXE, Foreman, HA/DR architectures (e.g., Pacemaker; multi-AZ/multi-region patterns). • Automate everything: golden images, OS baselines, config compliance (Puppet/Ansible pipelines). • Build & run : Golden Image creations with lifecycle management and push to Infrastructure. • Security hardening at OS and network layers (SELinux, firewalld/nftables/IPSec, CIS baselines). • Performance engineering: low-latency tuning, kernel parameters, NUMA, I/O optimization. • Observability: Prometheus/Grafana dashboards knowledge. • Lead incident response and postmortems with blameless RCAs and concrete action plans. • Mentor engineers; contribute to architecture standards, design reviews, and runbooks. • Partner with Security, Networking, App, and DevOps teams on cross-functional deliverables. Mandatory Qualifications • 15+ years Linux engineering in production at scale (RHEL). • Expert in systemd, boot process, LVM, XFS/ext4, iSCSI/multipath, RAID; tcpdump, iproute2. • Strong Bash and Python; comfortable with regex and writing robust automation. • Hands-on with experts role in Puppet, Ansible, Cloud Migration (Azure and AWS), Git, and CI/CD tools (GitHub Actions). • Infrastructure as Code with Puppet and Ansible. • Security: SELinux, SSH hardening, vulnerability scanning (OpenSCAP/Qualys). • Cloud: experience with AWS/Azure, VPC/networking, IAM, On-prem to cloud migrations and managed K8s desirable. • Demonstrated leadership in complex troubleshooting, incident management, and mentoring. Success Metrics (First 6–12 Months) • MTTR reduced by ≥30% through runbooks/automation. • Baseline OS compliance ≥95% within 90 days; patch SLAs consistently met. • Automation coverage (≥70% of repetitive ops codified). • Documentation (golden images, infra patterns, puppet modules, playbooks) published and adopted. • Cloud Migration on-prem to AWS/Azure

Key Responsibilities

Key Responsibilities • Own Linux infrastructure lifecycle: provisioning, configuration, patching, upgrades, and capacity planning. • Design & implement: Server provisioning (Kickstart, PXE, Foreman, HA/DR architectures (e.g., Pacemaker; multi-AZ/multi-region patterns). • Automate everything: golden images, OS baselines, config compliance (Puppet/Ansible pipelines). • Build & run : Golden Image creations with lifecycle management and push to Infrastructure. • Security hardening at OS and network layers (SELinux, firewalld/nftables/IPSec, CIS baselines). • Performance engineering: low-latency tuning, kernel parameters, NUMA, I/O optimization. • Observability: Prometheus/Grafana dashboards knowledge. • Lead incident response and postmortems with blameless RCAs and concrete action plans. • Mentor engineers; contribute to architecture standards, design reviews, and runbooks. • Partner with Security, Networking, App, and DevOps teams on cross-functional deliverables. Mandatory Qualifications • 15+ years Linux engineering in production at scale (RHEL). • Expert in systemd, boot process, LVM, XFS/ext4, iSCSI/multipath, RAID; tcpdump, iproute2. • Strong Bash and Python; comfortable with regex and writing robust automation. • Hands-on with experts role in Puppet, Ansible, Cloud Migration (Azure and AWS), Git, and CI/CD tools (GitHub Actions). • Infrastructure as Code with Puppet and Ansible. • Security: SELinux, SSH hardening, vulnerability scanning (OpenSCAP/Qualys). • Cloud: experience with AWS/Azure, VPC/networking, IAM, On-prem to cloud migrations and managed K8s desirable. • Demonstrated leadership in complex troubleshooting, incident management, and mentoring. Success Metrics (First 6–12 Months) • MTTR reduced by ≥30% through runbooks/automation. • Baseline OS compliance ≥95% within 90 days; patch SLAs consistently met. • Automation coverage (≥70% of repetitive ops codified). • Documentation (golden images, infra patterns, puppet modules, playbooks) published and adopted. • Cloud Migration on-prem to AWS/Azure

Skill Requirements

Key Responsibilities • Own Linux infrastructure lifecycle: provisioning, configuration, patching, upgrades, and capacity planning. • Design & implement: Server provisioning (Kickstart, PXE, Foreman, HA/DR architectures (e.g., Pacemaker; multi-AZ/multi-region patterns). • Automate everything: golden images, OS baselines, config compliance (Puppet/Ansible pipelines). • Build & run : Golden Image creations with lifecycle management and push to Infrastructure. • Security hardening at OS and network layers (SELinux, firewalld/nftables/IPSec, CIS baselines). • Performance engineering: low-latency tuning, kernel parameters, NUMA, I/O optimization. • Observability: Prometheus/Grafana dashboards knowledge. • Lead incident response and postmortems with blameless RCAs and concrete action plans. • Mentor engineers; contribute to architecture standards, design reviews, and runbooks. • Partner with Security, Networking, App, and DevOps teams on cross-functional deliverables. Mandatory Qualifications • 15+ years Linux engineering in production at scale (RHEL). • Expert in systemd, boot process, LVM, XFS/ext4, iSCSI/multipath, RAID; tcpdump, iproute2. • Strong Bash and Python; comfortable with regex and writing robust automation. • Hands-on with experts role in Puppet, Ansible, Cloud Migration (Azure and AWS), Git, and CI/CD tools (GitHub Actions). • Infrastructure as Code with Puppet and Ansible. • Security: SELinux, SSH hardening, vulnerability scanning (OpenSCAP/Qualys). • Cloud: experience with AWS/Azure, VPC/networking, IAM, On-prem to cloud migrations and managed K8s desirable. • Demonstrated leadership in complex troubleshooting, incident management, and mentoring. Success Metrics (First 6–12 Months) • MTTR reduced by ≥30% through runbooks/automation. • Baseline OS compliance ≥95% within 90 days; patch SLAs consistently met. • Automation coverage (≥70% of repetitive ops codified). • Documentation (golden images, infra patterns, puppet modules, playbooks) published and adopted. • Cloud Migration on-prem to AWS/Azure

Other Requirements

Key Responsibilities • Own Linux infrastructure lifecycle: provisioning, configuration, patching, upgrades, and capacity planning. • Design & implement: Server provisioning (Kickstart, PXE, Foreman, HA/DR architectures (e.g., Pacemaker; multi-AZ/multi-region patterns). • Automate everything: golden images, OS baselines, config compliance (Puppet/Ansible pipelines). • Build & run : Golden Image creations with lifecycle management and push to Infrastructure. • Security hardening at OS and network layers (SELinux, firewalld/nftables/IPSec, CIS baselines). • Performance engineering: low-latency tuning, kernel parameters, NUMA, I/O optimization. • Observability: Prometheus/Grafana dashboards knowledge. • Lead incident response and postmortems with blameless RCAs and concrete action plans. • Mentor engineers; contribute to architecture standards, design reviews, and runbooks. • Partner with Security, Networking, App, and DevOps teams on cross-functional deliverables. Mandatory Qualifications • 15+ years Linux engineering in production at scale (RHEL). • Expert in systemd, boot process, LVM, XFS/ext4, iSCSI/multipath, RAID; tcpdump, iproute2. • Strong Bash and Python; comfortable with regex and writing robust automation. • Hands-on with experts role in Puppet, Ansible, Cloud Migration (Azure and AWS), Git, and CI/CD tools (GitHub Actions). • Infrastructure as Code with Puppet and Ansible. • Security: SELinux, SSH hardening, vulnerability scanning (OpenSCAP/Qualys). • Cloud: experience with AWS/Azure, VPC/networking, IAM, On-prem to cloud migrations and managed K8s desirable. • Demonstrated leadership in complex troubleshooting, incident management, and mentoring. Success Metrics (First 6–12 Months) • MTTR reduced by ≥30% through runbooks/automation. • Baseline OS compliance ≥95% within 90 days; patch SLAs consistently met. • Automation coverage (≥70% of repetitive ops codified). • Documentation (golden images, infra patterns, puppet modules, playbooks) published and adopted. • Cloud Migration on-prem to AWS/Azure

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.