Senior Technical Lead
United Arab Emirates
Job Description
Senior Technical Lead
Others, Abu Dhabi

Job Summary

Job Title: Senior Engineer – Kubernetes Administration, EDGE AI CoE – T&I Cluster Job Purpose Design, build, secure, and operate on‑premises Kubernetes platforms hosting AI/ML workloads across the group. Own end‑to‑end cluster lifecycle (provisioning through decommission), GPU enablement, performance, reliability, and cost efficiency to support training, inference, data pipelines, and MLOps at scale, assuming on-prem deployments without managed cloud services. Key Responsibilities Cluster Architecture & Build (On‑Prem) Design and provision production‑grade clusters on bare‑metal and/or virtualized environments (e.g., kubeadm‑based; OpenShift, Rancher or upstream distributions), including HA control planes, etcd health, and node pools. Implement secure bootstrapping in connected and air‑gapped environments (private registries, image mirrors, OS hardening, CIS benchmarks). Configure networking: CNI (e.g., Cilium/Calico), ingress (NGINX/HAProxy), internal/external load‑balancing (MetalLB/BGP/ECMP), DNS/CoreDNS, IPv4/IPv6, and, where needed, Multus for secondary interfaces. Lifecycle Operations Plan and execute upgrades, patching, and capacity management; manage taints/affinities, quotas, and multi‑tenancy isolation. Implement autoscaling/right‑sizing (HPA/VPA; cluster autoscaling via on‑prem integrations) and node lifecycle automation. Maintain DR strategy, backups, and restores (e.g., Velero/restic), with documented runbooks and regular tests. AI/ML Workload Enablement Enable and manage specific workloads within each Kubernetes deployment that needs access to GPUs. Operate model serving (e.g., KServe/Seldon), batch/streaming pipelines, and ML orchestration (Kubeflow/Argo Workflows); optimize for throughput/latency and GPU utilization. Storage & Data Services Manage CSI drivers and dynamic provisioning; integrate with Ceph/Longhorn/Rook/NetApp/NFS as applicable; support stateful workloads with proper SLAs. Security & Compliance Enforce RBAC, Pod Security, NetworkPolicies, image signing/scanning (e.g., cosign/Trivy), admission controls/policy (OPA Gatekeeper/Kyverno), secrets management (Vault/Sealed Secrets/KMS). Integrate enterprise identity (LDAP/AD/Keycloak OIDC), certificate management /PKI, and audit logging; maintain evidence for audits. Platform Engineering & GitOps Standardize delivery with Helm/Kustomize; implement GitOps (Argo CD/Flux) and IaC (Terraform/Ansible) for repeatable environments. Operate on‑prem CI/CD runners (e.g., GitLab Runners/K8s executors) and artifact registries (Harbor) including air‑gapped workflows. Observability & SRE Implement metrics, logs, tracing (Prometheus/Grafana, EFK/ELK/Loki, OpenTelemetry); define SLIs/SLOs and capacity/health dashboards. Lead incident response, root cause analysis, and continuous improvement; participate in an on‑call rotation. Education & Experience Bachelor’s in Computer Science, Engineering, or related field (or equivalent experience). 7–10 years in platform/SRE/DevOps/systems roles, including 4+ years administering production Kubernetes. Proven experience designing, building, and operating on‑prem Kubernetes clusters (kubeadm/upstream; Rancher, OpenShift acceptable) in data‑center environments—no reliance on managed services. Hands‑on with GPUs for AI workloads, air‑gapped operations, private registries, and enterprise identity/security integrations. Certifications: CKA required (or obtained within 6 months); CKS/CKAD preferred; relevant Linux/networking certifications a plus. Key Skills Kubernetes core: control plane ops, etcd care, scheduling, admission controllers, multi‑tenancy (namespaces, quotas), PDBs, affinities, taints/tolerations. Linux/Containers: containerd/CRI‑O, systemd, kernel/GPU drivers, filesystems, OS hardening. Networking: Cilium/Calico, NetworkPolicies, ingress/egress; optional service mesh (Istio/Linkerd). Storage: CSI, PV/PVC, snapshots, performance tuning, backup/restore, DR patterns; Ceph/Rook/NetApp/NFS familiarity.

Key Responsibilities

  • Security: RBAC, Pod Security, image scanning/signing, OPA/Kyverno, Vault/Sealed Secrets, OIDC/LDAP/AD, audit and compliance practices.
  • Observability: Prometheus/Grafana, EFK/ELK/Loki, OpenTelemetry; capacity/performance tuning.
  • Delivery/IaC: Helm/Kustomize, Argo CD/Flux, Terraform/Ansible; GitLab CI/GitHub Actions; scripting in Bash/Python; strong YAML hygiene.
  • Soft skills: ownership, clear documentation, incident communication, stakeholder management, mentoring.

Education & Experience

  • Bachelor’s in Computer Science, Engineering, Data/AI, or related field; Master’s preferred.
  • 8+ years in software/ML engineering with 3+ years building GenAI/LLM solutions and production services.
  • Hands-on expertise with on-prem/air-gapped deployments, Kubernetes, service meshes, and secure networking.
  • Demonstrated experience with data engineering at scale: streaming/batch pipelines, metadata/lineage, governance.
  • Practical LLMOps/MLOps experience: CI/CD, model registries, evaluation, monitoring/observability.
  • Work in regulated, defense, or mission-critical environments preferred; eligibility for security clearance is a must.

Key Skills

  • Software/Platform: Python/Go/TypeScript, microservices, APIs, gRPC/REST, containerization, Kubernetes (plus kubectl and Helm).
  • Observability & Reliability: tracing/logging/metrics, SLO/SLA, autoscaling, HA/DR, incident management, cost governance.
  • Security & Compliance: data-flow controls, access management, encryption, secrets management, PDPL/NIST-aligned practices.
  • LLMOps/MLOps: model/prompt versioning, evaluation pipelines, canary/AB testing, drift detection, feedback loops.
  • Performance Engineering: latency/throughput tuning, memory/cost optimization, GPU/accelerator pipelines.
  • Collaboration: technical leadership, design reviews, mentoring, clear documentation, and stakeholder communication.

Skill Requirements

Other Requirements

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.