Job Summary
GKE SME
Key Responsibilities
GKE Architecture & Cluster Management: Design, deploy, and scale enterprise GKE environments (Standard and Autopilot modes). Manage multi-cluster control planes, node pool configurations, auto-scaling mechanisms (Horizontal/Vertical Pod Autoscalers), and zero-downtime cluster upgrades.Container Networking & Ingress Architecture: Architect highly secure container networking structures within Google Cloud. Design and tune GCP VPC-native clusters, GKE Ingress controllers, Gateway API, internal/external HTTP(S) Load Balancing, and Service Mesh technologies (e.g., Anthos Service Mesh / Istio).GitOps & Pipeline Automation: Establish global standards for declarative application management. Build and maintain GitOps workflows using tools like ArgoCD or Flux, and integrate infrastructure deployments using Terraform or Config Connector.Platform Security & Governance: Enforce strict cloud-native security controls. Implement GKE Workload Identity, Kubernetes RBAC integrated with IAM, Network Policies (Calico/Cilium), Binary Authorization pipelines, and continuous pod security standards audits.Observability & Performance Engineering: Design comprehensive telemetry dashboards and alert thresholds. Leverage Google Cloud Backup for GKE along with Cloud Logging and Cloud Monitoring (Stackdriver) or Prometheus/Grafana to track cluster health, optimize node utilization, and minimize MTTR.Cost Optimization & Resource Allocation: Analyze compute and memory footprint metrics to optimize infrastructure costs. Design efficient multi-tenancy models using namespaces, resource quotas, and limit ranges to eliminate wasted cloud spend.
Skill Requirements
Experience: 8+ years of professional experience spanning systems engineering or DevOps, with at least 4+ years of dedicated, hands-on experience engineering and scaling production GKE (Google Kubernetes Engine) environments.Google Cloud Mastery: Deep engineering knowledge of GCP core services, including VPCs, Cloud IAM, Interconnect/Cloud VPN, Cloud Storage, and Compute Engine underlays.Kubernetes Internals: Expert-level understanding of Kubernetes core architecture, scheduling mechanics, storage primitives (CSI drivers, Persistent Volumes), and container runtimes (containerd).Infrastructure as Code (IaC): High proficiency in writing modular, clean Terraform configurations to provision complex GCP and GKE infrastructure.Automation & Scripting: Strong scripting capabilities using Python, Go, or heavy Bash to interact with the Google Cloud SDK (gcloud) and Kubernetes API.
Other Requirements
NA