Job Summary
Role Summary
Provide Level 3 technical leadership and engineering support for Hybrid Cloud infrastructure. Responsible for complex troubleshooting, architecture governance, automation strategy, platform optimization, RCA, and technology lifecycle management.
Key Responsibilities
Technical Skills
- Windows Server 2016/2019/2022, Active Directory, DNS, DHCP, DFS, GPO, ADFS, SCCM
- VMware vSphere, ESXi, vCenter, DRS, HA, SRM
- Nutanix Prism, AHV, AOS, CVM, LCM
- AWS: EC2, VPC, IAM, ELB, EBS, S3, Route53, CloudWatch
- Azure: Virtual Machines, VNET, Storage, Key Vault, Azure AD, Azure Site Recovery
- Patch Management, Vulnerability Remediation, Disaster Recovery and Capacity Management
Automation Skills
- Hands-on experience in Ansible and/or Terraform for infrastructure provisioning, configuration management, and automation.
- Strong scripting skills using PowerShell, Python, and/or Shell Scripting for operational automation.
- Experience developing and maintaining automation workflows, reusable playbooks, and Infrastructure-as-Code (IaC) solutions to improve operational efficiency and reduce manual effort.
Key Responsibilities
- Manage Windows, VMware, Nutanix, AWS and Azure environments.
- Perform incident, problem, change and service request management.
- Execute patching, vulnerability remediation and compliance activities.
- Participate in disaster recovery planning, testing and execution.
- Maintain SOPs, runbooks, CMDB and operational documentation.
- Coordinate with OEM vendors and cloud providers for issue resolution.
- Lead RCA for major incidents and drive permanent fixes.
- Design platform standards, automation frameworks and cloud governance controls.
- Lead lifecycle upgrades, cloud transformations and performance optimization initiatives.
Skill Requirements
Experience
L3: 8-12+ Years
Other Requirements
24x7 Support Coverage
Provide 24x7 operational support, monitoring, incident management, and troubleshooting to ensure service availability and business continuity.
Participate in shift-based support, on-call rotations, and major incident management activities to meet SLA commitments.