Senior Site Reliability Engineer Lead
Canada
Job Description
Senior Site Reliability Engineer Lead
Toronto, Ontario

Job Summary

The Senior Support Lead in Site Reliability engineering (SRE) will be responsible for overseeing the support and reliability operations within the organization. This role will focus on ensuring the stability, performance, and efficiency of the systems while leading a team of support engineers to provide exceptional service.

Key Responsibilities

1. Lead and manage a team of support engineers in resolving incidents, requests, and problems to ensure system uptime and reliability.
2. Collaborate with the engineering and development teams to implement efficient and scalable solutions that enhance system performance.
3. Develop and maintain support documentation, standard operating procedures, and best practices for the support team.
4. Identify opportunities for automation and implement tools to streamline support processes.
5. Monitor system performance and provide recommendations for improvements to optimize system reliability.
6. Participate in on call rotations to address critical incidents and ensure 24/7 system availability.
7. Conduct regular performance evaluations, provide feedback, and mentor team members to promote professional growth.

Skill Requirements

1. In-depth knowledge of site reliability engineering (sre) principles and best practices.
2. Proficiency in system monitoring, incident management, and performance tuning tools.
3. Strong understanding of cloud services, microservices architecture, and containerization technologies.
4. Excellent problem-solving skills and the ability to troubleshoot complex technical issues.
5. Experience with scripting languages (e.g., python, bash) for automation and tool development.
6. Familiarity with agile methodologies and devops practices for continuous integration and delivery.
7. Strong communication and leadership skills to effectively lead a support team and collaborate with cross functional teams.
8. Ability to work under pressure, prioritize tasks, and manage multiple projects simultaneously.

Other Requirements

Key Responsibilities:

  • Deliver 24×7 monitoring, incident response, and problem management; drive MTTA/MTTR reduction and SLO/SLI adherence.
  • Perform preventive health checks; analyze ticket trends to implement continual service improvements and automation to reduce toil.
  • Execute blameless postmortems and high-quality RCA; maintain SOPs/runbooks and reliability dashboards.
  • Configure/tune observability (Dynatrace, CloudWatch, ELK); enable self-healing workflows and workload optimizations.
  • Support change/service requests within agreed SLAs; collaborate during transitions and onboard new AWS services.

1.Relevant certifications in Site Reliability Engineering (SRE) or Cloud Services are a plus.

Core Skills & Tools

  • AWS: Lambda, ECS/Fargate/EC2, API Gateway, SNS/SQS, Kinesis, RDS; IAM/KMS foundations.
  • Observability & ITSM: Dynatrace, CloudWatch, ELK; ServiceNow for incidents/changes; SLI/SLO dashboards.
  • Reliability Practices: Error budgets, capacity/performance benchmarking, automation/runbook execution, FinOps awareness.

 

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.