SRE Technical Specialist
Canada
Job Description
SRE Technical Specialist
Toronto, Ontario

Job Summary

The Technical Support Specialist in Site Reliability engineering (SRE) will be responsible for ensuring the reliability and stability of the systems and applications. The role involves providing technical support, troubleshooting issues, and implementing solutions to optimize system performance and availability. The Specialist will play a crucial role in monitoring, analyzing, and enhancing system reliability to meet business objectives effectively.

Key Responsibilities

1. Provide technical support and assistance to ensure the reliability and stability of systems and applications
2. Troubleshoot and resolve technical issues related to system performance and availability
3. Collaborate with cross functional teams to implement and optimize solutions for system reliability
4. Monitor system performance metrics and develop strategies to enhance system reliability
5. Implement automation tools and scripts to improve system monitoring and maintenance processes

Skill Requirements

1. Strong knowledge of site reliability engineering (sre) principles and practices
2. Proficiency in scripting languages like python, shell, or powershell
3. Experience with monitoring tools such as prometheus, grafana, or nagios
4. Familiarity with cloud platforms like aws, azure, or google cloud
5. Excellent problem-solving and analytical skills
6. Strong communication and teamwork abilities
7. Ability to work in a fast paced and dynamic environment

Other Requirements

Core Skills & Tools

  • AWS: Lambda, ECS/Fargate/EC2, API Gateway, SNS/SQS, Kinesis, RDS; IAM/KMS foundations.
  • Observability & ITSM: Dynatrace, CloudWatch, ELK; ServiceNow for incidents/changes; SLI/SLO dashboards.
  • Reliability Practices: Error budgets, capacity/performance benchmarking, automation/runbook execution, FinOps awareness.

1.Relevant certifications in Site Reliability Engineering (SRE) or related fields are a plus

  • Deliver 24×7 monitoring, incident response, and problem management; drive MTTA/MTTR reduction and SLO/SLI adherence.
  • Perform preventive health checks; analyze ticket trends to implement continual service improvements and automation to reduce toil.
  • Execute blameless postmortems and high-quality RCA; maintain SOPs/runbooks and reliability dashboards.
  • Configure/tune observability (Dynatrace, CloudWatch, ELK); enable self-healing workflows and workload optimizations.
  • Support change/service requests within agreed SLAs; collaborate during transitions and onboard new AWS services.

 

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.