Job Summary
Key Responsibilities
2. Troubleshoot and resolve technical issues related to system performance and availability
3. Collaborate with cross functional teams to implement and optimize solutions for system reliability
4. Monitor system performance metrics and develop strategies to enhance system reliability
5. Implement automation tools and scripts to improve system monitoring and maintenance processes
Skill Requirements
1. Strong knowledge of site reliability engineering (sre) principles and practices
2. Proficiency in scripting languages like python, shell, or powershell
3. Experience with monitoring tools such as prometheus, grafana, or nagios
4. Familiarity with cloud platforms like aws, azure, or google cloud
5. Excellent problem-solving and analytical skills
6. Strong communication and teamwork abilities
7. Ability to work in a fast paced and dynamic environment
Other Requirements
Core Skills & Tools
- AWS: Lambda, ECS/Fargate/EC2, API Gateway, SNS/SQS, Kinesis, RDS; IAM/KMS foundations.
- Observability & ITSM: Dynatrace, CloudWatch, ELK; ServiceNow for incidents/changes; SLI/SLO dashboards.
- Reliability Practices: Error budgets, capacity/performance benchmarking, automation/runbook execution, FinOps awareness.
1.Relevant certifications in Site Reliability Engineering (SRE) or related fields are a plus
- Deliver 24×7 monitoring, incident response, and problem management; drive MTTA/MTTR reduction and SLO/SLI adherence.
- Perform preventive health checks; analyze ticket trends to implement continual service improvements and automation to reduce toil.
- Execute blameless postmortems and high-quality RCA; maintain SOPs/runbooks and reliability dashboards.
- Configure/tune observability (Dynatrace, CloudWatch, ELK); enable self-healing workflows and workload optimizations.
- Support change/service requests within agreed SLAs; collaborate during transitions and onboard new AWS services.