Job Summary
Job Summary We are seeking a Site Reliability Engineer (SRE) L1 Generalist with 5–7 years of experience in infrastructure operations, cloud services, production support, monitoring, incident management, and reliability-focused operations. The role serves as the first line of operational support for critical services, with responsibility for monitoring, initial diagnosis, runbook-based remediation, evidence collection, ticket documentation, and timely escalation to L2/L3 or specialist teams. The successful candidate will work across infrastructure, cloud, network, platform, database, middleware, and application support teams to improve service availability, operational consistency, and customer experience. The position requires a broad technical foundation, strong troubleshooting skills, disciplined process execution, and an automation-first mindset. Key Responsibilities Production Monitoring and Operations • Monitor infrastructure, applications, cloud services, platforms, and dependent components using enterprise monitoring and observability tools. • Respond to alerts, events, incidents, and service degradation within agreed response targets. • Validate alerts, assess impact and urgency, and perform initial technical diagnosis. • Execute approved standard operating procedures, runbooks, and known-error resolutions. • Monitor service health, availability, performance, capacity, and operational trends. Incident Management and Triage • Act as the first technical responder for production incidents and operational issues. • Analyze logs, metrics, traces, dashboards, system events, and configuration information to isolate probable causes. • Classify incidents accurately by service, technology domain, severity, and business impact. • Resolve incidents within the authorized L1 support boundary or escalate with complete diagnostic evidence. • Coordinate with L2/L3, engineering, vendor, and service-management teams during incident resolution. • Provide clear and timely technical updates and maintain complete incident records. Reliability and Continuous Improvement • Support the measurement and reporting of service-level indicators, service-level objectives, and service-level agreements. • Participate in problem reviews, post-incident reviews, and corrective-action tracking. • Identify recurring alerts, incidents, manual activities, and operational toil. • Recommend runbook, monitoring, process, and automation improvements. • Contribute to service resilience, operational readiness, and knowledge-management initiatives. Automation and Tooling • Develop or enhance basic operational scripts using PowerShell, Python, Bash, or equivalent technologies. • Automate repetitive checks, data collection, health validation, and routine remediation where approved. • Support CI/CD, configuration management, Infrastructure as Code, and self-service initiatives as applicable. • Use version control and established development practices for operational scripts and automation artifacts. Documentation and Operational Readiness • Create and maintain SOPs, runbooks, troubleshooting guides, knowledge articles, and shift handover records. • Document diagnostic steps, evidence, actions, outcomes, and escalation details in the ITSM platform. • Participate in knowledge-transfer, shadow, reverse-shadow, and operational-readiness activities. • Support planned changes, maintenance activities, deployments, disaster-recovery exercises, and on-call operations. Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infr
Key Responsibilities
Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.
Skill Requirements
Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.
Other Requirements
Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.