SeniorAdministrator - Monitoring Tools, Event Monitoring
India
Job Description
SeniorAdministrator - Monitoring Tools, Event Monitoring
Hyderabad, Telangana

Job Summary

 Job Summary We are seeking a Site Reliability Engineer (SRE) L1 Generalist with 5–7 years of experience in infrastructure operations, cloud services, production support, monitoring, incident management, and reliability-focused operations. The role serves as the first line of operational support for critical services, with responsibility for monitoring, initial diagnosis, runbook-based remediation, evidence collection, ticket documentation, and timely escalation to L2/L3 or specialist teams. The successful candidate will work across infrastructure, cloud, network, platform, database, middleware, and application support teams to improve service availability, operational consistency, and customer experience. The position requires a broad technical foundation, strong troubleshooting skills, disciplined process execution, and an automation-first mindset. Key Responsibilities Production Monitoring and Operations • Monitor infrastructure, applications, cloud services, platforms, and dependent components using enterprise monitoring and observability tools. • Respond to alerts, events, incidents, and service degradation within agreed response targets. • Validate alerts, assess impact and urgency, and perform initial technical diagnosis. • Execute approved standard operating procedures, runbooks, and known-error resolutions. • Monitor service health, availability, performance, capacity, and operational trends. Incident Management and Triage • Act as the first technical responder for production incidents and operational issues. • Analyze logs, metrics, traces, dashboards, system events, and configuration information to isolate probable causes. • Classify incidents accurately by service, technology domain, severity, and business impact. • Resolve incidents within the authorized L1 support boundary or escalate with complete diagnostic evidence. • Coordinate with L2/L3, engineering, vendor, and service-management teams during incident resolution. • Provide clear and timely technical updates and maintain complete incident records. Reliability and Continuous Improvement • Support the measurement and reporting of service-level indicators, service-level objectives, and service-level agreements. • Participate in problem reviews, post-incident reviews, and corrective-action tracking. • Identify recurring alerts, incidents, manual activities, and operational toil. • Recommend runbook, monitoring, process, and automation improvements. • Contribute to service resilience, operational readiness, and knowledge-management initiatives. Automation and Tooling • Develop or enhance basic operational scripts using PowerShell, Python, Bash, or equivalent technologies. • Automate repetitive checks, data collection, health validation, and routine remediation where approved. • Support CI/CD, configuration management, Infrastructure as Code, and self-service initiatives as applicable. • Use version control and established development practices for operational scripts and automation artifacts. Documentation and Operational Readiness • Create and maintain SOPs, runbooks, troubleshooting guides, knowledge articles, and shift handover records. • Document diagnostic steps, evidence, actions, outcomes, and escalation details in the ITSM platform. • Participate in knowledge-transfer, shadow, reverse-shadow, and operational-readiness activities. • Support planned changes, maintenance activities, deployments, disaster-recovery exercises, and on-call operations. Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infr

Key Responsibilities

Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.

Skill Requirements

Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.

Other Requirements

Required Experience • 5–7 years of experience in SRE, infrastructure operations, cloud operations, NOC, platform support, DevOps operations, or enterprise production support. • Hands-on experience supporting business-critical production environments. • Demonstrated experience in alert handling, incident triage, troubleshooting, escalation, and technical documentation. • Experience working in a 24x7 support model, including rotational shifts or on-call coverage. • Experience collaborating with cross-functional infrastructure, application, cloud, network, database, security, and service-management teams. Required Technical Skills Operating Systems: Working knowledge of Linux and Windows Server administration, system services, processes, logs, file systems, resource utilization, and basic performance troubleshooting. Cloud Platforms: Hands-on exposure to Microsoft Azure, AWS, or Google Cloud, including compute, storage, networking, identity, monitoring, and basic cloud-service troubleshooting. Monitoring and Observability: Experience with tools such as Splunk, Dynatrace, Datadog, Grafana, Prometheus, Azure Monitor, AppDynamics, New Relic, or equivalent platforms. Networking: Understanding of TCP/IP, DNS, DHCP, routing, firewalls, VPN, proxies, load balancers, ports, protocols, and connectivity troubleshooting. ITSM and Operations: Experience with ServiceNow or a comparable ITSM tool, including Incident, Problem, Change, Request, and Knowledge Management processes. Scripting and Automation: Basic to intermediate scripting skills using PowerShell, Python, Bash, Shell, or equivalent technologies. DevOps and Platforms: Working understanding of Git, CI/CD concepts, containers, Kubernetes fundamentals, configuration management, and Infrastructure as Code concepts. Reliability Practices: Understanding of SLIs, SLOs, SLAs, error budgets, operational toil, incident response, post-incident reviews, and continuous improvement. Preferred Qualifications • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline, or equivalent practical experience. • ITIL Foundation certification. • Cloud certification in Azure, AWS, or Google Cloud. • Kubernetes, DevOps, automation, or SRE-related certification. • Experience with enterprise-scale, hybrid-cloud, or distributed technology environments. Professional Competencies • Strong analytical, troubleshooting, and problem-solving skills. • Ability to remain organized and effective during high-priority incidents. • Clear verbal and written communication with technical and nontechnical stakeholders. • Disciplined documentation, ownership, follow-through, and shift-handover practices. • Collaborative approach with a strong customer-service and reliability mindset. • Willingness to learn new technologies and continuously improve operational practices. Key Performance and Success Measures • Timely alert acknowledgment, incident response, and escalation. • Accurate incident classification, evidence collection, and ticket documentation. • Compliance with SOPs, runbooks, change controls, SLAs, and governance requirements. • Quality and effectiveness of first-line diagnosis and remediation. • Reduction of repeat incidents, false alerts, and manual operational toil. • Contribution to automation, knowledge quality, service reliability, and operational improvement.

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.