Job Summary
We are seeking an experienced Prometheus Lead (L2) to lead the architecture, administration, governance, and optimization of the enterprise Prometheus monitoring platform. The candidate will be responsible for designing metrics collection frameworks, monitoring strategies, alerting mechanisms, and platform scalability across on-premises, cloud, hybrid, and containerized environments. The role requires strong expertise in Prometheus architecture, exporters, service discovery, metric storage, alerting, and performance optimization. The candidate will act as the technical owner of the Prometheus platform and work closely with Infrastructure, Cloud, SRE, Application, and Operations teams to deliver proactive monitoring and observability capabilities.
Key Responsibilities
Key Responsibilities • Lead the administration, governance, and lifecycle management of the Prometheus monitoring platform. • Design and maintain enterprise-wide monitoring solutions using Prometheus. • Define monitoring standards, metric collection policies, alerting strategies, and best practices. • Implement scalable monitoring architectures across infrastructure, applications, cloud, databases, and container platforms. • Configure and manage service discovery mechanisms and exporters. • Develop monitoring frameworks to provide visibility into system health, performance, availability, and capacity. • Establish monitoring baselines, thresholds, and alerting standards. • Collaborate with infrastructure, cloud, application, and SRE teams to onboard new services into monitoring. • Drive platform optimization, scalability improvements, and monitoring standardization. • Provide technical leadership and mentorship to monitoring engineers and platform administrators. • Create technical documentation, operational procedures, and knowledge articles.
Skill Requirements
Required Skills & Experience • Strong hands-on experience with Prometheus administration and monitoring architecture. • Deep understanding of metrics collection, monitoring frameworks, and alerting concepts. • Experience with PromQL query development and optimization. • Strong knowledge of Prometheus federation and scalability concepts. • Experience configuring exporters and monitoring integrations. • Knowledge of Linux administration and troubleshooting. • Understanding of infrastructure monitoring, application monitoring, and cloud monitoring. • Experience with containerized environments and Kubernetes monitoring. • Scripting knowledge in: o Python o Bash o PowerShell o YAML • Excellent troubleshooting, analytical, and problem-solving skills. • Strong communication and stakeholder management capabilities. Classification: Internal Job Description HCLTech • Experience leading technical teams and enterprise monitoring initiatives. Preferred Skills • Experience with Kubernetes and OpenShift environments. • Knowledge of Grafana dashboard integration. • Exposure to cloud monitoring technologies (Azure, AWS, GCP). • Experience with ServiceNow integration and ITSM processes. • Knowledge of SRE and observability practices. • Exposure to OpenTelemetry and modern monitoring architectures. • Experience supporting large-scale enterprise monitoring environments. • Understanding of capacity planning and performance engineering.
Other Requirements
Technical Core sExpectations 1. Prometheus Platform Administration • Installation, configuration, upgrade, and maintenance of Prometheus environments. • Configure Prometheus servers, federation, remote storage, and retention policies. • Manage monitoring architecture for high availability and scalability. • Troubleshoot Prometheus performance issues and optimize resource utilization. • Implement backup, recovery, and lifecycle management procedures. 2. Metrics Collection & Monitoring • Design and implement metrics collection strategies across enterprise environments. • Configure scrape jobs, targets, labels, and monitoring policies. • Establish monitoring coverage for: o Windows Servers o Linux Servers o Virtualization Platforms o Network Devices o Databases o Middleware o Cloud Services o Applications • Define performance indicators, health metrics, and operational KPIs. • Ensure monitoring data accuracy and consistency. 3. Exporters & Service Discovery • Deploy, configure, and manage Prometheus exporters. • Administer exporters such as: o Node Exporter o Windows Exporter o Blackbox Exporter o SNMP Exporter o Database Exporters o Application Exporters • Configure dynamic service discovery mechanisms. • Enable monitoring for cloud-native and enterprise infrastructure environments. • Ensure exporter health, optimization, and lifecycle management. 4. Alert Management • Design and implement enterprise alerting strategies. • Define alert rules, thresholds, severity levels, and escalation procedures. • Configure alert routing and notification mechanisms. • Reduce alert noise through tuning and threshold optimization. • Support proactive incident detection and operational response. 5. Performance & Capacity Monitoring • Develop monitoring solutions for performance analysis and capacity planning. • Create dashboards and reports to support operational teams. • Analyze trends and historical metrics to identify risks and bottlenecks. • Support availability reporting, SLA compliance, and performance management initiatives. • Drive continuous monitoring improvements through data-driven insights.