Job Summary
Monitor infrastructure, applications, cloud services, and endpoint health using DataDog dashboards, alerts, and observability tools. • Investigate and troubleshoot alerts, performance degradation, and service availability issues. • Create and maintain dashboards, monitors, reports, and service health views. • Analyze logs, metrics, traces, and events to identify trends and root causes of incidents. • Support alert tuning, noise reduction, threshold optimization, and monitoring standardization initiatives. • Collaborate with ServiceNow, infrastructure, cloud, and application teams to support incident resolution and operational improvements • Assist in observability integrations with ServiceNow ITSM/ITOM, Nexthink, cloud monitoring platforms, and enterprise tools. • Perform operational reporting, trend analysis, capacity monitoring, and service performance reviews. • Support implementation of runbooks, automation workflows, and operational best practices.
Key Responsibilities
We are seeking a DataDog & Observability Engineer (L2) to support enterprise monitoring, observability, and operational intelligence platforms. The role will focus on monitoring infrastructure, applications, endpoints, and digital services using DataDog, while ensuring proactive detection, analysis, and resolution of performance and availability issues. The engineer will work closely with operations, infrastructure, cloud, and service management teams to improve system visibility, reduce incidents, and enhance service reliability through observability best practices
Skill Requirements
Good to Have Requirements • Experience with ServiceNow ITSM/ITOM, Event Management, and CMDB integrations. • Exposure to Nexthink, ThousandEyes, Splunk, Dynatrace, Grafana, or Azure Monitor. • Knowledge of cloud platforms such as Azure, AWS, or GCP. • Basic scripting experience using PowerShell, Python, or Shell scripting.
Other Requirements
Experience & Skill Requirements • 3–6 years of experience in Monitoring, Infrastructure Operations, NOC, Cloud Operations, or Observability Support. • Hands-on experience with DataDog administration, monitoring, dashboards, and alert management. • Strong understanding of observability concepts including: o Metrics o Logs o Traces o Events o Telemetry Management • Experience in performance, availability, and capacity monitoring.