Job Summary
Tool SME
Key Responsibilities
Observability Architecture & Strategy: Design, build, and maintain the enterprise-wide architecture for monitoring, alerting, and log aggregation across hybrid-cloud, containerized, and on-premises environments.APM & Infrastructure Monitoring: Deploy and optimize Full-Stack Observability and Application Performance Monitoring (APM) tools (e.g., Dynatrace, Datadog, New Relic, ScienceLogic) to track transaction traces, infrastructure health, and end-user experiences.Synthetic & Real-User Monitoring: Establish proactive digital experience monitoring (DEM) strategies by programming complex synthetic user journeys and real-user monitoring (RUM) tests to alert on application degradation before customers feel the impact.Alert Orchestration & Noise Reduction: Architect and fine-tune intelligent alerting configurations and synthetic thresholds. Integrate monitoring tools with incident orchestration platforms (e.g., PagerDuty, Opsgenie, xMatters) to reduce alert fatigue and route high-fidelity notifications directly to
Skill Requirements
Experience: 7+ years of dedicated experience engineering, managing, and architecting enterprise infrastructure tools, monitoring platforms, or observability frameworks.Platform Expertise: Advanced, internal-level technical engineering knowledge of at least one premium enterprise visibility stack (most notably Dynatrace, Datadog, ScienceLogic, or New Relic).Telemetry Protocols: Deep understanding of modern telemetry frameworks and standards, including OpenTelemetry (OTel), Prometheus metrics, SNMP, WMI, SSH, and raw log parsing.API Ingestion & Webhooks: High proficiency in leveraging REST/SOAP APIs to extract custom data sets, push external metrics, and configure cross-platform automation pipelines.Automation & Scripting: Strong proficiency in scripting languages (Python, PowerShell, or Bash) along with experience automating monitoring agent deployments through tools like Ansible, Terraform, or SCCM.
Other Requirements
NA