Job Summary
• 5–8 years of experience in infrastructure or application monitoring, with 3+ years in Datadog.
• Strong implementation experience across Datadog APM, Logs, Metrics, RUM, and Synthetics modules.
• Hands-on experience with cloud environments (AWS, Azure, or GCP).
• Knowledge of scripting (Python, PowerShell, or Bash) for automation of Datadog configurations.
• Familiarity with CI/CD pipelines, containerized workloads (Docker, Kubernetes), and distributed systems.
• Strong understanding of modern observability practices (metrics, logs, traces).
Key Responsibilities
• Experience integrating Datadog with Terraform or configuration management tools.
• Knowledge of ITIL event and incident processes.
Skill Requirements
• Own Datadog full-stack observability across infrastructure, application, network, and cloud components.
• Fine-tune and optimize existing Datadog monitoring configurations, dashboards, alerts, and synthetics.
• Implement advanced Datadog functionalities such as APM distributed tracing, RUM, log analytics, and CI/CD pipeline integrations.
• Define observability KPIs and SLIs/SLOs aligned to business and IT service health.
• Develop and maintain custom monitors, service maps, and metric dashboards for proactive monitoring.
• Integrate Datadog with incident management systems (ServiceNow, OpsGenie, PagerDuty, etc.).
• Collaborate with developers, SREs, and infrastructure teams to trace, diagnose, and resolve performance bottlenecks.
• Mentor team members and document observability use cases and best practices.
Other Requirements
• Experience integrating Datadog with Terraform or configuration management tools.
• Knowledge of ITIL event and incident processes.