Job Summary
Role Summary
The Observability Operations Engineer will implement and operate early-stage observability integrations across Kubernetes, Airflow, GCP, Databricks, and Azure environments. This role deploys Datadog agents, integrates Airflow health and pipeline signals, and builds practical Python and SQL checks for Datadog and Dynatrace, with strong attention to telemetry quality, ingest volume, retention, and cost.
Key Responsibilities
Key Responsibilities
- Deploy, configure, upgrade, and troubleshoot Datadog Agent and Cluster Agent in Kubernetes/GKE using Helm and manifests
- Integrate Airflow scheduler, workers, DAGs, task runs, retries, duration, SLA, and failure signals into Datadog and/or Dynatrace
- Build Python and SQL checks/extractors for health, data freshness, pipeline status, utilization, and cost signals
- Configure Datadog and Dynatrace metrics, logs, traces, tags, dashboards, and monitors; validate end-to-end telemetry
- Support integrations across GCP, Databricks, and Azure, with attention to ingest volume, retention, and cost controls
- Partner with SRE and the Datadog Administrator / Observability SME on incident workflows, runbooks, and deployment patterns
- Document reusable integration and operations procedures for owning teams
Skill Requirements
Required Skills
- Lead: Kubernetes/GKE agent deployment, Datadog/Dynatrace operations, Airflow integrations, telemetry troubleshooting
- Strong: Python, SQL, Helm/manifests, logs, metrics, traces, API/integration development, monitoring
- Working Knowledge: GCP(BigQuery, Dataflow, Airflow, Cloud Run, Cloud SQL, Redis, BigTable), Databricks, Azure, OpenTelemetry, cost/right-sizing signals, SLO and alerting concepts